~/wiki

Architecture Migration

Confiance : high
architecture-migrationsystem-evolutioncode-consolidationdependency-managementproduction-safetybranching-strategyrollback-capabilityautonomous-modules

Systematic process of evolving software architectures from complex, interdependent systems to cleaner, more maintainable structures. Particularly critical when consolidating multiple system versions or preparing for team handovers.

Core Methodology

Migration Planning

Dependency Mapping: Before touching any code, create a comprehensive map of all imports, shared components, and cross-module dependencies. This reveals the true scope of changes needed.

Target Architecture Definition: Design the end-state architecture with clear principles - autonomous modules, simplified interfaces, reduced line count, eliminated legacy code paths.

Phase-Based Execution: Break the migration into logical phases that can be implemented incrementally with rollback points at each stage.

Branch Strategy for Safe Migration

Isolation Principle: Create a dedicated migration branch from the stable production state. This provides complete isolation from ongoing development work.

Naming Convention: Use descriptive branch names like clean/v3-handover that clearly indicate the migration purpose and target state.

Commit Granularity: Commit by logical phases rather than individual files to enable meaningful rollback points if issues arise.

Implementation Patterns

Dependency Internalization

Rather than refactoring shared dependencies, copy and simplify required components directly into the target module:

# Before: Complex shared library
src/shared/
├── embedder.py     (425 lines, handles 5 providers)
├── reranker.py     (450 lines, supports 3 algorithms)  
└── config.py       (941 lines, manages 50+ settings)

# After: Internalized and simplified
src/autonomous_module/
├── embedder.py     (200 lines, 2 providers needed)
├── reranker.py     (120 lines, 1 algorithm used)
└── config.py       (200 lines, 15 essential settings)

Benefits:

  • Eliminates complex import paths
  • Enables aggressive simplification for actual use cases
  • Removes risk of breaking changes in shared components
  • Reduces cognitive load for new developers

Component Consolidation

Merge related functionality that was previously separated for reusability but adds complexity:

# Before: Multiple related modules
query_processor.py      (479 lines)
intent_classifier.py    (400 lines)  
acronym_expander.py     (200 lines)

# After: Unified functionality  
query_processor.py      (300 lines, merged capabilities)

Legacy Code Archival

Archive Strategy: Move deprecated code to archive directories rather than deleting:

src/_archive/
├── rag/         # Original implementation
├── rag_v2/      # Second iteration  
└── rag_v3/      # Complex third version

Consumer Migration: Update all UI components and entry points to use the new autonomous module exclusively, eliminating references to archived systems.

Real-World Case Study: RAG System Migration

The assistant-rh project provides a comprehensive example of architecture migration under time pressure:

Initial State:

  • 3 parallel RAG implementations with overlapping functionality
  • Complex interdependencies preventing clean deployment
  • 5700+ lines spread across multiple modules
  • 34 UI pages importing from different system versions

Migration Process:

Phase 1: Autonomous Core Creation

  • Built rag_v3_clean with internalized dependencies
  • Consolidated 879 lines (query_processor + intent_classifier) → 300 lines
  • Simplified embedder from 425 → 200 lines (removed unused providers)
  • Created unified LLM client (100 lines) replacing multiple implementations

Phase 2: Consumer Updates

  • Rewrote main chatbot UI (4500 → 1500-2000 lines)
  • Updated admin pages to use autonomous module exclusively
  • Migrated evaluation tools to new architecture

Phase 3: Archive and Cleanup

  • Moved src/rag/, src/rag_v2/, src/rag_v3/ to src/_archive/
  • Categorized UI pages: 4 essential for production, 30 moved to development folder
  • Archived ~147 notebooks to reduce repository size

Results:

  • 53% reduction in active codebase size (5700 → 2700 lines)
  • Zero dependencies on archived modules
  • Complete functional preservation with improved performance
  • Dramatically simplified onboarding for next engineer

Migration Challenges

Incomplete Dependency Discovery

Problem: Missing a subtle import dependency that only manifests in specific code paths or production conditions.

Solution: Use dependency analysis tools and comprehensive testing across all code paths before declaring migration complete.

Consumer Update Complexity

Problem: UI and integration code may have deep dependencies on the old architecture that aren't immediately obvious.

Solution: Plan for consumer updates to potentially require significant rewrites rather than simple import path changes.

Feature Regression Risk

Problem: Aggressive simplification may accidentally remove functionality that wasn't obviously used but is actually required.

Solution: Maintain comprehensive test coverage and feature inventories throughout migration process.

Success Metrics

Code Reduction: Measure actual lines of code eliminated while preserving functionality. 30-50% reductions are typical for successful migrations.

Dependency Elimination: Count imports from external modules. Target is zero imports from deprecated systems.

Onboarding Time: Measure how long it takes a new engineer to understand and modify the migrated system compared to the original.

Deployment Simplicity: Compare the number of steps and potential failure points in deployment before and after migration.

See also