Refactoring and legacy modernization π§
Legacy code refactoring is the use case where the agent is most useful and most dangerous at the same time. Useful because it can analyze hundreds of lines in seconds and propose migration patterns. Dangerous because refactoring without tests is a bug with a different name. The golden rule: safety net first, scissors second.
Refactoring without tests is like repairing a bridge while someone is driving on it.
When to use the agent for this case π―
- Legacy module modernization β migrating from an obsolete pattern to a modern one.
- Interface extraction β defining clear contracts from monolithic code.
- Complexity reduction β breaking giant functions into small, testable units.
- Incremental migration β strangler fig: wrapping the legacy with new layers, one piece at a time.
When NOT to use the agent β
- There are no regression tests β without a safety net, every modification is a gamble. Add tests first (see testing and TDD), then refactor.
- The legacy is critical for production β if a bug can cause financial losses, refactoring must be planned with robust integration tests.
- You do not understand the observable behavior β if you do not know what the legacy code does, you cannot verify that the refactoring preserves behavior.
Opening prompt π
“Refactor [module] preserving behavior. Safety net: [tests to pass]. Constraints: do not modify [public APIs / external behavior]. Plan for incremental sub-tasks.”
Concrete example:
“Refactor
legacy/auth.gopreserving observable behavior. Safety net: all tests inauth_test.gomust pass. Constraints: do not modify the HTTP API contract, do not change JWT token format. Plan: extract dependencies first, then introduce interfaces, finally replace implementation.”
Context setup π§
- Identify existing tests β they are your safety net. If there are none, create a minimal regression test suite first.
- Document observable behavior β what does the module do? What inputs, what outputs? What are the edge cases?
- Define constraints β what must NOT change? Public APIs, data formats, external behavior.
Workflow π‘
- Plan by sub-task: each sub-task must be individually verifiable. One sub-task = one issue = one acceptance criterion.
- Extend test coverage before modifying code. The more tests you have, the less risk you take.
- Execute in micro-increments: after each modification, run the test suite. If something fails, stop and analyze.
- Use plan mode for every sub-task that crosses multiple modules. Context must be focused.
- Consider delegating to sub-agents for independent tasks β it reduces duration and context consumption of the main session.
The strangler fig pattern is your ally: do not rewrite everything at once. Wrap the legacy with new layers, one piece at a time, until the old one can be removed.
Characterization testing: the real safety net π§ͺ
Before refactoring anything, you must lock down the current behavior β bugs included. Characterization tests (or golden master tests) do not document how the code should work, but how it works now, bugs and all.
The generation cycle operates in 3 phases:
- Scaffolding: the agent creates tests with placeholder assertions (e.g.,
assert result == None). - Execution: the test fails, revealing the legacy code’s actual output.
- Capture: the agent captures the actual output and converts it into the immutable baseline.
A test that captures the current behavior β bugs and all β is more valuable than a test that describes ideal behavior. The first tells you when something changes. The second tells you nothing.
To verify behavioral equivalence post-refactoring, use a dual-run harness: legacy and refactored code executed in parallel against the same characterization suite. Equivalence is proved only when output, state mutations, and emitted events exactly match the golden masters.
Database schema evolution ποΈ
Database migrations are the most critical failure point. The Expand and Contract pattern operates in 4 sequential phases:
| Phase | Objective | Mechanism |
|---|---|---|
| Expand | Non-breaking addition | New nullable columns/tables, default values |
| Dual-Write + Backfill | Concurrent writes to both schemas | Async job with rate limiting |
| Reader Migration | Read from new schema via feature flag | Verification with characterization tests |
| Contract | Legacy schema removal | Drop columns, NOT NULL constraints |
The rule is: never combine an application change with a schema migration in the same deploy. Each phase must be independently deployable and verifiable.
Anti-patterns in agentic refactoring π«
| Anti-pattern | Symptom | Remediation |
|---|---|---|
| Business logic hallucination | Agent “cleans up” code that actually implements complex business rules | Document every behavior in plain language before refactoring |
| Disguised big-bang rewrite | PR with thousands of lines, impossible to review | 400-line PR limit; decompose into atomic sub-tasks |
| Feature + refactoring pollution | Mixed commits: structural refactor + new feature | Dedicated refactor branches; merge before starting features |
| Excessive fragmentation | Too-small microservices, distributed monolith | Start with modular monolith; extract only on clear domain boundaries |
Task decomposition and agent-sized units π§©
Every refactoring task must be an atomic unit with an explicit contract:
|
|
Isolated environments (like OpenHands sandboxes) execute tasks in containers, preventing out-of-scope modifications. For large-scale refactoring, an Architect Agent analyzes dependency graphs and assigns independent modules to Worker Agents in separate sandboxes.
LSTs vs long-context LLMs π¬
For precise code transformation, enterprise platforms combine agentic reasoning with Lossless Semantic Trees (LSTs). LSTs represent code at full compiler fidelity, retaining type bindings, formatting, and transitive dependency graphs β without requiring compilation during iterations.
Long-context LLMs (1M+ tokens) are useful for high-level architectural mapping, but suffer from retrieval degradation on large prompts and lack compiler-level precision for variable scoping and type attribution.
Acceptance criteria β
- Regression suite green at every increment.
- No change in observable behavior.
- No unplanned public API changes.
- Each sub-task independently verifiable.
Further reading π
- Testing and TDD β the fundamental safety net for refactoring.
- Adding features β the complementary use case.
- Surviving sessions β how to handle errors and loops during refactoring.