Refactoring Legacy Code with AI Coding Agents: A Practical Approach
Most AI coding content focuses on greenfield development. Build a new feature, scaffold a new project, generate a fresh component. That is the easy case. The hard case — the one that consumes far more engineering time — is the codebase that already exists. The one with ten years of history, undocumented conventions, and implicit contracts between modules that nobody fully understands.
Refactoring legacy code with AI agents is different from writing new code with them. The agent is not creating in a vacuum. It is working inside a system with constraints it cannot see, assumptions it was not told about, and behavior that depends on things that were never made explicit. Get the workflow wrong and the agent will cheerfully modernize your codebase into something that does not work.
Get it right and you have a collaborator that can process repetitive changes across hundreds of files without getting bored, losing focus, or introducing the kind of inconsistencies that happen when a human is on hour six of a migration.
Why Legacy Refactoring Is Hard for Agents
Before reaching for an AI agent, understand the failure modes.
Agents do not know about implicit contracts. A function might return null in a specific edge case, and seventeen callers depend on that behavior. No documentation says so. No test covers it. The agent sees a function that should probably throw an error instead of returning null, and "fixes" it. The fix is locally correct and globally catastrophic.
Style inconsistency is invisible risk. Legacy codebases often have multiple coding styles because they were written by different people over many years. An agent might normalize everything to one style in a file it touches, creating a diff that mixes the actual refactor with cosmetic noise. Reviewing a 200-line diff where 180 lines are formatting changes and 20 are logic changes is worse than reviewing 20 lines of pure logic change.
Dependency chains are deep. Changing one module often requires changes in files that import it, files that import those files, and so on. An agent that refactors a database access module might not realize it also needs to update the integration test helpers, the CLI scripts, and the admin panel that all use the same internal API.
Tests may not exist. The biggest problem. Without tests, there is no automated way to verify that a refactor preserved behavior. The agent cannot run a check. You cannot run a check. You are both guessing.
The Workflow: Safe Refactoring With AI Agents
Step 1: Establish a Safety Net
Do not start refactoring without a verification mechanism. If the codebase has tests, make sure they pass before you change anything. If it does not, write tests first.
This is where AI agents are immediately useful. Before touching the legacy code, use the agent to generate characterization tests — tests that capture current behavior regardless of whether that behavior is correct.
Working in src/billing/invoice-calculator.ts:
Before making any changes to this file, write characterization tests
in src/billing/__tests__/invoice-calculator.test.ts that capture the
current behavior of every exported function.
Include edge cases:
- zero-value invoices
- negative line items
- missing tax configuration
- invoices with more than 100 line items
Run the tests to verify they all pass against the current code.
Do not modify the source file.
The key constraint: "do not modify the source file." You want a test suite that locks in current behavior first. Only then is it safe to refactor.
Step 2: Scope Ruthlessly
The most common refactoring failure with AI agents is scope creep. You ask for one change and the agent improves six related things you did not ask about.
Scope every refactoring task with explicit boundaries:
Migrate the logging calls in src/billing/ from console.log to
the structured logger in src/lib/logger.ts.
Constraints:
- Only modify files in src/billing/
- Replace console.log, console.warn, and console.error calls
- Use logger.info, logger.warn, and logger.error respectively
- Preserve the exact log message strings
- Do not change any other logic in these files
- Do not rename variables, reformat code, or fix unrelated issues
After each file, run: npm test -- --testPathPattern=billing
Six constraints. Each one prevents a specific failure mode. Without them, the agent might rename variables for consistency, fix a linting issue it noticed, or restructure a function while it is in there. Each of those changes is individually reasonable and collectively a nightmare to review.
Step 3: Work in Layers
Refactoring a legacy codebase is not one task. It is a sequence of small, independently verifiable changes. Each layer should leave the codebase in a working state.
Layer 1: Infrastructure. Set up the new pattern alongside the old one. If you are migrating from callbacks to promises, create the promise-based version of the core functions but do not change any callers yet.
Create promise-based wrappers for the callback functions exported
by src/data/legacy-db.ts. Put them in src/data/db.ts.
Do not modify legacy-db.ts.
Do not change any code that imports legacy-db.ts.
Export the same function names with the same parameters,
but returning promises instead of accepting callbacks.
Add tests that verify the promise wrappers produce identical
results to the callback originals.
Layer 2: Migration. Move callers from old to new, one module at a time.
Migrate src/api/routes/users.ts from importing src/data/legacy-db
to importing src/data/db (the promise-based version).
Replace callback-style calls with async/await.
Keep the same route behavior and response format.
Run: npm test -- --testPathPattern=users
Layer 3: Cleanup. Once all callers are migrated, remove the old code.
Verify that src/data/legacy-db.ts has no remaining importers.
Run: grep -r "legacy-db" src/
If no imports remain, delete legacy-db.ts and its tests.
Run the full test suite.
Each layer is a clean commit. Each is reviewable independently. If something goes wrong in layer 2, you can revert to layer 1 without losing the infrastructure work.
Step 4: Use Multiple Agent Sessions
Large refactors benefit from parallel agent sessions. Not parallel edits to the same files — that creates conflicts — but parallel work on independent modules.
If you are migrating logging across the entire codebase:
- Session 1: Migrate
src/billing/ - Session 2: Migrate
src/auth/ - Session 3: Migrate
src/api/
Each session works in its own directory scope, runs its own tests, and produces its own commits. You review each independently and merge them together.
This is where terminal organization becomes important. Managing three concurrent agent sessions plus a test runner plus your git workflow requires either a well-configured tmux setup or a terminal designed for multi-session work. Jumping between unlabeled terminal tabs while trying to remember which agent is working on billing versus auth is how mistakes happen.
Step 5: Verify at Every Level
Run tests after every change, not just at the end. The verification sequence should escalate from narrow to broad:
# After each file change
npm test -- --testPathPattern=billing/invoice
# After each module is done
npm test -- --testPathPattern=billing
# After the full migration
npm test
npm run build
npm run lint
If a narrow test fails, you know exactly which change caused it. If you wait until the end, you have a 40-file diff and a failing test suite with no clear attribution.
Common Refactoring Tasks and Agent Prompts
Dependency Replacement
Replacing one library with another across the codebase.
Replace all uses of moment.js with date-fns in src/.
For each file:
1. Replace moment imports with equivalent date-fns imports
2. Convert moment() calls to date-fns function calls
3. Preserve existing date format strings where possible
4. Add a comment if a format string needed to change
Do not change test files yet. We will migrate tests separately
after verifying the source changes pass existing tests.
Constraints:
- One file at a time
- Run tests after each file
- Stop and report if a conversion is ambiguous
API Contract Modernization
Moving from REST callbacks to async handlers.
Convert the Express route handler in src/api/routes/reports.ts
from callback style to async/await.
Current pattern:
router.get('/reports', (req, res, next) => {
getReports(req.query, (err, data) => {
if (err) return next(err);
res.json(data);
});
});
Target pattern:
router.get('/reports', async (req, res, next) => {
try {
const data = await getReports(req.query);
res.json(data);
} catch (err) {
next(err);
}
});
Keep error handling behavior identical.
Run route-specific tests after.
Configuration Extraction
Moving hardcoded values to configuration.
Scan src/services/ for hardcoded configuration values:
- URLs and hostnames
- Timeout values
- Retry counts
- Feature flags
List them with file paths and line numbers.
Do not make any changes yet. Just produce the inventory.
Start with discovery. Let the agent scan and report. Then decide which values to extract and where to put them. Do not let the agent make that architectural decision autonomously on a legacy codebase.
When Not to Use Agents for Refactoring
Agents are not a universal tool for legacy code. Some situations require human judgment that agents cannot reliably provide.
Architectural redesign. If the refactoring involves changing the fundamental structure of the application — splitting a monolith, changing data flow patterns, restructuring the module hierarchy — the agent should not drive those decisions. Use the agent for the mechanical work after you have decided the target architecture.
Ambiguous behavior. If the code does something and nobody knows whether that behavior is intentional, an agent cannot tell you either. It will guess, and its guess will sound confident. When you encounter "mystery behavior" in legacy code, investigate with the agent but make the keep-or-remove decision yourself.
Performance-critical paths. A refactored version that is functionally identical but 3x slower is not a successful refactoring. Agents do not naturally consider performance characteristics unless you explicitly tell them to benchmark.
The Realistic Expectation
AI coding agents do not magically modernize legacy codebases. What they do is collapse the most tedious part of the work: the repetitive, file-by-file, line-by-line mechanical changes that make refactoring feel like a grind.
The architectural decisions, the scoping, the verification strategy, and the judgment calls about ambiguous behavior — that is still your job. The agent handles the volume. You handle the direction.
For a solo developer or a small team facing a large legacy codebase, that division of labor changes what is feasible. A migration that would take two weeks of grinding can be done in two days of focused work: a few hours of planning and scoping, and then letting the agents execute while you review.
That is not magic. It is leverage. And on legacy code, leverage is what you need most.
Working on a large refactoring with multiple agent sessions? Agents UI provides persistent, project-organized terminal sessions designed for exactly this workflow.