Google used large language models to migrate tens of thousands of code locations internally, and reported that 80% of the changes in one migration were fully AI-authored, cutting total migration time by an estimated 50%. A controlled trial published six months later found experienced developers were 19% slower when allowed to use AI tools, while believing they had been 20% faster.

Both results are real. AI legacy code modernization is the use of large language models and agents to read, translate, refactor and test legacy code. Whether it accelerates your programme depends on factors such as task type, test coverage, code complexity, team expertise and verification practices.

This guide sets out what the research actually found, the conditions that separate the two outcomes, the tool categories, a workflow with defensible quality gates, and how to measure whether AI is helping you rather than assuming it is.

What is AI Legacy Code Modernization?

AI legacy code modernization is the use of large language models, agents and code-transformation tooling to analyse, document, translate, refactor and test existing code that has become costly or difficult to maintain manually. It spans comprehension of unfamiliar code, mechanical transformation at scale, test generation, and migration between languages or frameworks.

AI is one route through a wider modernization decision, and our legacy application modernization guide covers the approaches it sits inside.

The important distinction is between comprehension and transformation. Comprehension is reading a codebase nobody understands and producing an explanation, a dependency map, or a specification of what a module actually does. Transformation is changing the code itself.

Comprehension is often a lower-risk starting point for legacy code modernization using ai, because engineers can review generated explanations before using them to guide changes. Transformation carries more risk, because a plausible-looking change that alters behaviour can pass review and reach production.

What the Evidence Shows about Legacy Code Modernization using AI

Much of the discussion relies on vendor claims. Three useful sources provide a more grounded view, and their findings differ.

Google’s migrations, and what the authors admit: In How is Google using AI for internal code migrations?, published January 2025, engineers reported concrete results across several internal migrations. In an Int32 to Int64 migration, 80% of the code modifications in landed changelists were fully AI-authored, and total migration time fell by an estimated 50%. A JUnit3 to JUnit4 migration changed 5,359 files and more than 149,000 lines of code in three months, with around 87% of AI-generated code committed without alteration.

The authors are careful about what this proves. They describe the work as “an experience report” and state plainly that they “do not carry out comparisons against other approaches or evaluate research questions/hypotheses.” They also note that “code reviews and change rollouts still require a human operator.”

The controlled trial found the opposite: In July 2025, METR published a randomised controlled trial of 16 experienced open-source developers working on 246 real issues in large, mature repositories, using frontier models of the time. Developers predicted AI would make them 24% faster. They were measured as 19% slower. After the study, having done the work, they still estimated AI had made them 20% faster.

METR is equally careful about scope. They state the result does not show that AI fails to speed up most developers, and note their setting involved unusually high-quality codebases and experienced maintainers.

The organisational picture: The 2025 DORA research found that 90% of technology professionals now use AI at work and more than 80% believe it has increased their productivity, while 30% report little or no trust in AI-generated code. Higher AI adoption correlated with increases in both delivery throughput and delivery instability. DORA’s framing is that AI acts as an amplifier: it magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones.

What the three together mean: AI delivered large, measurable gains on mechanical, well-specified, heavily tested transformations at enormous scale. It did not speed up experienced engineers doing judgment-heavy work in complex unfamiliar territory. And the people doing the work were badly wrong about which situation they were in, by a margin of roughly 39 percentage points.

Where AI for Legacy Code Modernization Works, and Where it Does Not

The pattern across the evidence is consistent enough to plan against. The question is not whether ai for legacy code modernization works, but whether your specific task has the properties that make it work.

Three properties predict success: the change is mechanical and well-specified rather than requiring judgment per instance, an automated check can verify the result, and the same change repeats often enough that setup cost is recovered. Google’s migrations had all three. The METR tasks had none of them.

Task AI suitability Why
Explaining undocumented code High A wrong explanation is caught by the human reading it. No production risk.
Generating documentation and dependency maps High Output is reviewed before use, errors are cheap
Mechanical, repetitive transformation at scale High The Google pattern: one well-specified change, applied thousands of times, verified by existing tests
Generating tests for untested code Medium-high Valuable, but tests written from current behaviour encode existing bugs as expected results
Language translation, for example COBOL to Java Medium Can generate syntactically plausible target code, but semantic accuracy and maintainability require substantial validation
Refactoring with behaviour preservation Medium Depends entirely on test coverage to verify nothing moved
Judgment-heavy changes in complex systems Low The METR setting. Review and correction overhead can exceed the time saved
Deciding what to modernize and in what order Low Requires business context the model does not have

Table: Eight legacy modernization tasks ranked by how reliably AI helps, based on where the published evidence is strongest and weakest.

Read the table as a sequencing tool. Start where suitability is high, build confidence and measurement, and move down only once you can tell the difference between working and not working.

The Precondition for Successful Legacy Code Modernization

Google’s migrations worked because the changed files could be built and their unit tests run automatically, with the model attempting repairs when builds or tests failed. The verification loop is what made AI-authored change safe to land at that volume.

Legacy code often lacks enough automated verification to safely support large-scale automated changes. A common challenge in legacy estates is insufficient automated test coverage, which makes large-scale changes harder to verify.

This is the central practical problem with AI-driven legacy system modernization, and most articles on the subject skip it. AI can generate transformations far faster than a human can verify them by hand. Without automated verification, you have not removed the bottleneck; you have moved it to review, and made it worse by increasing the volume of change arriving there.

The sequencing follows directly. Use AI first to build the safety net: generate characterisation tests that capture what the code currently does, have engineers review those tests against business expectations, and only then use AI to change the code the tests now protect. Inverting this order can increase review and correction work, while the METR result shows that developers may underestimate that overhead.

Legacy Code Modernization AI Tools: What to Use and When

The legacy code modernization ai tools market splits into three categories that fail in different ways, and picking the wrong category for the work is a common and expensive error.

Category What it is Best used for Main limitation
Deterministic codemods Rule-based transformation tools, AST-driven, no model involved Known, repeatable syntactic changes Cannot handle anything requiring judgment or context
AI assistants In-editor models suggesting or writing code under direct supervision Comprehension, test generation, single-file work Output quality drops as context grows beyond what the developer can hold
AI agents Multi-step systems that plan, edit across files, run tests and iterate Repetitive multi-file transformation with a verification loop Failure modes can be harder to detect because agents perform multiple actions across files and tools before presenting the final result

Table: Three tool categories for AI legacy code modernization, with the work each suits and the way each one fails.

Deterministic codemods are underrated for legacy work. Where a transformation can be expressed as a rule, a codemod does it identically every time and the result is reviewable as a rule rather than as thousands of individual diffs. Consider ai agents for legacy code modernization when individual changes require contextual judgment that deterministic rules cannot reliably express.

Step-By-Step AI Legacy Code Modernization Process

This eight-step process puts verification before transformation, so every AI-generated change can be checked automatically before it lands.

1. Freeze the behavioural contract

Establish what the system does now, from its actual behaviour rather than its documentation. Capture inputs and outputs for the paths that matter.

2. Generate characterisation tests, then review them

AI can be particularly useful here, making characterisation-test generation a strong starting point for many modernization programmes. Engineers must review the tests against business expectations, because tests derived from current behaviour will faithfully encode current bugs.

3. Establish the verification loop before any transformation

Build, test and lint must run automatically on changed files. Without this, stop here and fix it first.

4. Start with one narrow, repetitive transformation

Pick something mechanical with a clear success criterion, so you learn how your tooling behaves on your codebase before the stakes rise.

5. Keep change sets small enough to review properly

Google explicitly split results into smaller sets to avoid overwhelming reviewers. Volume defeats review, and defeated review is where defects enter.

6. Review AI changes as changes, not as AI output

Same standard as any other commit. A reviewer who assumes the model was right provides no gate at all.

7. Measure against the baseline you took in step three

Compare cycle time, rework rate and review load against the pre-AI numbers, per task type rather than across the whole programme. Self-reported speed is not evidence, and the next section sets out what to track instead.

8. Keep a human decision point on rollout

Google’s own report notes that code reviews and change rollouts still require a human operator, at their scale and maturity.

How to Measure whether AI Legacy Code Modernization is Actually Helping

This is the section most articles on legacy code modernization leave out, and the METR result makes it the most important one on the page. Developers in a controlled trial were wrong about their own speed by roughly 39 percentage points, in the direction of optimism, after doing the work. Self-report is not evidence.

Measure four things, against a baseline captured before AI is introduced.

  • Cycle time per unit of change, end to end: From work starting to change landing in production, including review and rework. Measuring only generation time is how teams conclude AI is helping when it is not: generation gets faster while review and correction absorb the difference.
  • Rework rate: What proportion of AI-assisted changes needed correction after review or after landing. A rising rework rate against a falling generation time is the specific signature of the METR effect.
  • Escaped defects and instability: DORA found AI adoption correlating with higher throughput and higher instability. Track change failure rate alongside speed, because the two moving together is a known pattern, not bad luck.
  • Review load: Hours spent reviewing per change landed. If AI triples the volume arriving at review and review capacity is unchanged, the constraint has simply moved and the queue is now somewhere less visible.
Metric Baseline to capture Warning signal What it means
End-to-end cycle time Median time from work starting to change in production, by task type Flat or rising while generation time falls Review and rework are absorbing the gains
Rework rate Proportion of changes needing correction after review or after landing Rising alongside faster generation The METR pattern: speed traded for correction load
Change failure rate Failures per change reaching production Rising with throughput The DORA instability pattern, not bad luck
Review hours per change landed Reviewer time per merged change Rising, or review queue lengthening Constraint has moved to review rather than disappearing

Table: Four measurements for AI-assisted modernization, each with the baseline to capture before starting and the signal that AI is shifting work rather than removing it.

Run the comparison per task type rather than across the programme. The evidence says AI helps enormously on some categories and hurts on others, so a blended average across both will show a small, uninformative number and hide the decision you actually need to make.

One further discipline is worth the effort: record what people expected before each phase, alongside what was measured. The gap between the two is the most useful governance signal available, and on the METR evidence it will be large.

What AI Legacy Code Modernization Costs

The cost here has an unusual shape. Tooling may be only part of the overall investment, with engineering preparation, verification and review often requiring significant effort.

The real cost sits in the preconditions and the verification: building the automated test coverage that legacy code lacks, standing up the build and test loop, and the engineering review time that every AI-generated change still requires. On a codebase with no test suite, expect the majority of early effort to go into the safety net rather than into transformation.

For broader software modernization projects, our guide puts the modernization scope between $20,000 and $400,000 or more, depending on complexity. An ai-powered legacy code modernization service can shift spending from manual transformation toward testing, verification and engineering review, particularly during the first engagement.

Challenges of AI Legacy Code Modernization and their Solutions

Five problems account for most failed AI modernization efforts. Each has a practical control, and none of them is a better prompt.

Challenge 1: Plausible wrongness: The failure mode is not code that breaks the build. It is code that compiles, passes weak tests, looks correct in review, and behaves differently in one edge case discovered months later.

Solution: Gate on behaviour rather than on appearance. Characterisation tests written before transformation, run automatically on every change, catch what review does not. Where coverage is thin, restrict AI to comprehension work until the tests exist.

Challenge 2: Behaviour drift during translation: Cross-language migration is where this is concentrated. Semantics that differ subtly between languages, around numeric precision, date handling, null semantics and error propagation, are exactly where models produce confident and wrong output.

Solution: Run both implementations against the same inputs and compare outputs, rather than reviewing the translated code by eye. Build the comparison harness before the translation starts, and treat the known-divergent areas as a specific test checklist.

Challenge 3: Review capacity as the real constraint: Generation scales; human review does not. Any plan that increases change volume without increasing review capacity has relocated the bottleneck rather than removed it.

Solution: Cap the volume of AI-generated change entering review per cycle, and measure review hours per change landed alongside delivery speed. If the queue lengthens, the constraint has moved rather than disappeared.

Challenge 4: Trust without verification: DORA found 30% of practitioners report little or no trust in AI-generated code, while more than 80% believe AI has made them more productive. Holding both positions at once is reasonable, and it means governance cannot rely on practitioner confidence as a signal.

Solution: Base governance on measured outcomes rather than on team sentiment. Record what was expected before each phase alongside what was measured, and use the gap as the review signal. Our AI consulting practice covers AI readiness assessment and AI governance consulting for exactly this.

Challenge 5: Data and IP exposure: Legacy code often contains embedded credentials, customer data in test fixtures, and proprietary business logic.

Solution: Confirm where the code is processed, how it is retained, whether it can be used for model training, and what access controls apply before the first file leaves your environment. Scan for embedded secrets and scrub test fixtures as a precondition, not as a follow-up.

When to Bring in Outside Support for AI Legacy Code Modernization

Organisations without in-house experience in characterisation testing, verification pipelines or AI tooling often bring in external support for the first engagement. The evidence above suggests what to look for in that support. A good partner starts with comprehension and test coverage before transformation. It measures outcomes against a baseline instead of relying on self-reported speed. It keeps change sets small enough to review properly and leaves a human decision point on roll out. Given the data and IP risks covered earlier, it’s also worth asking how and where your code will be processed and retained.

The work usually spans several disciplines. Legacy software modernization services typically include application modernization consulting, software re-engineering, code refactoring, data modernization and API upgradation. An AI development practice covers model selection and integration, which determines how AI tooling fits into the verification loop. Where assessment shows that a rebuild makes more sense than a refactor, a custom software development practice handles the new build.

SparxIT works across these areas, drawing on 19+ years of software engineering and a decade of AI development, including more than 20 production-level AI use cases.

Product Design

Partner with Experts

Frequently Asked Questions

Can AI modernize legacy code without rebuilding the entire application?

open-icon close-icon

Yes, and incremental use is where it performs best. AI can document modules, generate tests around existing behaviour, and transform code section by section while the application keeps running. Full rewrites can increase verification risk, because they remove the existing implementation that teams could use as a behavioural reference.

What types of legacy code can AI modernize?

open-icon close-icon

Models handle widely represented languages best, including Java, C#, JavaScript, Python and PHP. COBOL, RPG and PL/I are supported but with less reliable idiomatic output, because training data is thinner. Practical suitability depends more on test coverage and code structure than on the language itself.

Is AI-generated code safe to deploy in a regulated environment?

open-icon close-icon

It can be, under the same controls as any other change: traceable review, test evidence, and an audit trail showing who approved what. The additional questions regulators ask concern where code was sent for processing and what the retention terms were. Settle both before the first file leaves your environment.

How long does AI-assisted legacy code modernization take?

open-icon close-icon

Timelines depend on whether the verification loop already exists. On a codebase with automated builds and good test coverage, mechanical transformations can move very quickly, as Google’s 149,000-line migration in three months shows. On an untested codebase, expect the early phase to go into building test coverage before any transformation starts.

How do you stop AI from changing behaviour while modernizing legacy code?

open-icon close-icon

Capture current behaviour as characterisation tests before any transformation, review those tests against business expectations, and run them automatically on every change. Keep changesets small enough to review properly. Behaviour preservation is a verification problem rather than a prompting problem, and it is solved by the test suite rather than by the model.