Google used large language models to migrate tens of thousands of code locations internally, and reported that 80% of the changes in one migration were fully AI-authored, cutting total migration time by an estimated 50%. A controlled trial published six months later found experienced developers were 19% slower when allowed to use AI tools, while believing they had been 20% faster.
Both results are real. AI legacy code modernization is the use of large language models and agents to read, translate, refactor and test legacy code. Whether it accelerates your programme depends on factors such as task type, test coverage, code complexity, team expertise and verification practices.
This guide sets out what the research actually found, the conditions that separate the two outcomes, the tool categories, a workflow with defensible quality gates, and how to measure whether AI is helping you rather than assuming it is.
AI legacy code modernization is the use of large language models, agents and code-transformation tooling to analyse, document, translate, refactor and test existing code that has become costly or difficult to maintain manually. It spans comprehension of unfamiliar code, mechanical transformation at scale, test generation, and migration between languages or frameworks.
AI is one route through a wider modernization decision, and our legacy application modernization guide covers the approaches it sits inside.
The important distinction is between comprehension and transformation. Comprehension is reading a codebase nobody understands and producing an explanation, a dependency map, or a specification of what a module actually does. Transformation is changing the code itself.
Comprehension is often a lower-risk starting point for legacy code modernization using ai, because engineers can review generated explanations before using them to guide changes. Transformation carries more risk, because a plausible-looking change that alters behaviour can pass review and reach production.
Much of the discussion relies on vendor claims. Three useful sources provide a more grounded view, and their findings differ.
Google’s migrations, and what the authors admit: In How is Google using AI for internal code migrations?, published January 2025, engineers reported concrete results across several internal migrations. In an Int32 to Int64 migration, 80% of the code modifications in landed changelists were fully AI-authored, and total migration time fell by an estimated 50%. A JUnit3 to JUnit4 migration changed 5,359 files and more than 149,000 lines of code in three months, with around 87% of AI-generated code committed without alteration.
The authors are careful about what this proves. They describe the work as “an experience report” and state plainly that they “do not carry out comparisons against other approaches or evaluate research questions/hypotheses.” They also note that “code reviews and change rollouts still require a human operator.”
The controlled trial found the opposite: In July 2025, METR published a randomised controlled trial of 16 experienced open-source developers working on 246 real issues in large, mature repositories, using frontier models of the time. Developers predicted AI would make them 24% faster. They were measured as 19% slower. After the study, having done the work, they still estimated AI had made them 20% faster.
METR is equally careful about scope. They state the result does not show that AI fails to speed up most developers, and note their setting involved unusually high-quality codebases and experienced maintainers.
The organisational picture: The 2025 DORA research found that 90% of technology professionals now use AI at work and more than 80% believe it has increased their productivity, while 30% report little or no trust in AI-generated code. Higher AI adoption correlated with increases in both delivery throughput and delivery instability. DORA’s framing is that AI acts as an amplifier: it magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones.
What the three together mean: AI delivered large, measurable gains on mechanical, well-specified, heavily tested transformations at enormous scale. It did not speed up experienced engineers doing judgment-heavy work in complex unfamiliar territory. And the people doing the work were badly wrong about which situation they were in, by a margin of roughly 39 percentage points.
The pattern across the evidence is consistent enough to plan against. The question is not whether ai for legacy code modernization works, but whether your specific task has the properties that make it work.
Three properties predict success: the change is mechanical and well-specified rather than requiring judgment per instance, an automated check can verify the result, and the same change repeats often enough that setup cost is recovered. Google’s migrations had all three. The METR tasks had none of them.
| Task | AI suitability | Why |
| Explaining undocumented code | High | A wrong explanation is caught by the human reading it. No production risk. |
| Generating documentation and dependency maps | High | Output is reviewed before use, errors are cheap |
| Mechanical, repetitive transformation at scale | High | The Google pattern: one well-specified change, applied thousands of times, verified by existing tests |
| Generating tests for untested code | Medium-high | Valuable, but tests written from current behaviour encode existing bugs as expected results |
| Language translation, for example COBOL to Java | Medium | Can generate syntactically plausible target code, but semantic accuracy and maintainability require substantial validation |
| Refactoring with behaviour preservation | Medium | Depends entirely on test coverage to verify nothing moved |
| Judgment-heavy changes in complex systems | Low | The METR setting. Review and correction overhead can exceed the time saved |
| Deciding what to modernize and in what order | Low | Requires business context the model does not have |
Table: Eight legacy modernization tasks ranked by how reliably AI helps, based on where the published evidence is strongest and weakest.
Read the table as a sequencing tool. Start where suitability is high, build confidence and measurement, and move down only once you can tell the difference between working and not working.
Google’s migrations worked because the changed files could be built and their unit tests run automatically, with the model attempting repairs when builds or tests failed. The verification loop is what made AI-authored change safe to land at that volume.
Legacy code often lacks enough automated verification to safely support large-scale automated changes. A common challenge in legacy estates is insufficient automated test coverage, which makes large-scale changes harder to verify.
This is the central practical problem with AI-driven legacy system modernization, and most articles on the subject skip it. AI can generate transformations far faster than a human can verify them by hand. Without automated verification, you have not removed the bottleneck; you have moved it to review, and made it worse by increasing the volume of change arriving there.
The sequencing follows directly. Use AI first to build the safety net: generate characterisation tests that capture what the code currently does, have engineers review those tests against business expectations, and only then use AI to change the code the tests now protect. Inverting this order can increase review and correction work, while the METR result shows that developers may underestimate that overhead.
The legacy code modernization ai tools market splits into three categories that fail in different ways, and picking the wrong category for the work is a common and expensive error.
| Category | What it is | Best used for | Main limitation |
| Deterministic codemods | Rule-based transformation tools, AST-driven, no model involved | Known, repeatable syntactic changes | Cannot handle anything requiring judgment or context |
| AI assistants | In-editor models suggesting or writing code under direct supervision | Comprehension, test generation, single-file work | Output quality drops as context grows beyond what the developer can hold |
| AI agents | Multi-step systems that plan, edit across files, run tests and iterate | Repetitive multi-file transformation with a verification loop | Failure modes can be harder to detect because agents perform multiple actions across files and tools before presenting the final result |
Table: Three tool categories for AI legacy code modernization, with the work each suits and the way each one fails.
Deterministic codemods are underrated for legacy work. Where a transformation can be expressed as a rule, a codemod does it identically every time and the result is reviewable as a rule rather than as thousands of individual diffs. Consider ai agents for legacy code modernization when individual changes require contextual judgment that deterministic rules cannot reliably express.
This eight-step process puts verification before transformation, so every AI-generated change can be checked automatically before it lands.
Establish what the system does now, from its actual behaviour rather than its documentation. Capture inputs and outputs for the paths that matter.
AI can be particularly useful here, making characterisation-test generation a strong starting point for many modernization programmes. Engineers must review the tests against business expectations, because tests derived from current behaviour will faithfully encode current bugs.
Build, test and lint must run automatically on changed files. Without this, stop here and fix it first.
Pick something mechanical with a clear success criterion, so you learn how your tooling behaves on your codebase before the stakes rise.
Google explicitly split results into smaller sets to avoid overwhelming reviewers. Volume defeats review, and defeated review is where defects enter.
Same standard as any other commit. A reviewer who assumes the model was right provides no gate at all.
Compare cycle time, rework rate and review load against the pre-AI numbers, per task type rather than across the whole programme. Self-reported speed is not evidence, and the next section sets out what to track instead.
Google’s own report notes that code reviews and change rollouts still require a human operator, at their scale and maturity.
This is the section most articles on legacy code modernization leave out, and the METR result makes it the most important one on the page. Developers in a controlled trial were wrong about their own speed by roughly 39 percentage points, in the direction of optimism, after doing the work. Self-report is not evidence.
Measure four things, against a baseline captured before AI is introduced.
| Metric | Baseline to capture | Warning signal | What it means |
| End-to-end cycle time | Median time from work starting to change in production, by task type | Flat or rising while generation time falls | Review and rework are absorbing the gains |
| Rework rate | Proportion of changes needing correction after review or after landing | Rising alongside faster generation | The METR pattern: speed traded for correction load |
| Change failure rate | Failures per change reaching production | Rising with throughput | The DORA instability pattern, not bad luck |
| Review hours per change landed | Reviewer time per merged change | Rising, or review queue lengthening | Constraint has moved to review rather than disappearing |
Table: Four measurements for AI-assisted modernization, each with the baseline to capture before starting and the signal that AI is shifting work rather than removing it.
Run the comparison per task type rather than across the programme. The evidence says AI helps enormously on some categories and hurts on others, so a blended average across both will show a small, uninformative number and hide the decision you actually need to make.
One further discipline is worth the effort: record what people expected before each phase, alongside what was measured. The gap between the two is the most useful governance signal available, and on the METR evidence it will be large.
The cost here has an unusual shape. Tooling may be only part of the overall investment, with engineering preparation, verification and review often requiring significant effort.
The real cost sits in the preconditions and the verification: building the automated test coverage that legacy code lacks, standing up the build and test loop, and the engineering review time that every AI-generated change still requires. On a codebase with no test suite, expect the majority of early effort to go into the safety net rather than into transformation.
For broader software modernization projects, our guide puts the modernization scope between $20,000 and $400,000 or more, depending on complexity. An ai-powered legacy code modernization service can shift spending from manual transformation toward testing, verification and engineering review, particularly during the first engagement.
Five problems account for most failed AI modernization efforts. Each has a practical control, and none of them is a better prompt.
Challenge 1: Plausible wrongness: The failure mode is not code that breaks the build. It is code that compiles, passes weak tests, looks correct in review, and behaves differently in one edge case discovered months later.
Solution: Gate on behaviour rather than on appearance. Characterisation tests written before transformation, run automatically on every change, catch what review does not. Where coverage is thin, restrict AI to comprehension work until the tests exist.
Challenge 2: Behaviour drift during translation: Cross-language migration is where this is concentrated. Semantics that differ subtly between languages, around numeric precision, date handling, null semantics and error propagation, are exactly where models produce confident and wrong output.
Solution: Run both implementations against the same inputs and compare outputs, rather than reviewing the translated code by eye. Build the comparison harness before the translation starts, and treat the known-divergent areas as a specific test checklist.
Challenge 3: Review capacity as the real constraint: Generation scales; human review does not. Any plan that increases change volume without increasing review capacity has relocated the bottleneck rather than removed it.
Solution: Cap the volume of AI-generated change entering review per cycle, and measure review hours per change landed alongside delivery speed. If the queue lengthens, the constraint has moved rather than disappeared.
Challenge 4: Trust without verification: DORA found 30% of practitioners report little or no trust in AI-generated code, while more than 80% believe AI has made them more productive. Holding both positions at once is reasonable, and it means governance cannot rely on practitioner confidence as a signal.
Solution: Base governance on measured outcomes rather than on team sentiment. Record what was expected before each phase alongside what was measured, and use the gap as the review signal. Our AI consulting practice covers AI readiness assessment and AI governance consulting for exactly this.
Challenge 5: Data and IP exposure: Legacy code often contains embedded credentials, customer data in test fixtures, and proprietary business logic.
Solution: Confirm where the code is processed, how it is retained, whether it can be used for model training, and what access controls apply before the first file leaves your environment. Scan for embedded secrets and scrub test fixtures as a precondition, not as a follow-up.
Organisations without in-house experience in characterisation testing, verification pipelines or AI tooling often bring in external support for the first engagement. The evidence above suggests what to look for in that support. A good partner starts with comprehension and test coverage before transformation. It measures outcomes against a baseline instead of relying on self-reported speed. It keeps change sets small enough to review properly and leaves a human decision point on roll out. Given the data and IP risks covered earlier, it’s also worth asking how and where your code will be processed and retained.
The work usually spans several disciplines. Legacy software modernization services typically include application modernization consulting, software re-engineering, code refactoring, data modernization and API upgradation. An AI development practice covers model selection and integration, which determines how AI tooling fits into the verification loop. Where assessment shows that a rebuild makes more sense than a refactor, a custom software development practice handles the new build.
SparxIT works across these areas, drawing on 19+ years of software engineering and a decade of AI development, including more than 20 production-level AI use cases.






Yes, and incremental use is where it performs best. AI can document modules, generate tests around existing behaviour, and transform code section by section while the application keeps running. Full rewrites can increase verification risk, because they remove the existing implementation that teams could use as a behavioural reference.












Models handle widely represented languages best, including Java, C#, JavaScript, Python and PHP. COBOL, RPG and PL/I are supported but with less reliable idiomatic output, because training data is thinner. Practical suitability depends more on test coverage and code structure than on the language itself.












It can be, under the same controls as any other change: traceable review, test evidence, and an audit trail showing who approved what. The additional questions regulators ask concern where code was sent for processing and what the retention terms were. Settle both before the first file leaves your environment.












Timelines depend on whether the verification loop already exists. On a codebase with automated builds and good test coverage, mechanical transformations can move very quickly, as Google’s 149,000-line migration in three months shows. On an untested codebase, expect the early phase to go into building test coverage before any transformation starts.












Capture current behaviour as characterisation tests before any transformation, review those tests against business expectations, and run them automatically on every change. Keep changesets small enough to review properly. Behaviour preservation is a verification problem rather than a prompting problem, and it is solved by the test suite rather than by the model.