A widely quoted claim that most data migration projects fail or exceed their budgets traces back to a 2007 white paper. It has been repeatedly recycled, rounded up and reattributed, while many references omit its publication date.
The current picture is less dramatic and more useful. Standish Group figures for the 2020 to 2024 cycle put 31% of IT projects as successful, 50% as challenged and 19% as outright failures. Most migrations do not collapse. They arrive late, over budget, or with quality problems found after go-live.
That matters, because the real risk in a migration is not a statistic. Legacy data migration is the work of moving business records out of a system that can no longer hold them and into one that can, with the meaning, relationships and completeness of those records intact. The failure mode is rarely a dramatic collapse. It is a migration that appears to succeed, goes live, and is found six weeks later to have silently dropped a category of records nobody thought to check.
Everything below is built around preventing that specific outcome: deciding what actually needs to move, moving it in a sequence you can reverse, and proving it arrived.
Legacy data migration is the process of transferring data from an outdated system into a modern platform while preserving its accuracy, structure, relationships and auditability. It covers extraction, cleansing, mapping, transformation, loading, reconciliation and the decision about what to do with records that cannot be moved.
A data migration from legacy to new system is not a copy operation, and neither is a legacy database migration. The source encodes meaning in ways the target will not accept: status flags that mean different things depending on which decade the record was created, free-text fields carrying structured data, customer identifiers that were merged during an acquisition, and business rules that exist only as code rather than as documented logic.
Data migration is usually one workstream inside a wider programme, and our legacy application modernization guide covers the application side of the same effort.
The work is therefore mostly archaeology and verification. Extraction and loading are the easy parts. Understanding what the data means, deciding what it should become, and proving the result matches is where the effort goes, and where budgets are lost.
Search this topic and you will meet the same claim in several forms: 83% of data migrations fail, or more than 80% fail or exceed budget, or 60% fail on the first attempt. These get attributed to Gartner, to analyst firms generally, or to nobody at all.
The traceable origin is a Bloor Research white paper by Philip Howard, published in September 2007, which found that more than 60% of data migration projects overran on time, budget, or both. The same paper estimated industry spend on data migration at over $5 billion a year and argued the root cause was that “the techniques and disciplines of data migration are not treated seriously enough or are not well enough understood.”
Two things follow. The widely quoted 83% is higher than what the original research reported, which is what happens to a number passed between marketing pages for eighteen years. And the underlying finding is now old enough that it predates cloud data platforms, modern ETL tooling and most current database engines.
For something current, the broader project data is more defensible. Standish Group CHAOS figures for the 2020 to 2024 cycle, reported by Giuseppe Arcidiacono in PM World Journal in January 2026, put 31% of IT projects as successful, 50% as challenged, and 19% as outright failures, with what the paper describes as remarkably little movement across the period.
Treat that as the honest baseline. Many migration projects avoid outright failure but still face schedule, budget or data-quality problems before or after go-live, and that is the outcome the rest of this guide is written to avoid.
The single biggest cost control in any legacy data migration strategy is deciding, early and explicitly, that not all of the data is moving.
Most teams default to migrating everything, because it feels safer and because nobody wants to be the person who deleted something. That default is what turns a six-month project into an eighteen-month one. Every additional record class carries mapping, transformation, testing and reconciliation effort, and much of it is data no user has touched in a decade.
| Data class | Test | Destination |
| Active operational records | In use now, or needed for live processing | Migrate to the new system |
| Recent historical records | Queried regularly for service, reporting or support | Migrate, possibly in a simplified structure |
| Retained-but-dormant records | Not queried, but held under a retention obligation | Archive in a queryable read-only store |
| Reference and lookup data | Codes, hierarchies and mappings the new system needs | Migrate, after rationalising retired values |
| Superseded or duplicate records | Replaced by a later record, or a known duplicate | Resolve before migration, do not carry forward |
| Data past its retention period | No legal, regulatory or business reason to keep it | Dispose under your retention policy, with sign-off |
Table: A six-class triage for legacy data. Deciding this before mapping begins is the difference between migrating what matters and migrating everything.
Two rules keep this defensible. Get the retention decision signed off by whoever owns records management rather than by the project, because disposal is a governance act and not a technical one. And document the test applied to each class, so that a question six months later about why a record set was not migrated has a written answer.
This is the sequence for a data migration from a legacy system to a modern platform. Steps four through eight repeat per wave rather than running once. Where the target is cloud infrastructure, our guide to migrating legacy applications to the cloud covers the platform move that runs alongside it.
Run the data, not the documentation. Count records per table, measure null rates per field, find the distinct values in every status and code column, and identify the fields whose format changes partway through the history. Documentation for systems of this age describes intent; profiling describes reality.
Agree what migrates, what archives and what is disposed of, and get records management sign-off before mapping starts.
For each target field, record the source field, the transformation rule, the handling for nulls and out-of-range values, and the named owner of that rule. This document is the migration; the code is just its implementation.
Fixing a duplicate customer record in the legacy system fixes it once. Fixing it in the transformation logic handles the issue consistently during migration, but it can leave the underlying source-data problem unresolved.
You will run it many times. A migration that cannot be safely rerun makes testing, recovery and failure handling much harder.
Sub-sampled runs hide timing problems, and timing is what forces the cutover window. A trial run that does not use production-scale volumes has not tested the thing that will go wrong.
Covered in detail in the next section. This is the step most commonly compressed under schedule pressure and most commonly regretted.
Sequence by business domain rather than by table, so each wave delivers something a user can verify.
Agree in advance what is checked daily for the first fortnight, who checks it, and what threshold triggers a rollback rather than a fix-forward.
A legacy system data migration is won or lost at reconciliation, and it is the part most guides reduce to “test thoroughly.” Testing confirms the process ran. Reconciliation proves the data arrived intact and still means what it meant.
Define these checks before the first trial run, with a written acceptance threshold for each. A check with no agreed threshold becomes a negotiation at two in the morning on cutover weekend.
| Check | What it proves | Typical acceptance criterion |
| Record counts by entity | Nothing was silently dropped in transit | Exact match, or every variance individually explained and signed off |
| Control totals | Numeric data arrived with its values intact | Financial and quantity totals match to the penny or unit |
| Referential integrity | Relationships between records survived | Zero orphaned child records against migrated parents |
| Field-level completeness | Mandatory data did not arrive empty | Null rate in target no higher than profiled null rate in source |
| Value distribution | Transformations did not distort the data | Distribution of key coded fields matches source within an agreed tolerance |
| Business rule equivalence | The new system reaches the same answers | Sample transactions produce identical outputs in both systems |
| Reporting equivalence | Downstream numbers still tie out | Key operational and regulatory reports match across both systems for the same period |
| Sample record inspection | The data is correct, beyond being numerically consistent | A stratified sample reviewed field by field by a business user, not by the migration team |
Table: Eight reconciliation checks for a legacy data migration, with the acceptance criterion each one needs agreed in advance.
Three disciplines make this work. Reconcile at every run rather than only the final one, so you see trends rather than a single verdict. Have a business owner rather than the migration team sign off the sample inspection, because the team that wrote the mapping is the least likely to notice a mapping assumption that is wrong. And keep the reconciliation output, because it is the evidence that the migration was sound when someone asks about a specific record two years from now.
Every migration of real data produces records that will not move. The target validates a field the source never populated. A customer exists on a transaction but not in the customer table. A date sits in the year 1900 because a legacy screen defaulted it. Two records claim the same unique identifier.
The failure is not that these exist. It is treating them as defects to be fixed one at a time under cutover pressure, which is how a migration weekend overruns.
Handle them as a designed path instead. Give every rejected record a machine-readable reason code rather than a log line, so exceptions can be counted, grouped and trended between runs, then agree the disposition per category before cutover rather than during it.
| Exception category | Typical example | Disposition |
| Correctable at source | Duplicate customer records, missing mandatory field | Fix in the legacy system and re-extract |
| Correctable by rule | Placeholder dates, retired status codes with a known successor | Transform under a documented default, signed off by the data owner |
| Needs a business decision | Two records claiming the same unique identifier | Quarantine, route to the named business owner, resolve before go-live |
| Orphaned relationships | Transaction referencing a customer that no longer exists | Migrate with a placeholder parent, or exclude with sign-off |
| Beyond retention | Records past their retention period surfacing as errors | Dispose under the retention policy rather than remediating |
| Genuinely unreadable | Corrupt records, unrecoverable proprietary formats | Document, exclude, and report the count to the data owner |
Table: Six exception categories in a legacy data migration and the disposition each needs agreed before cutover, not during it.
Set an exception threshold that halts the cutover rather than hoping the number stays small. And publish exception counts after every trial run, because a category that grows between runs is telling you the source system is still changing in ways your mapping does not handle.
Format obsolescence is the long-term version of the same problem. Even the US National Archives, in its digital preservation programme, performs format transformations when it receives material it cannot process, and runs media migration on a multi-year cycle. Data you cannot read is not retained data, whatever your retention schedule says.
The steps above describe the sequence. These legacy data migration best practices describe the disciplines that run across all of it, and they are mostly about people rather than tooling.
The one that saves the most time is the least technical: a single named person, on the business side, who can settle what a field means without convening a meeting.
Most legacy data migration challenges trace back to three things: undocumented sources, data quality nobody owns, and reconciliation time that gets compressed.
The people who designed the source system have usually left. Business rules exist as code, and the documented rules describe a system that changed years ago. Profiling and sample inspection are the only reliable ways to recover the truth, and both take longer than teams plan for.
The legacy system is still in production while you migrate it. Records change, new ones arrive, and a mapping built against a snapshot drifts out of date. Either freeze structural change or design for delta handling from the start.
Profiling reveals duplicates, orphans and invalid values that predate the project. Fixing them is genuinely valuable and was genuinely not in the estimate. Decide early whether the migration owns data remediation or simply carries the problems across with documentation.
It sits at the end of the plan, which is where the schedule pressure lands. Protecting reconciliation time is the highest-leverage scheduling decision on the project. Our write-up of cloud data migration challenges covers the infrastructure-side version of the same squeeze.
Discovered during the full-volume trial run, if you do one, and on cutover weekend if you do not.
Rejected records accumulate without a decision-maker, and the backlog becomes a go-live blocker in the final fortnight.
Migration cost depends on data complexity, validation requirements and data volume, with complexity often having a greater impact than raw volume. A terabyte of clean, well-structured records is cheaper to move than a hundred gigabytes carrying thirty years of merged, undocumented history.
For a reference point on adjacent work, our cloud migration process guide puts cloud migration between $5,000 and $100,000 or more, depending on location, expertise, team structure and the degree of application modernization required. Pure data migration costs can fall within a similar range, but the final estimate depends on source complexity, data quality, transformation depth and validation requirements.
Six factors move a quote:
We help organizations profile, cleanse, map, move and reconcile data out of legacy systems and into modern platforms, backed by 19+ years of software engineering experience and 2000+ completed projects. Our solutions for legacy data migration sit inside our legacy software modernization services, which include data modernization and API upgradation alongside application re-engineering, so the migration is planned against the target architecture rather than in isolation.
Our ERP modernization practice has modernized over 40 enterprise ERP systems, with data modernization and data mapping as named parts of the process. For teams whose destination is an analytics platform rather than a transactional one, our big data analytics practice covers data ingestion, warehousing, cleaning and modeling.






AI tooling is strongest in profiling and mapping: inferring the structure of undocumented sources, proposing field-level mappings, spotting anomalies across millions of records, and generating reconciliation test cases. It compresses the discovery phase, which is where migrations lose the most time. Proposed mappings still need business sign-off, because a plausible mapping can be confidently wrong.












Duration tracks source complexity rather than data volume. A single well-documented source with clean data can move in weeks. A multi-source migration with undocumented business rules, heavy remediation and regulatory reconciliation runs several quarters. The number of trial runs required before reconciliation passes is usually the best predictor available.












Overruns concentrate in three places: discovery taking longer than planned because the source is undocumented, data quality remediation that surfaces during profiling and was never estimated, and reconciliation cycles repeating because each run reveals new exception categories. Full-volume trial runs early are what turn those surprises into known quantities.












You profile it rather than read about it. Extract the real data, count records and null rates, enumerate distinct values in every coded field, trace relationships between tables, and reconstruct business rules from the data patterns and from interviews with long-serving users. Treat any surviving documentation as a hypothesis to be tested, not as fact.












Near-zero downtime is achievable using change data capture to keep the target synchronised with the source, then switching once the delta is small enough to apply in minutes. For many back-office systems, a planned cutover window can reduce implementation complexity and cost when the business can accommodate scheduled downtime.