Data foundations Guide
Data Warehouse Migration in 2026: When to Move and How to Cut Over Safely
Support deadlines, open table formats and AI have put the old warehouse on every data leader's agenda. The platform is the easy choice. What decides the outcome is how much you retire before moving, how converted code is proven, and when the old system is switched off. This guide sets out the evidence for each.
For CIOs, CDOs and heads of data deciding whether and how to move a Teradata, Oracle, Netezza, SAP BW, Hadoop or Synapse estate to a cloud warehouse or lakehouse.
The short answer
Migrate a data warehouse when a support deadline, an unsupported platform or an AI program forces the question, and treat the platform as the smaller decision. Retire unused reports and tables first, keep data in an open table format, convert code with tools but accept it only when both systems return the same results, and switch the old platform off by domain on dates set in advance.
Key takeaways
- SAP NetWeaver 7.5, which carries SAP BW 7.5, leaves mainstream maintenance at the end of 2027; Hadoop and Netezza appliances are already past vendor support.123
- The biggest saving comes before anything moves: LinkedIn cut about 70% of its migration workload by consolidating 1,424 datasets into 450.4
- Every major platform now reads Apache Iceberg, so lock-in has moved from the storage format to the catalog, the metric definitions and the contract.567
- Vendors claim their free tools automate 80 to 96% of conversion, while independent benchmarks put language models at 38 to 55% accuracy on cross-dialect SQL.891011
- The widely quoted 83% failure rate cannot be traced to Gartner; Bloor Research's surveys found 84% of data migrations overran in 2007 and nearly 62% delivered on time and budget by 2011.12
- Consumption platforms are built for spend to grow: Snowflake's existing customers spent 26% more year on year, and most built-in limits alert rather than stop.1314
Most data warehouse migrations are pitched as a choice between platforms. The record of companies that have done it says the platform matters less than three decisions made around it: how much of the old estate is retired before anything moves, how converted code is proven to give the same answers, and when the old system is finally switched off. This guide covers why 2026 forces the question, what the named migrations moved and gained, how far AI conversion tools can be trusted, and how to cut over without breaking the reports a business runs on.
- ~70%of LinkedIn's migration workload removed by consolidating 1,424 datasets into 450 before moving4
- End of 2027mainstream maintenance ends for SAP NetWeaver 7.5, the platform under SAP BW 7.51
- Under 38.5%average accuracy of language models translating SQL between database systems in an independent benchmark10
Why 2026 is a decision year
For many estates the vendor has already set the date. SAP's mainstream maintenance for NetWeaver 7.5, the platform SAP BW 7.5 runs on, ends at the end of 2027, with optional extended maintenance to the end of 2030 at a premium, while BW/4HANA follows S/4HANA's commitment to 2040.1 Others are past their deadline and running on borrowed time.
| Platform | Support position | What it means for a migration |
|---|---|---|
| SAP BW 7.5 on NetWeaver 7.5 | Mainstream maintenance to the end of 2027, extended to the end of 2030 at extra cost1 | A dated choice between BW/4HANA, SAP's cloud data products or another platform |
| Cloudera CDH and HDP | HDP 3.1 support ended in December 2021; CDH 6 and HDP 3.1 had a last six months of limited support in 20222 | Estates still in production carry unsupported risk today |
| Cloudera platform 7.x | Version 7.3.1 support ends December 2026; Base 7.1.9 in October 20282 | Even current Hadoop customers face an upgrade or a move inside three years |
| IBM Netezza N3001 appliances | Support ended April 30, 2023, and IBM says service extensions are not available3 | Hardware failure has no vendor remedy |
| Azure Synapse dedicated SQL pools | Supported, with Microsoft's own overview steering new and existing workloads to Fabric Data Warehouse15 | No deadline yet, and new features are going elsewhere |
AI adds a second reason. In a Gartner survey of 1,203 data management leaders, 63% said their organizations either lack or are unsure they have the right data management practices for AI, and Gartner predicts organizations will abandon 60% of AI projects unsupported by AI-ready data through 2026.16 The prediction is a forecast, and its direction matches where the money is going. Snowflake's product revenue grew 37% year on year to $1.49 billion in its quarter to July 2026, and Databricks reported a revenue run-rate above $7 billion, growing more than 80%.1317 In a Dremio survey of 563 IT decision-makers, 55% already ran most of their analytics on a lakehouse and 67% expected to within three years.18 Dremio sells lakehouse software, so the direction is the useful part.
Open table formats moved the lock-in to the catalog
The biggest change since 2024 is where the data sits. Databricks bought Tabular, the company founded by Apache Iceberg's creators, to reconcile Iceberg with its own Delta Lake, conceding that the two formats "became incompatible due to their independent development."19 Snowflake made Iceberg tables generally available in June 2024 and its Open Catalog, previously named Polaris, in October 2024.520 AWS launched S3 Tables with built-in Iceberg support in December 2024, and Google made its BigLake Iceberg tables and metastore generally available in May 2025.216 A table written once in customer-owned storage can now be read by several engines.
That shifts the lock-in instead of ending it. The Apache Polaris proposal itself says that "many interdependent limitations exist between engines and catalogs, which create lock-in."7 The two leading catalogs are both open source and compete: Databricks released Unity Catalog under Apache 2.0 in June 2024 with an Iceberg REST interface, and Apache Polaris became a top-level Apache project in February 2026.2223 Permissions, lineage and metric definitions live in that layer, so choosing the catalog deserves the care buyers used to give the database.
Open formats also make the next move cheaper, including a move off a cloud warehouse. New Relic moved more than 1,000 datasets from Snowflake to Iceberg and reports a 35 to 52% reduction in annual data platform spend, after finding itself "deeply locked into a single vendor's storage, compute, and query model."24 Notion built its own lake in 2022 because 90% of its upserts were updates, a pattern most warehouses are not built for, and reports net savings of over a million dollars that year.25
What the named migrations moved and gained
Companies publish what they moved and how fast it runs now. None of the companies below publishes the total cost of the project.
| Company | What moved | Reported result |
|---|---|---|
| LinkedIn, Teradata to an open-source stack | Its analytics warehouse, after mapping lineage across every dataset | 1,424 datasets consolidated to 450, about 70% less to migrate; savings described as millions in licensing and support4 |
| PayPal, a dozen systems to BigQuery | About 400 PB across a dozen systems, including what PayPal believes was the world's largest Teradata deployment | More than 300 PB moved, about 25% of workloads decommissioned, four infrastructure vendors cut to one, queries 2.5 to 10 times faster, with zero downtime26 |
| Uber, Hadoop to Google Cloud | More than 1 exabyte in each of two regions | Storage and the existing stack move first, as they are, so dashboard owners and pipeline authors change nothing27 |
| Children's Hospital of Philadelphia, Netezza to Snowflake | A 15-year-old on-premises warehouse | Moved over 18 months with an in-house validator checking rows and aggregates28 |
| Notion, warehouse ingestion to its own lake | Update-heavy Postgres data that was slow and costly to load | Net savings of over a million dollars in 2022, and higher since25 |
| New Relic, Snowflake to Iceberg | More than 1,000 batch and streaming datasets | 35 to 52% lower annual data platform spend24 |
Three lessons repeat. Scope comes down before anything moves: LinkedIn used lineage to collapse its datasets, and PayPal switched off a quarter of its workloads.426 At very large scale, the data moves first and the code is modernized later, so users see no change on day one.27 And running two warehouses has a price that grows with time. LinkedIn's first, unplanned attempt left both systems running with data copied between them, which it says "resulted in double the maintenance cost and complexity, along with very confused consumers."4
The failure statistic everyone quotes
The figure that 83% of data migrations fail or overrun is usually credited to Gartner, and we could not trace it to any Gartner publication. The traceable source is Bloor Research. Its 2007 survey found that 84% of data migration projects ran over time, over budget or both. Its 2011 survey found nearly 62% delivered on time and on budget, and 30% delayed, by about four months on average and sometimes a year or more.12 Both surveys predate cloud warehouses, and no current independent survey replaces them.
The most expensive recent failure shows where the risk actually sits. When TSB moved its customer services to a new platform in April 2018, the Bank of England noted that the data itself migrated successfully. The platform then failed, all of TSB's branches and a large share of its 5.2 million customers were affected, and service did not return to normal until December 2018. Regulators fined TSB £48.65 million, on top of £32.7 million paid to customers in redress.29 Moving the data is the part teams plan for. Proving the new system behaves the same under real load, and having a way back, is the part that fails.
AI conversion: fast on syntax, unproven on meaning
Every major platform now gives away its migration tooling, and each has added a language model on top of a rules engine. Snowflake made SnowConvert free in January 2025 and says it automates more than 96% of code and object conversion.8 Google says its Gemini-enhanced BigQuery translation delivers over 95% accuracy, Databricks says Lakebridge automates up to 80% of migration tasks, and AWS says its generative AI conversion handles up to 90% of schema objects for moves to PostgreSQL.30931 None of these figures comes with a published method, and each measures something different.
The independent numbers are lower because they check results. On PARROT, a benchmark of SQL translation across 22 database systems, language models averaged under 38.53% accuracy.10 On UniQL, the best model answered 54.63% of questions correctly across 16 dialects, scored lowest on Teradata at 37.74%, and got only 20.14% of questions right in every dialect.11 Enterprise schemas are harder still. On ESQ-Bench, built on populated enterprise Oracle schemas, GPT-4o's execution accuracy fell from 79.8% to 57.2% as complexity rose, against more than 89% reported on academic benchmarks, and most failures at the harder tiers were queries that ran and returned wrong results.32
The vendors' own documentation points the same way. Microsoft's Fabric Migration Assistant warns that "mistakes can happen as Copilot uses AI, so verify code suggestions before running them," and notes that T-SQL compatibility is incomplete.33 AWS lists generative AI conversion as unavailable for Oracle to Redshift, its only warehouse target.34 Snowflake's newer verification runs converted logic on both the source system and Snowflake with synthetic data and compares the results.35 That is the right test for any tool: a converted query is accepted when it returns the same answer as the original on the same data.
Reconciliation is the acceptance test
Data checks run in tiers. Row counts by table and partition come first, then column aggregates, then row-level hash comparisons on the tables that feed money, risk and regulatory figures. Google's open-source Data Validation Tool covers all three levels across sources including Teradata, Oracle, Snowflake and BigQuery.36 Datafold sunset its open-source data-diff package in May 2024 and kept diffing in its commercial product, a reminder to check the maintenance status of any free tool a migration depends on.37
The acceptance threshold belongs to the business. In Faire's move of more than 5,000 tables from Redshift to Snowflake, each business group agreed a target level of accuracy with its data engineers before signing off, according to a case study published by the diffing vendor.38 Databricks' migration team runs parallel testing and reconciliation with business experts in user acceptance testing, starting from a production pilot of one end-to-end use case.39 Agreeing tolerances before the first comparison keeps sign-off from turning into a negotiation at the end.
The cheapest table to reconcile is the one that never moves. In a migration Microsoft documents, a consumer goods company found that about half of its thousands of reports had been opened in the previous year, and about half of those delivered significant value. Microsoft's advice on the rest is blunt: "Sometimes, the cheapest and easiest thing to do is nothing."40 Metric definitions deserve the opposite treatment. They drift quietly when SQL is rewritten, so writing each one down once, in a portable form such as the open-source Open Semantic Interchange specification or dbt's MetricFlow, makes them testable and keeps them out of the next lock-in.41
How the bill behaves after go-live
Consumption platforms charge for capacity running, and the idle traps differ by platform. Snowflake bills per second with a 60-second minimum each time a warehouse starts, and its resource monitors cover warehouses only, leaving serverless and AI features to separate budgets.4243 BigQuery charges for the slots it scales up to, not the slots a job uses, even if the job fails, and keeps scaled capacity for at least 60 seconds.44 Redshift Serverless keeps billing for open transactions, cancelled queries and the health-check queries connection pools send.45 Databricks says its budgets should not be used "as a way to ensure an absolute spend cap."14
Growth after adoption is the design. Snowflake's net revenue retention of 126% means its existing customers spent 26% more than a year earlier.13 Instacart's 2023 filing showed payments to Snowflake of about $13 million, $28 million and $51 million from 2020 to 2022; Snowflake replied that actual usage was $28 million in both 2021 and 2022, and that Instacart had optimized its Snowflake costs by 50%.4647 Prepaid commitments and real consumption can tell very different stories, and tuning moves the bill as much as the platform does. Flexera's 2026 respondents put wasted cloud spend at 29%, the first rise in five years, and the FinOps Foundation finds data cloud platforms and AI are now the most actively managed areas because of "less predictable usage patterns."4849
In regulated firms, the migration is a supervised change
| Rule | What it asks of a migration |
|---|---|
| ECB guide on risk data aggregation and reporting, May 2024 | Data lineage at attribute level, starting from data capture and including extraction, transformation and loading, and data governance taking part in IT change initiatives50 |
| DORA, applying since January 17, 2025 | Changes to ICT systems recorded, tested, assessed, approved, implemented and verified; tested exit plans that can move the relevant data to another provider or back in-house51 |
| EU Data Act | Switching charges between data processing services banned from January 12, 202752 |
| HIPAA technical safeguards | Electronic mechanisms to corroborate that health information has not been altered or destroyed in an unauthorized manner53 |
The ECB lists reconciliation errors among the causes of large-scale miscalculations of key risk ratios it has found at banks.50 DORA's exit requirement also turns this migration into a design brief for the next one: open formats, a catalog that other engines can read, and code kept in a repository make a tested exit possible.51
Lift and shift
Copy or virtualize as-is
- Fastest route off a dying platform
- Unused reports and tables move too
- Old cost patterns land on a metered bill
- Modernization deferred, often for years
Rebuild at once
New models, one cutover
- A clean design on paper
- Every report changes on the same day
- Reconciliation becomes one huge sign-off
- The highest chance of a TSB-style weekend
Domain by domain
How we advise
- Unused estate retired before conversion
- Each domain reconciled against agreed tolerances
- Old platform switched off on dates set in advance
- Open formats keep the next move cheap
How to run a data warehouse migration
- Inventory usage and retire first Read the audit logs and lineage of the old platform, list who uses each report and table, and retire what nobody opened in a year. Every object retired is one fewer to convert, test and reconcile.
- Choose the format and catalog before the engine Keep tables in an open format such as Iceberg in storage you own, and pick the catalog for permissions, lineage and the engines that must read it.
- Write the metrics down once Define revenue, margin, active customer and every other headline number in one portable semantic layer before any SQL is rewritten, so old and new can be compared on the same definitions.
- Convert with tools, accept on results Let the rules engine and the language model do the first pass, then accept code only when the converted version returns the same results as the original on the same data.
- Reconcile in a parallel run Run counts, aggregates and row-level hashes on every load cycle, against tolerances the business owner agreed in advance, until each domain has a run of clean cycles.
- Cut over by domain and switch off on a date Move one domain at a time with a tested way back, and set the date the old platform stops for that domain before the work starts, so two systems never run indefinitely.
- Turn on spend controls on the first day Set warehouse monitors, maximum capacity limits and budgets per team before users arrive, and review the first month's bill line by line.
Questions before signing a migration plan
- Which reports and tables were used in the last year, and who owns each one that stays?
- When does vendor support for our current platform end, and what does running past it risk?
- Which catalog will hold permissions and lineage, and which engines can read our tables through it?
- How will converted code be proven to return the same results, and who signs that off?
- What tolerance will each business owner accept, and how many clean cycles before cutover?
- On what date does each part of the old platform switch off?
The platforms will keep converging, and the free tools will keep improving at the parts that are easy to measure. What separates a migration that pays from one that drags on is still the unglamorous work: an estate made smaller before it moves, numbers that match to an agreed tolerance, and an old system that is actually switched off.
This is how we run data engineering work: an audit of what is used and what each number means, open formats in your own cloud account, reconciliation built into the pipeline, and the old platform retired source by source. On one hospital ERP, putting reporting on a single layer took administrative reports from days to minutes.
Questions leaders ask
What is data warehouse migration?
Data warehouse migration is moving an organization's analytical data, the code that transforms it and the reports that read it from one platform to another, such as from Teradata, Oracle, Netezza, SAP BW or Hadoop to a cloud warehouse or lakehouse. It covers schemas, data, SQL and ETL logic, security, metric definitions and the switch-off of the old system.
When should we migrate our data warehouse?
When a support deadline, an unsupported platform or an AI program makes the old warehouse a risk. SAP NetWeaver 7.5, which carries SAP BW 7.5, leaves mainstream maintenance at the end of 2027, and Hadoop distributions and Netezza appliances are already past vendor support. A promised saving alone is a weaker reason, because consumption bills grow after go-live.
Should we move to a data warehouse or a lakehouse?
It depends on how the data is read. A cloud warehouse suits modeled reporting with modest data science. A lakehouse keeps one copy in an open table format that SQL, machine learning and AI workloads can all read. Since 2024 the major warehouses also read Apache Iceberg, so the choice of catalog now matters as much as the engine.
Can AI convert our SQL and ETL code automatically?
It converts most syntax and saves real time, but its output needs proof. Vendors claim 80 to 96% automation, while independent benchmarks put language models at 38 to 55% accuracy on cross-dialect SQL, lowest on Teradata. Accept converted code only when it returns the same results as the original on the same data.
How long does a data warehouse migration take?
It varies with the size of the estate and how much is retired first. Published examples range from defined first phases to multi-year programs: a children's hospital moved a 15-year-old Netezza warehouse to Snowflake over 18 months, and Uber is moving more than an exabyte in stages. Bloor found delayed projects slipped by about four months on average.
How do you validate data after a warehouse migration?
In tiers, on every load cycle of a parallel run: row counts by table and partition, then column aggregates, then row-level hash comparisons on critical tables. Each business owner agrees a tolerance in advance, and a domain cuts over after a run of clean cycles. Google's open-source Data Validation Tool covers all three levels.
How do we avoid lock-in on the new platform?
Keep tables in an open format such as Apache Iceberg in storage you own, choose a catalog that other engines can read, write metric definitions in a portable semantic layer, and keep all code in your own repository. In the EU, the Data Act also bans switching charges between data processing services from January 12, 2027.
Sources
- SAP NetWeaver 7.5 Maintenance StrategySAP Community
- Cloudera Support Lifecycle PolicyCloudera
- PureData System for Analytics N3001-020_1.0.x - Withdrawal notificationIBM
- Evolving LinkedIn's analytics tech stackLinkedIn Engineering, December 7, 2021
- Open Storage with Iceberg Tables Now Generally AvailableSnowflake, June 2024
- BigLake evolved: Build open, high-performance, enterprise Iceberg-native lakehousesGoogle Cloud, May 30, 2025
- Polaris ProposalApache Software Foundation Incubator
- Snowflake Accelerates Self-Serve Data Warehouse Migrations with Free SnowConvert for AllSnowflake, January 28, 2025
- Introducing Lakebridge: Free, Open Data Migration to Databricks SQLDatabricks, June 12, 2025
- PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL TranslationZhou et al., NeurIPS 2025
- UniQL: Towards Dialect-Universal Benchmarking for Text-to-SQLGao et al., June 2026
- White Paper: Data MigrationBloor Research, May 2011
- Snowflake Reports Financial Results for the Second Quarter of Fiscal 2027Snowflake via SEC, September 2, 2026
- Create and monitor budgetsDatabricks Documentation
- What is dedicated SQL pool (formerly SQL DW)?Microsoft Learn
- Lack of AI-Ready Data Puts AI Projects at RiskGartner, February 26, 2025
- Databricks Grows >80% YoY, Surpasses $7B Revenue Run-RateDatabricks, August 13, 2026
- State of the Data Lakehouse in the AI Era reportDremio, January 9, 2025
- Databricks + TabularDatabricks, June 2024
- Apache Iceberg tables: Support for Snowflake Open Catalog, General AvailabilitySnowflake, October 18, 2024
- AWS expands Amazon S3 with features to support Apache Iceberg and metadata managementSiliconANGLE, December 3, 2024
- Databricks Open Sources Unity CatalogDatabricks, June 12, 2024
- Apache Polaris Graduates to Top Level Project!Apache Polaris, February 19, 2026
- Escaping the Snowflake Tax: How we cut data costs by over 50% at New RelicNew Relic, April 29, 2026
- Building and scaling Notion's data lakeNotion, July 1, 2024
- PayPal's historic data migration is the foundation for its gen AI innovationGoogle Cloud, February 27, 2026
- Modernizing Uber's Batch Data Infrastructure with Google Cloud PlatformUber Engineering, May 30, 2024
- Untangling a Legacy Data Migration: Lessons from Netezza to SnowflakeDatafold
- TSB fined £48.65m for operational resilience failingsBank of England, December 20, 2022
- Simplify your data platform migration with AI-powered BigQuery Migration ServicesGoogle Cloud, April 11, 2025
- AWS Database Migration Service now automates time-intensive schema conversion tasks using generative AIAWS, December 1, 2024
- ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic DivergenceMishra, Chukkapalli and Naik, 2026
- Migration Assistant for Fabric Data WarehouseMicrosoft Learn
- Converting database schemas using DMS Schema ConversionAWS Documentation
- What's New in SnowConvert AI: February 2026Snowflake, February 9, 2026
- Automate data validation with DVTGoogle Cloud
- Sunsetting open source data-diffDatafold, 2024
- Faire Redshift to Snowflake migrationDatafold
- Databricks Migration Strategy: Lessons LearnedDatabricks, October 23, 2024
- Learn from customer Power BI migrationsMicrosoft Learn
- What the Open Semantic Interchange (OSI) spec means for metrics, semantics, and AIdbt Labs, January 29, 2026
- Overview of warehousesSnowflake Documentation
- Working with resource monitorsSnowflake Documentation
- Introduction to slots autoscalingGoogle Cloud Documentation
- Billing for Amazon Redshift ServerlessAWS Documentation
- Snowflake gets sensitive about Instacart's $100M paymentsThe Register, August 31, 2023
- Snowflake and Instacart: The FactsSnowflake, August 30, 2023
- Flexera 2026 State of the Cloud Report: The convergence of cloud and valueFlexera, 2026
- State of FinOps 2026 ReportFinOps Foundation, 2026
- Guide on effective risk data aggregation and risk reportingEuropean Central Bank, May 2024
- Regulation (EU) 2022/2554 (Digital Operational Resilience Act)EUR-Lex
- Regulation (EU) 2023/2854 (Data Act)EUR-Lex
- 45 CFR 164.312, Technical safeguardsLegal Information Institute, Cornell Law School
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 4, 2026. No client data appears in our insights.