Moving to a lakehouse without stopping the business

Data EngineeringFrom migrating legacy pipelines to Databricks on AWS2 min read
Legacy warehousemove step by stepLakehouseDatabricksComparecounts, totalsMonday reportstill on time

A step by step way to migrate legacy pipelines to Databricks on AWS while reports keep running.

Migrations rarely fail because of technology. They fail because the business cannot pause while you rebuild everything. Reports still need to go out on Monday. This is the approach I use to move legacy data pipelines to a lakehouse, in this case Databricks on AWS, without a big bang switch.

Why move at all

The usual reasons: old servers that are slow and expensive, pipelines nobody wants to touch, and new needs like machine learning that the old setup cannot support. Write these reasons down with numbers. You will need them to prove the migration worked.

Step 1: Make a list and rank it

List every pipeline, who uses its output, how often it runs and how often it breaks. Start with something important enough to matter but simple enough to finish in a few weeks. An early win buys trust for the hard parts.

Step 2: Run old and new side by side

SourcesLegacy pipelinesstill serving reportsLakehouseDatabricks on AWSComparerow counts, totalsSwitch
Both systems run in parallel until the outputs match.

Build the new pipeline next to the old one. Both run on the same inputs. A simple comparison job checks row counts, totals and a sample of records every day. Only when they match for a while do you move consumers over.

Step 3: Move consumers last

Dashboards and downstream jobs switch one by one, each with an owner who signs off. Keep the old pipeline running in read only mode for a short time, then turn it off for real. Leaving it running forever is how you end up paying for two systems.

Lessons

  • Do not lift and shift bad models. A migration is the cheapest moment to fix naming and logic.
  • Measure before and after. In one migration like this we saw about 60% better performance, and having that number made every later conversation easier.
  • Train the people who will use and maintain the new system. A platform nobody understands is not finished.
  • Infrastructure as code from day one. Terraform for the workspace, clusters and permissions saves you from snowflake environments.
Faizan Khan

Faizan Khan is an AI and data engineer in Berlin. Working on something like this? Book a 30 minute call or email hello@faizankhan.me.