Skip to main content
Back to Blog
Data Migration

Mission Critical Database Migration Blueprints You Can Copy

October 202616 min read
Mission-critical database migration planning dashboard with cutover waves and replication flow

Expand-contract should be your safest default for refactor projects. CDC/blue-green combo gets you closest to zero downtime for customer facing systems with hard SLAs. Big-bang cutover is frequently the fastest approach for small low traffic systems, assuming you can heavily test it. Always run your lowest risk waves first, but do throw a representative complex workload in early so you work through the hard problems before your mission critical systems rely on the solution. Runbooks, tests and our own lessons learned follow.

TL;DR:

  • Use expand/contract when possible when the store does not change; combine blue/green with CDC for hardened SLAs, but account for replication lag and write pauses.
  • Start with low-risk, low-effort workloads, but place one representative complex system in the first wave so you expose hard problems early.
  • Expect short locks during initial snapshots; snapshots of running PostgreSQL databases can lock each table for less than ten seconds, and Amazon RDS blue/green switchover defaults to a 300 second timeout.
  • Validate before cutover with production-scale data, row counts, checksums, and a dry-run rollback in staging against written numeric failure thresholds.

Plan for a safer infrastructure migration

PODTECH designs and develops bespoke software applications for mission critical infrastructure modernization. We specialize in legacy technology refresh and critical system integrations.

View PODTECH's Solutions

Table of Contents

Migration patterns and when to choose them

Each migration strategy involves a trade-off of speed vs risk vs operational complexity. Choosing an inappropriate strategy for your workload is the number one reason migrations fail, so take the time to match the strategy to your system before you modify a line of schema.

Big-bang (offline) migration migrates all the data in a single step during a maintenance window. It's fine for small databases, internal tools, or systems with a legitimate period of low traffic. The downside is your one chance to migrate will be in production, with an aggressive deadline, and no partially migrated fallback option. Mitigation strategies: fully practice the migration in staging with production data volumes, and provide a cutoff time at which the migration will automatically rollback if it doesn't successfully validate.

Phased or incremental migration: Moves data and traffic gradually, table by table or service-by-service. This approach is appropriate when your system has strong module boundaries and the teams benefit from learning through early waves to de-risk migrations. Risk: migration timelines can extend indefinitely due to maintaining two schemas in production. Mitigation: set a hard deadline for the phased migration period and measure schema drift weekly.

Expand-Contract (parallel change) strategy introduces new schema elements in parallel with old ones. It then incrementally migrates writes and reads until everything is using the new shape. Finally it removes the old schema elements. This is the go-to strategy for refactor/schema changes where the datastore itself isn't changing — only the shape of data stored within. Downsides include complexity creep if the contract part of the pattern keeps getting postponed. Mitigation: set an explicit expiration date on the old schema and treat it like a tracked deprecation, not a someday-maybe task.

Blue-green migration provisions an entire parallel environment (the “green” target), replicates into it, then flips traffic over during a single controlled cutover. Blue-green migrations work best for customer-facing systems that require a fast reversible switch and can tolerate a brief write pause. The primary risk with this method is replication lag near switchover time. Mitigation: enforce a switchover timeout and automated readiness checks. This is exactly what Amazon RDS blue/green switchover guardrails are built to do by running automated pre-switchover checks before renaming instances/endpoints.

Active/active migration keeps old and new systems both live and serving real traffic. This approach works well with systems that require very high availability and can justify the engineering effort needed to synchronize in both directions. The danger is conflict resolution when two write paths are both live. Mitigation: limit active/active use to read-heavy workloads initially, deferring resolution of write conflicts to a limited and well-tested set of operations.

With dual-write migration your application writes to both the old and new databases during a window of time instead of relying on database level replication. Ideal for teams that want application level control over how consistency is handled. Risk: data can silently go out of sync if one side of the write succeeds and the other fails. Mitigation: log the result of every dual-write and perform reconciliation checks constantly. Don't rely on your application layer to prevent divergence.

Suitability checklist for any pattern:

  • Confirm downtime tolerance in minutes, not vague terms, before picking a pattern.
  • Check whether the schema itself changes or only the data store.
  • Assess whether your team can realistically operate two live systems during the transition.

Test the workload that you think will most likely violate the pattern first.

Sequencing and wave planning for your migration

A migration roadmap without sequencing is merely a wishlist. Wave planning prioritizes your list of systems into a sequenced, de-risked timeline.

  1. Map out technical dependencies first. List all the services, jobs and integrations that write to or read from each database. Don't forget batch jobs and reporting pipelines that teams forget about until cutover day.
  2. Map business dependencies second. Mark off which systems impact billing, compliance reporting or customer-facing SLAs, as these have larger blast radiuses even when the technical dependency graph may seem trivial.
  3. Batch cautiously. If you're not sure about a workload, assign it to a later wave than the ones you feel confident with. Waves with variable confidence are where slips occur.
  4. Rank each candidate based on business value versus effort to migrate. Move quick wins early to gain momentum and validate your blueprint; major transformations should be migrated only after the blueprint has been validated on quick wins.
  5. Build in scheduling buffers. The Microsoft Cloud Adoption Framework advises workload groupings should be determined by dependency type and criticality. Workloads with quick wins should be prioritised early and critical systems should only be scheduled after the migration capability has been proven.

Here's a realistic example of a wave template. Wave 0 encompasses shared infrastructure and truly non-critical systems, the ones where a failure costs you an afternoon's work, not a customer. Wave 1 includes deliberately one representative complex workload. This workload should be chosen such that it closely mirrors your most difficult remaining systems in terms of schema shape, traffic pattern and integration count. Subsequent waves increase in criticality only after Wave 1 has exercised the runbook, rollback path and monitoring setup under realistic conditions.

Tip: Always approach Wave 1 like it's a practice run for your toughest system. Don't blow through it on your easiest system. You will learn about your true risks on your biggest wave, not your second-biggest wave.

Zero-downtime and near-zero strategies in practice

Zero-downtime migration isn't really one technique so much as three: synchronization, cutover and safeguards to prevent a failed cutover from turning into an outage.

Change data capture (CDC) approaches stream from reading the source database' transaction log and applying the changes to the target database in near real time. Per IBM's guidance on zero-downtime migration, this is typically done in conjunction with a historical backfill and a synchronisation pipeline. Both hosted services and home-grown tooling are capable of the replication and schema transformation required. From an operations standpoint, the backfill phase is critical because it frequently needs to run against a production system: always run bulk exports during your off-hours, and measure the time taken per-table since most migrations will require a short lockout on each table even for continuous migrations, though Google Cloud's documentation for CDC migrations against PostgreSQL mentions a window of less than ten seconds per table for the initial snapshot.

Source DBWrites activeCDC stream openCatch-upPause writesLag → 0Green DBPromote targetResume writesBackfill + syncGuardrails + timeoutEndpoint switch

Blue-green orchestration happens in the same exact order every time once CDC has brought your target all the way up to your source: stop writes against the source; wait for replication to catch up 100%; rename instances/endpoints; unleash your writes against your new production environment. Amazon RDS enforces the strict ordering here with a switchover timeout defaulted at 300 seconds, and with a suite of readiness guardrails that will outright block the switchover if the environments aren't in sync. This is the type of behaviour you want to see: if your migration tool knows better than to proceed, that's safer than if it proceeded anyway.

Bidirectional sync plus active/active setups allow both databases to be writable during transition, eliminating single-direction coupling at the expense of elevating conflict resolution to a first-class engineering problem. Don't make this trade unless your availability requirements absolutely call for it; while there's an upfront cost to split-brain resolutions, the complexity cost is ongoing.

Metrics to watch continuously during any sync-based migration:

  • Replication lag, with an alert threshold well below your cutover tolerance.
  • Error rate on the replication pipeline itself, not just on application traffic.
  • CPU and I/O on source and target, because if your source is saturated it will silently increase your lag.
  • Row-count and checksum drift between source and target tables.

One switchover guardrail to steal outright for your own runbook: Amazon RDS blue/green deployments have a default switchover timeout of 300 seconds, after which they will not complete the cutover if replication has not genuinely caught up. This gives you a built-in, enforced floor under your worst-case downtime window.

Tooling, automation and integration considerations

The correct automation can eliminate entire classes of human error, not just engineer hours, from the cutover.

The automation candidates worth building or buying first:

Version control of your schema, so that every change to the structure is tracked, reviewed and reversible rather than made by hand.

  • Automated backfill jobs that can restart from a checkpoint instead of zero upon failure.
  • Cutover and rollback scripts developed and tested as code, not cut and paste from a document.
  • Traffic forwarding APIs allowing you to programmatically swap traffic between source and target instead of changing DNS or config manually.

Run them alongside your delivery pipeline. Don't set migration up as its own project lane. Integrate schema changes into the same CI/CD pipeline as your application code. Gate promotions behind the same monitoring queries you already rely on. Control which services read from the new database before you cut over completely with feature flags. The goal is to keep your migration work surfacing on the same dashboards your team stares at every day, not in its own spreadsheet.

Vendor familiarity is less important than three factors when choosing tooling: does it allow you to enforce switchover guardrails and configurable timeouts? Does it expose a promotion API you can script against? Does it provide managed CDC capabilities or allow you to insert your own pipeline if your transformation requirements are unique? Teams may find putting a lightweight API wrapper around their legacy system while they do an expand-contract migration buys them time to incrementally migrate consumers while continuing to push features.

Tip of the day: Create your rollback script prior to creating your cutover script. If you cannot automate how to get back to a known-good state, you are not ready to automate moving forward.

Testing, validation and fail-forward rollback planning

Untested rollback plans aren't rollback plans, they are wishful thinking. The space between “we have a rollback plan” and “we have tested our rollback plan” is where most cutover-day disasters occur.

Mandatory tests before any production cutover:

Dry-run your schema against a full production copy of your data, not against a trimmed down sample set. Constraint violations and long index rebuild times may not surface until production scale.

  1. Shadow traffic testing. Route a copy of live production requests to your target system, but don't serve responses from it.
  2. Data reconciliation. Compare row counts, checksums and sampled record contents between source and target after backfill and again after sync.
  3. Load testing on the target by itself. Verify that it can handle expected production traffic volume prior to being put in that position.
  4. Chaos and rollback drills. Intentionally fail the cutover in staging so you can verify the rollback path actually restores a known-good state, instead of just confirming that rollback doesn't error.

Rollback and fail-forward criteria must be numerical values instead of ad-hoc decisions made under stress. The Microsoft Cloud Adoption Framework advises a fail-forward approach: specify deterministic failure criteria ahead of time, such as CPU utilization sustained above x% for y minutes or error rate exceeding z% baseline, and stage test those conditions exactly to ensure your rollback automation actually restores to a known-good state as opposed to believing it will.

Thresholds need to be set against your baseline, not treated as magic numbers that apply everywhere. But having a documented, numeric threshold instead of “we'll see how it feels” is what separates rollback from panic.

Don't use your fail-forward path in staging until you've practised failing forward at least once. Also automate the verification step so that you don't have a human making the judgement call, under pressure, if the rollback succeeded.

Runbook and migration checklist for cutover day

A well-designed runbook is intentionally boring: each step is documented, assigned, and tested before go-live day.

Pre-migration:

  • Ensure full backups are created and restored at least once to prove they work, not just take up space.
  • Obtain formal stakeholder agreement on cutover window, rollback criteria and decision authority for the day of cutover.
  • Capture baseline values for replication lag, error rate and resource usage dashboards prior to migration.
  • Finalise schemas and perform an end-to-end full replication test against production-scale data.

Cutover:

  • Pause or reroute writes if necessary to meet your pattern's write pause requirement, then verify that the write pause is actually occurring.
  • Run the final backfill pass and confirm reconciliation checks pass clean.
  • Promote the target system using your tested script instead of manually typing commands.
  • Run health checks against the new production environment before switching endpoints or DNS.

Post-migration:

Quiesce the old system and ensure replication is stopped safely. This should only happen after the new system has run cleanly in the face of live traffic for a negotiated period of time — the stabilisation period.

  • Archive rather than delete the old database in case something bad is discovered late.

Record lessons learned as soon as possible, and plan a post-mortem even if the migration went cleanly.

PhasePrimary ownerKey gate before proceeding
Pre-migrationEngineering leadBackup restore verified, sign-off received
CutoverOn-call migration teamReconciliation checks pass, health checks green
Post-migrationPlatform/ops teamStabilisation period complete, old system archived

How we structure migrations for mission-critical systems

Mission critical systems don't tolerate an untested cutover. That's why we treat migrations as a series of steps instead of a point-in-time activity. Our approach to datacenter platforms, financial systems and regulated spaces looks the same from engagement to engagement.

  • Evaluate first. We need to understand dependencies, level of criticality and compliance constraints before we can recommend a blueprint. The optimal pattern for a billing system is not often the optimal pattern for an internal reporting tool.
  • Pilot wave. Execute a representative, intentionally hard workload early. Think of Wave 1 from the sequencing framework above. This reveals hard problems early when there's less at risk.
  • Wave execution. Expand the validated pattern to remaining systems, modulating speed based on each wave's comfort level.
  • Managed cutover and support. We don't hand over a runbook and walk away. We remain involved through cutover and the stabilisation period that follows.

It's why we involve monitoring and ops teams from day one instead of cutover time, covered in more detail in our article about datacenter mobilisation: telemetry arriving on day of cutover can't tell you anything useful about the weeks of sync leading up to it. For companies where no downtime is not negotiable, it's why blue-green and CDC-based cutovers are actually safe and not just theoretically well planned. Migrating SCADA/OT telemetry into the cloud follows the same process, which we outline in our guide to migrating SCADA workloads into the cloud, and you can see how well we execute on it through our 99.9% uptime SLA and track record of delivering over 250 projects across regulated industries.

Choosing a risk profile and aligning stakeholders

It doesn't matter what plan you choose. What matters is whether your sponsors have agreed to, in writing, what acceptable risk looks like before cutover day. Define technical thresholds based on business tolerance levels upfront: if a sponsor says “we can't have any downtime,” explain to them in advance, and in numbers, what replication lag or error rate will actually cause a rollback.

Let your first successful wave prove your case, don't just show progress. A successful Wave 1 where nothing broke is the key card that will allow you to play your high-risk waves later on — it is worth infinitely more than any slide deck you could present. Report status at regular intervals on cutover day, not just if there is a problem: lack of communication causes sponsor anxiety, and anxious sponsors make bad decisions.

— Harry

How we can help with your migration

Reading a blueprint is one thing. Implementing against a live financial system, datacenter monitoring platform or regulated compliance environment is something entirely different, and that is where we spend most of our time. We provide assessment, pilot wave execution, managed cutover support and legacy modernisation as standalone services built around these sequencing and guardrail principles — not a one-size-fits-all checklist.

If your migration involves BMS, PMS or NMS integration, or you're modernising a legacy platform where an uncontrolled outage is not an option, our legacy modernisation team can scope a pilot wave against your actual dependency graph rather than one we hypothesize. Teams that are considering a larger rebuild in conjunction with the migration usually have our SaaS development work happening in parallel to these cutovers.

The following logical step is a brief assessment call to map out your dependencies and select the appropriate blueprint for your unique systems. Click over to PODTECH to schedule one.

FAQ

How is data migration done?

Data migration generally occurs in four steps: analyzing source data and dependencies and mapping them to the target system, selecting a migration pattern such as big-bang, phased, or CDC-based blue-green, performing backfill and synchronisation, and finally validating and cutting over to the new system. The order of these steps may vary based on acceptable downtime and the number of downstream dependencies on the source database.

Which is better, Flyway or Liquibase?

Flyway and Liquibase are both database schema migration tools. They keep track of, and automatically apply, changes in your database structure as part of your migration pipeline. The choice between Flyway and Liquibase mostly comes down to your stack and tooling rather than one being better than the other. Development teams that are invested in Java-centric tooling tend to prefer Liquibase due to its wider format support. Development teams looking for simpler, SQL-based migrations tend to prefer Flyway's lightweight approach.

What are the 7 migration strategies?

Definitions differ slightly by source, however a generally accepted list consists of: big-bang, phased/incremental, expand-contract, blue-green, active/active, dual-write and trickle migration. AWS offers prescriptive advice on strategies, defining several of these strategies as offline, flash-cut, active/active and incremental cutover strategies depending on the acceptable downtime window and application architecture.

How do you transfer data between databases?

Migration between databases typically takes place either through a one-off bulk export/import process, suitable for smaller, less mission critical systems, or through a live replication pipeline with change data capture if some degree of downtime is unacceptable. The appropriate strategy depends on the volume of data being moved, your tolerance for downtime, and whether you are also modifying the schema at the time of migration.

How long does a database migration typically take?

Timescale will differ wildly depending on data size and pattern selected, however most zero-downtime strategies pack most of the risk into small periods of time: rolling migrations typically have less than 10 seconds of per-table lockout during the initial copy, while blue/green migrations are limited to the default timeout of roughly 5 minutes for the actual switchover. Backfill and sync work can range from hours to days depending on data size, while user-visible downtime remains small.

Sources

Recommended