Kubernetes
Zero-Downtime Kubernetes Migration: The Complete Checklist
We recently migrated a fintech company from VMs to GKE with an 11-minute cutover window. Users didn't notice. This is the checklist we used — 50 items across 7 phases that make the difference between a smooth migration and a 3 AM incident.
Why Migrations Fail
Most failed Kubernetes migrations share common causes:
- Insufficient rehearsal — First cutover attempt is the real one
- Data migration afterthought — Applications move, but databases don't sync properly
- No rollback plan — When things go wrong, there's no quick way back
- DNS propagation surprise — TTLs weren't lowered, cutover takes hours instead of minutes
- Missed dependencies — That hardcoded IP nobody remembered
This checklist addresses each of these failure modes.
Phase 1: Pre-Migration Assessment
Before writing a single Dockerfile, understand what you're migrating.
Application Inventory
- ☐ List all services with their current infrastructure (VMs, containers, serverless)
- ☐ Document dependencies between services (who calls whom)
- ☐ Identify stateful vs. stateless services
- ☐ Map external dependencies (third-party APIs, SaaS services)
- ☐ Document current resource usage (CPU, memory, disk, network)
Data Assessment
- ☐ Inventory all databases (type, size, replication status)
- ☐ Identify data that must migrate vs. data that stays (logs, caches)
- ☐ Document current backup and recovery procedures
- ☐ Assess database schema compatibility with managed services
Network Assessment
- ☐ Document current DNS configuration and TTLs
- ☐ Map all network paths (internal, external, VPN)
- ☐ Identify firewall rules and security groups
- ☐ List all hardcoded IPs in configuration and code
Phase 2: Kubernetes Cluster Setup
Build your target environment before touching production.
Cluster Configuration
- ☐ Choose managed Kubernetes (EKS/GKE/AKS) vs. self-managed
- ☐ Configure node pools (instance types, autoscaling limits)
- ☐ Set up cluster networking (VPC CNI, Calico, Cilium)
- ☐ Configure ingress controller (nginx, Traefik, cloud-native ALB/GCLB)
- ☐ Set up cert-manager for TLS certificates
- ☐ Configure cluster autoscaler with appropriate limits
Security Baseline
- ☐ Configure RBAC with least-privilege roles
- ☐ Set up network policies (deny-all default + explicit allows)
- ☐ Configure pod security policies/standards
- ☐ Set up secrets management (External Secrets Operator + Vault/cloud secrets)
- ☐ Enable audit logging
Observability
- ☐ Deploy monitoring stack (Prometheus/Grafana or Datadog)
- ☐ Set up log aggregation (Loki, CloudWatch, etc.)
- ☐ Configure alerting with appropriate thresholds
- ☐ Create dashboards for migration monitoring
Phase 3: Application Containerization
Convert applications to containers with production-ready configurations.
Dockerfile Best Practices
- ☐ Use multi-stage builds to minimize image size
- ☐ Pin base image versions (no
:latest) - ☐ Run as non-root user
- ☐ Scan images for vulnerabilities (Trivy, Snyk)
- ☐ Build multi-arch images if using Graviton/ARM
Kubernetes Manifests
- ☐ Define resource requests and limits for all containers
- ☐ Configure liveness and readiness probes
- ☐ Set up horizontal pod autoscaling (HPA)
- ☐ Configure pod disruption budgets (PDBs)
- ☐ Use ConfigMaps and Secrets (not hardcoded values)
- ☐ Set appropriate replica counts for HA
CI/CD Pipeline
- ☐ Automated image building on push
- ☐ Image scanning in pipeline
- ☐ Automated deployment to staging
- ☐ Deployment approval gates for production
Phase 4: Data Migration Strategy
The hardest part. Plan this carefully.
Database Migration
- ☐ Set up database replication from source to target (DMS, native replication)
- ☐ Verify replication lag monitoring
- ☐ Test failover procedure in non-production
- ☐ Plan for schema changes needed for managed services
- ☐ Document connection string changes needed
File Storage Migration
- ☐ Sync files to cloud storage (S3, GCS) if not already there
- ☐ Set up continuous sync for files created during migration window
- ☐ Update application code to use new storage paths
Cache Strategy
- ☐ Plan for cache warming on new infrastructure
- ☐ Or: accept cold cache and plan for increased DB load during cutover
Phase 5: Pre-Cutover Preparation
Everything that needs to happen before the actual migration.
DNS Preparation
- ☐ Lower DNS TTLs to 60 seconds (do this 48+ hours before cutover)
- ☐ Prepare DNS changes but don't apply them yet
- ☐ Verify DNS propagation monitoring is in place
Traffic Management
- ☐ Set up traffic splitting capability (weighted routing)
- ☐ Configure health checks on new infrastructure
- ☐ Test load balancer failover
Rehearsals
- ☐ Complete full migration rehearsal in staging
- ☐ Time each step of the cutover procedure
- ☐ Practice rollback procedure
- ☐ Do at least 2 full rehearsals before production
Communication
- ☐ Create status page update templates
- ☐ Notify stakeholders of maintenance window (if any)
- ☐ Set up war room communication channel
- ☐ Document escalation contacts
Phase 6: Cutover Execution
The actual migration. This should be boring because you've rehearsed.
Pre-Cutover Checks
- ☐ Verify all pods are healthy on new cluster
- ☐ Confirm database replication is caught up (lag = 0)
- ☐ Verify monitoring and alerting is active
- ☐ Confirm rollback procedure is ready
- ☐ All team members in war room
Traffic Cutover
- ☐ Start with small traffic percentage (5-10%) to new cluster
- ☐ Monitor error rates and latency
- ☐ Gradually increase traffic (25%, 50%, 75%, 100%)
- ☐ Watch for increased database load as caches warm
Database Cutover
- ☐ Stop writes to old database
- ☐ Verify replication is complete
- ☐ Promote new database as primary
- ☐ Update connection strings
- ☐ Verify writes are going to new database
Post-Cutover Verification
- ☐ Verify all endpoints are responding correctly
- ☐ Check error rates are within normal bounds
- ☐ Verify data integrity (spot checks)
- ☐ Confirm no traffic is going to old infrastructure
Phase 7: Post-Migration
The migration isn't done when traffic moves. These steps complete it.
Stabilization
- ☐ Keep old infrastructure running for 24-48 hours (rollback safety)
- ☐ Monitor closely for 24 hours
- ☐ Address any issues that emerge under real traffic
- ☐ Tune autoscaling based on real usage patterns
Cleanup
- ☐ Decommission old infrastructure (after bake period)
- ☐ Reset DNS TTLs to normal values
- ☐ Update documentation
- ☐ Archive migration runbooks
Retrospective
- ☐ Document what went well
- ☐ Document what could be improved
- ☐ Update checklist for next migration
Our 11-Minute Cutover: What Made It Work
For the fintech migration mentioned at the start, here's what made the difference:
- Three full rehearsals — By the third, we had the procedure down to 11 minutes
- Database replication running for 2 weeks — Zero data to sync at cutover
- DNS TTLs at 60 seconds for 3 days — Traffic shifted within 2 minutes
- Gradual traffic shift — 5% → 25% → 50% → 100% over 15 minutes
- Automated health checks — Traffic automatically shifted back if errors spiked
The 11 minutes was the database cutover window (stop writes, verify sync, promote). Everything else happened with zero downtime.
Download the Checklist
Want this checklist in a format you can actually check off? We have it as a Google Sheet you can copy.