Commerce · GCP · GKE
E-commerce Cloud Migration
A Magento platform serving 40 storefronts had to move from a provider-operated legacy stack into a dedicated GCP and GKE platform. I led the technical migration, implemented the critical paths, and put risks in front of the people who had to decide on them.
Production path
A controlled path from the edge to the database
The cutover connected traffic, runtime, data, and production signals in a verifiable order.
- 01Cloudflare edgeDNS, certificates, and controlled traffic switching
- 02GKE runtimeMagento, APIs, jobs, and 32 queue consumers
- 03Cloud servicesCloud SQL, Redis, search, and persistent media
- 04Cutover signalsMonitoring, rehearsal, and explicit approval gates
The scale behind the cutover
These six figures shaped the architecture, rehearsal, and cutover order.
- Production database
- 100+ GiB108.48 GiB in the recorded source baseline.
- Base tables
- 2,465Verified after the Cloud SQL DMS rehearsal.
- Storefronts
- 40Each with its own host, language, and store mapping.
- Queue consumers
- 32Controlled as independent writers during cutover.
- Media entries
- 1M+Covered by the preflighted synchronisation path.
- Less deployment configuration
- 85%Magento configuration values reduced from 1,739 to 254.
Starting point
The platform was much larger than a Magento deployment. Runtime, database, Varnish, search, media, edge rules, store mappings, and external integrations were split across systems and service providers. Some production behaviour existed only in accumulated configuration.
A company-controlled cloud platform reduced dependencies and improved control ahead of peak events. We first mapped the existing production behaviour, then migrated it in controlled steps.
Operating model
From a provider-dependent stack to a company-owned platform
The migration clarified technical ownership and operating boundaries across workloads, data, and traffic.
Starting point
- runtime and operations tied to the incumbent provider
- behaviour split across the edge, Nginx, Varnish, and Magento
- cutover and recovery paths lacked end-to-end evidence
Company-owned platform
- a dedicated GCP and GKE platform
- Controlled delivery with Terraform, Kustomize, and Jenkins
- rehearsed data, media, and writer transitions
What I owned
I led the technical migration and stayed in the implementation. My scope covered the GCP and GKE structure, Terraform stacks, Magento and Kubernetes overlays, Cloud SQL replication, media synchronisation, Cloudflare, Varnish, monitoring, and the cutover runbook.
Many dependencies crossed the application team, platform team, incumbent providers, and project leadership. I turned assumptions into checks, resolved contradictory requirements, and brought technical and business decision-makers the risk, its impact, and the next safe state.
The decisions that shaped the migration
Four decisions defined the migration path. Each reduced the blast radius and gave the team a verifiable recovery point.
Migration design
Four decisions instead of a high-risk big bang
The architecture made failures visible early and kept production states unambiguous.
- 01
Use a dedicated platform lane
Separate GCP projects and clusters kept the migration away from existing production namespaces.
- 02
Prove current behaviour
Routing, cache rules, store mappings, and integrations were tested before we kept or removed them.
- 03
Rehearse the data move
CDC and an isolated, promoted Cloud SQL rehearsal tested the data and application before cutover.
- 04
Encode writer states
Stopped, bootstrap, protected app, and background states made each transition and recovery boundary explicit.
When the migration got difficult
The rehearsals found real defects: incomplete dumps, wrong runtime assumptions, missing database grants, inherited configuration, and blocked external calls. I stopped affected rollouts before those failures could spread.
Every stop produced sanitised evidence, a corrected artefact, and a clear decision. Schedule pressure did not change the safety gates or the burden of proof. Decision-makers received the technical finding, its impact, and the next safe action.
Outcome and handover
The production path moved to the dedicated GCP and GKE platform. Cloud SQL took over the production database through the rehearsed replication path, more than one million media entries sat within the controlled synchronisation flow, and all 40 storefront mappings were checked against the target path.
Terraform, Kustomize, Jenkins, monitoring, and runbooks now support one operating process. The team received the running platform with its decisions, approval gates, and recovery procedures.
Need technical leadership that also delivers the critical path?
Describe the platform, its dependencies, and the current time pressure. I will assess which critical migration path I can own.