Cloud Migration Strategy and Execution Questions
Planning and executing a move to the cloud: the migration strategies (rehost, replatform, refactor, repurchase, retire, retain), legacy assessment, dependency mapping, cutover planning, and rollback. Covers phased migration roadmaps, workload modernization, risk management during cutover, and validating success post-migration. The end-to-end migration lifecycle, not steady-state operations.
Create a high-level cost estimation and sensitivity model for migrating 5 PB of storage plus associated compute workloads to a public cloud. Include assumptions for storage tiering (hot/warm/cold), expected ingress/egress patterns and costs, data transfer acceleration and appliance costs, compute sizing and licensing, expected growth rate, and sensitivity to changes in egress and storage-class pricing.
Sample Answer
Direct answer: Build the cost model as separate line items for one-time transfer cost and ongoing post-migration cost, each with an explicit sensitivity range (not a single point estimate), since a 5PB migration's dominant cost drivers (egress, transfer method, and post-migration storage tiering) each have wide enough uncertainty that a single number would be misleading.
Structured elaboration. Storage tiering assumptions: split the 5PB by access pattern (hot: actively queried, warm: occasionally accessed, cold: archival/compliance-retention-only), since storage-class pricing typically varies by an order of magnitude between hot and cold tiers, and getting the hot/warm/cold split roughly right matters far more to the total estimate than precision on any single tier's unit price. Ingress/egress patterns and costs: ingress (data coming IN to the cloud) is typically free or low-cost on major providers; egress (data leaving, including the one-time migration-out cost if data is being moved FROM another cloud, or ongoing egress for any workload that serves data back out) is usually the dominant and most-underestimated cost component, so the model needs an explicit egress-volume assumption with a stated confidence range. Data-transfer acceleration/appliance costs: compare a physical transfer appliance's flat cost against network-transfer cost at the org's actual available bandwidth; at 5PB, this comparison typically favors an appliance unless the org has substantial dedicated bandwidth already provisioned. Compute sizing and licensing: separate from the storage cost entirely, model the compute needed to actually USE the migrated data (query engines, processing clusters), including any licensing costs that don't disappear just because the data moved (some legacy software licenses are tied to deployment model, not just data location). Expected growth rate: the 5PB figure is a snapshot; the ongoing cost model needs a growth assumption (e.g., X% per quarter) since post-migration storage cost compounds. Sensitivity to egress and storage-class pricing changes: run the model at low/base/high assumptions for both egress volume and storage-class pricing, since these are the two inputs most likely to be wrong in the initial estimate and most likely to move the total by a large margin.
Worked example. At 5PB with an assumed 60% cold / 30% warm / 10% hot split: cold-tier storage cost is typically an order of magnitude cheaper per GB than hot, so the split assumption alone can swing the ongoing monthly storage estimate by 3-5x depending on whether the org's real access pattern is closer to 60/30/10 or, say, 20/30/50. A sensitivity table showing total cost at three different hot/cold splits (rather than one blended number) gives the actual decision-makers a much more honest picture of the range they're committing to.
Trade-offs & pitfalls. Presenting a single point-estimate total cost for a 5PB migration, rather than a range with the dominant sensitivity drivers called out explicitly, sets the project up to look like it's "over budget" the moment reality (an inevitably-imperfect initial hot/cold split, or higher-than-assumed egress) diverges even slightly from the point estimate; a sensitivity-based model manages that expectation upfront.
You need to migrate 10 PB of historical data and live ETL from a legacy on-prem Hadoop cluster to a cloud lakehouse (S3 + Delta Lake) with minimal downtime and no data loss. Prepare a detailed migration strategy: initial bulk transfer approach, live-sync strategy (CDC), validation checks, cutover plan, incremental testing phases, cost estimate, and mitigation for unexpected data drift or format incompatibilities.
Sample Answer
Direct answer: For 10PB of historical data and live ETL with zero data loss and minimal downtime, split the problem in two: bulk-transfer the HISTORICAL data using a physical or high-throughput parallel transfer method on its own timeline (it's static, so there's no downtime pressure on it), while migrating the LIVE ETL pipeline via change-data-capture (CDC)-based continuous sync with a short final cutover, and only cut over application traffic once both are verified caught up and consistent.
Structured elaboration. Initial bulk transfer approach: at 10PB, pure network transfer is almost always too slow and too costly (egress and time) compared to a physical transfer appliance (or, if network capacity genuinely supports it, a scheduled high-throughput parallel-transfer job run over weeks); this decision is a straightforward cost/time trade-off calculation once the org's actual available bandwidth is known. Live-sync strategy (CDC): once bulk transfer of the historical baseline completes, start CDC replication from the point the bulk transfer's snapshot was taken, so the target catches up to "now" and then stays current; this requires the source system to retain enough change-log history to bridge the gap between when the bulk snapshot was taken and when CDC starts consuming from it. Validation checks: row/record-count and checksum comparison at the PARTITION level (not the whole 10PB as one unit, which is both slow and gives poor failure localization), so a discrepancy can be traced to a specific partition/time-range rather than triggering a full re-transfer. Cutover plan: once CDC lag reaches near-zero and partition-level validation passes across the full dataset, pause the live ETL write path briefly, drain final lag, do one last validation pass, and cut the application/ETL over to the new lakehouse. Incremental testing phases: validate correctness incrementally as bulk-transfer batches complete (don't wait for all 10PB before testing anything), so schema or format incompatibilities are caught on batch 3 rather than batch 300. Cost estimate: physical appliance/transfer cost plus cloud storage cost for the target (accounting for storage-class strategy: cold-tier for genuinely historical/rarely-accessed data, not uniformly hot storage for all 10PB) plus the engineering cost of the CDC pipeline and validation tooling. Mitigation for data drift/format incompatibilities: run a schema-compatibility check as an early, cheap gate before committing to the full bulk transfer (validate against a representative sample first), and build the CDC pipeline to explicitly reject and quarantine (not silently drop or silently coerce) any record that doesn't match the expected schema, surfacing it for manual review. Metadata/catalog migration: for a Hadoop-based source, the Hive metastore (the catalog mapping table/partition names to their actual file locations and schemas) has to migrate alongside the raw data, not as an afterthought; a lakehouse target needs that same table-to-file mapping rebuilt or imported, or every downstream query breaks even though the underlying files transferred correctly. Source-specific tuning: if the source is a petabyte-scale system like Oracle Exadata, the target lakehouse layout should be designed around the ACCESS PATTERNS the source's own indexing and partitioning reveal (which columns Exadata's indexes were built on, and which partitioning scheme kept its hot queries fast) rather than a generic flat migration, since those patterns are the cheapest available evidence for how to lay out partitioning and file sizing in the new lakehouse. Business impact: track and report the business consequence of the migration in progress, not just its technical status, e.g. the specific reports or SLAs that depend on data freshness during the transition, so stakeholders see a business-relevant view (data as fresh as before vs. degraded) rather than only "batch N of 300 complete."
Worked example. A pragmatic 3-month incremental phasing: month 1, transfer and validate the metastore/catalog mapping plus the coldest, lowest-business-risk partitions first (proves the pipeline and the Hive-metastore rebuild work before anything business-critical rides on it); month 2, transfer the bulk of the remaining historical data in parallel batches by partition/time-range, each validated via checksum before being marked complete, while CDC replication starts catching up live ETL changes since the month-1 baseline snapshot; month 3, close the remaining gap, run the final partition-level validation pass across the full 10PB, and execute the brief write-pause cutover. Concretely, in calendar terms: weeks 1-8 cover the bulk historical transfer described above, week 9 onward is CDC catch-up running in parallel with the tail of the transfer, and the final validation-and-cutover window lands in month 3 once CDC lag is confirmed near-zero.
Trade-offs & pitfalls. Attempting to network-transfer all 10PB directly without first modeling the actual achievable throughput against available bandwidth is a common and expensive mistake; the appliance-vs-network decision should be made from a real bandwidth-and-cost calculation, not a default assumption that network transfer is always simpler.
Describe a practical, enterprise-grade IAM design that spans multiple clouds. Explain how you would implement SSO for human users, service identities for cloud-to-cloud access, centralized identity governance.
Sample Answer
Direct answer
An enterprise-grade identity and access management (IAM) design across AWS, Azure, and Google Cloud needs one identity plane feeding all three, not three separately administered ones. Human users authenticate through a single sign-on (SSO) flow from one upstream identity provider (IdP) into each cloud's native federation mechanism; workloads that call across cloud boundaries use short-lived, federated service identities instead of static long-lived keys; and a centralized governance layer, policy authored once and compiled to each cloud's native format, plus every cloud's access logs aggregated into one place, is what keeps least privilege actually true across three structurally different permission models as the organization grows, rather than only at the moment the design was drawn.
Structured elaboration
SSO for human users. A single upstream IdP, itself typically federated from the organization's existing on-premises directory so employees hired before the cloud migration and those onboarded natively afterward share one identity, federates via SAML or OpenID Connect (OIDC) into each cloud's native SSO integration: AWS's identity center, Azure's native directory, and Google Cloud's identity service. The critical design choice is that role assignment is driven by group membership in the upstream IdP, not by a separate account provisioned per employee per cloud: a group like "platform admins" maps to a defined permission set in AWS, a role assignment in Azure, and an IAM role binding in Google Cloud, all three kept in sync from the same source. This is what makes offboarding reliable across three clouds at once; the classic multi-cloud failure mode is a separate, unlinked account per employee per cloud that nobody remembers to disable in all three places.
Service identities for cloud-to-cloud access. No static, long-lived credential should ever cross a cloud boundary. Each cloud's native workload-identity federation lets a service authenticate using a short-lived token issued by its own cloud, exchanged for a scoped, temporary credential in the target cloud: a workload running in Google Cloud that needs to call an AWS API presents a Google-issued identity token to AWS's federation endpoint and receives a temporary, narrowly scoped AWS credential in return, with no persistent key stored anywhere. The mapping of exactly which workload identity in one cloud may assume which role in another needs to be centrally registered and reviewable, not provisioned ad hoc by whichever team needed cross-cloud access first; an unregistered or over-broad mapping here is functionally a standing credential even though no static key exists.
Centralized identity governance. AWS's policy documents, Azure's role-based access control, and Google Cloud's IAM bindings are three structurally different permission models, so governance cannot mean checking each cloud's console independently. Two mechanisms make this tractable: authoring access policy as code in one higher-level representation that compiles down to each cloud's native format, so a least-privilege change is reviewed once and applied consistently everywhere instead of three times by three different reviewers using three different mental models; and aggregating every cloud's identity and access logs, AWS's activity trail, Azure's activity log, and Google Cloud's audit logs, into one central pipeline for anomaly detection and periodic access review. A privilege-escalation pattern that spans two clouds is invisible if each cloud's logs stay in its own silo; centralizing them is what makes that pattern visible at all.
Worked example
"NimbusCo" runs this design across two stages of its own growth. In its earlier stage, NimbusCo operates with 200 teams sharing just 5 cloud accounts (one broad account per business unit across the three clouds), and the group-to-role mapping above is coarse: a handful of upstream IdP groups cover most of the organization's access needs, and centralized governance is mostly a matter of keeping those few mappings correct. As NimbusCo grows to 5000 users spread across 200 accounts (moving to one account per team and environment for real isolation), the same design is what keeps this from becoming unmanageable: without the policy-as-code layer, a least-privilege change would now mean touching some multiple of 200 separate account-level policies by hand across three different permission models, which is exactly the scale at which manual per-account administration stops being realistic and centralized governance stops being optional.
NimbusCo's upstream IdP is itself federated from its original on-premises Active Directory, so an engineer hired before the company adopted any cloud provider authenticates through the same SSO path, and is governed by the same group-to-role mapping, as someone onboarded directly into the cloud IdP years later.
The audit pipeline earns its cost the day a Google Cloud service account, registered only to assume one specific, narrowly scoped AWS role for a data-export job, is observed by the aggregated logs assuming a second, more privileged AWS role it was never registered for. Google Cloud's audit log shows the identity-token issuance for that service account; AWS's activity trail, ingested into the same pipeline, shows the corresponding role-assumption call landing on a role outside the centrally registered mapping. Neither log on its own, viewed inside its own cloud's console, would necessarily stand out: the Google-side event looks like routine token issuance, and the AWS-side event looks like a normal role assumption from a valid, federated identity. Correlated in the shared pipeline against the registered mapping, the combination flags as an escalation before the workload has a chance to use the unauthorized role, which is the entire argument for centralizing the logs instead of leaving them per-cloud.
flowchart TB
OnPrem[On-prem Active Directory]
IdP[Upstream identity provider]
AWS[AWS]
Azure[Azure]
GCP[Google Cloud]
Gov[Centralized governance: policy-as-code and log aggregation]
OnPrem -->|federated| IdP
IdP -->|SSO| AWS
IdP -->|SSO| Azure
IdP -->|SSO| GCP
GCP -->|workload identity federation| AWS
AWS -->|access logs| Gov
Azure -->|access logs| Gov
GCP -->|access logs| Gov
Gov -->|policy as code| AWS
Gov -->|policy as code| Azure
Gov -->|policy as code| GCP
Trade-offs and pitfalls
The upstream IdP becomes a single point of failure for sign-in across all three clouds at once, which is the correct trade for consistent governance but means it needs its own high-availability design and cannot be treated as an afterthought. Group-to-role mappings drift if they are maintained by hand in three different consoles rather than defined once in code and deployed to each cloud; the design only holds if the policy-as-code layer is the actual source of truth, not documentation describing what each console is supposed to contain. Workload-identity federation setup has real upfront complexity per cloud pair, and the most common pitfall is a team quietly falling back to a static, long-lived cross-cloud key "temporarily" to unblock a project, a shortcut that tends to outlive the project and becomes exactly the standing credential the design was built to avoid. A centralized governance layer can also become a bottleneck if every routine, already-compliant change has to wait on a manual review queue; the fix is to make the policy-as-code guardrails expressive enough that most changes are self-service within pre-approved bounds, with manual review reserved for anything outside them. Finally, a log-aggregation pipeline that can detect a cross-cloud escalation but has no paired incident-response runbook only produces a well-documented breach after the fact; detection and response need to be built together, not detection first with response deferred.
You are tasked with migrating petabytes of archival and active data to the cloud while minimizing downtime and cost. Compare strategies including online replication, offline transfer appliances, parallel bulk transfer with WAN acceleration, and staged migration with cutover windows. For each approach, discuss throughput, cost, security controls during transfer, and verification strategies.
Sample Answer
Direct answer: Compare online replication, offline transfer appliances, parallel bulk transfer with WAN acceleration, and staged migration with cutover windows primarily on the throughput-vs-downtime-vs-cost triangle, choosing based on the actual bandwidth available and the downtime budget rather than defaulting to whichever method is most familiar.
Structured elaboration. Online replication (continuous sync while the source stays live and serving): best throughput-to-downtime ratio since there's effectively no dedicated migration window, but requires the most sophisticated tooling (handling ongoing changes, not just a static copy) and the longest total elapsed time to reach full parity. Offline transfer appliances: best for very large datasets where available network bandwidth is the binding constraint; cost is the flat appliance/shipping fee, throughput is bounded by how fast data can be loaded onto and off the device, and security during transfer relies on the appliance's own encryption (verify it's encrypted at rest on the device, not just in transit before/after). Parallel bulk transfer with WAN acceleration (compression, deduplication, TCP optimization to better utilize available bandwidth): improves effective throughput over naive single-stream transfer without needing physical hardware, cost is primarily the acceleration tooling/service fee plus egress, security relies on standard in-transit encryption plus whatever the acceleration layer adds. Staged migration with cutover windows (migrate in scheduled batches, each with its own short downtime window): moderate throughput (bounded by whatever transfer method is used within each stage), predictable and boundable downtime PER STAGE even though the OVERALL project takes longer, cost is typically the lowest of the four since it uses straightforward transfer mechanics without specialized acceleration or hardware. For each approach: throughput, cost, security controls during transfer, verification strategies. Online replication: throughput moderate-to-high (continuous, but rate-limited to avoid impacting the source), cost is ongoing (replication infrastructure runs for the duration), security requires encrypting the replication stream and securing the ongoing connection, verification is continuous checksum/row-count comparison. Appliance: throughput bounded by device I/O and shipping time, cost is a flat fee (can be cost-effective at very large scale despite feeling old-fashioned), security requires validating the device's own encryption and chain-of-custody during physical transport, verification is a full checksum pass after data lands. WAN-accelerated parallel transfer: throughput improved but still bandwidth-bound, cost scales with data volume and acceleration-service fees, security is standard TLS in transit, verification is per-batch checksums. Staged with cutover windows: throughput determined by the underlying transfer method used per stage, cost lowest, security standard, verification per-stage plus a final full-parity check.
Worked example. Concretely: 500TB total, with 1Gbps genuinely available for the migration (measured, not assumed) and a generous 4-month overall timeline but zero tolerance for extended downtime at any single point. At 1Gbps, transferring the full 500TB over the network alone would take roughly 46 days at theoretical maximum (500,000,000MB / 125MB/s = 4,000,000 seconds ~= 46.3 days) -- comfortably inside the 4-month (~120 day) window with wide margin, which confirms bandwidth is NOT the binding constraint in this scenario; downtime-per-cutover is. That's exactly the case staged migration with small windows is built for: split the dataset into 50 x 10TB logical partitions, each transferred in the background (roughly 1 day, about 22 hours, per 10TB partition at the same 1Gbps) and each cut over independently with its own brief, validated 2-hour maintenance window, so the 50 partitions can be staggered across the 4-month window with room to spare, while no single cutover risks more than one partition's worth of downtime.
Trade-offs & pitfalls. Choosing WAN-accelerated parallel transfer purely because it sounds more sophisticated than staged batching, without first checking whether available bandwidth is actually the binding constraint, is a common mistake: if the real constraint is downtime tolerance per system rather than raw transfer speed, staged migration with small cutover windows solves the actual problem more directly than throughput optimization does. Conversely, if the numbers had gone the other way (bandwidth genuinely too low to finish within the deadline, as a naive 200Mbps assumption would show for this same 500TB), the correct move is not to force staged network migration through anyway -- it's to add a physical appliance for the bulk of the colder data or negotiate more bandwidth, since no amount of clever staging changes the total bytes-over-the-wire math.
Tell me about a cloud migration you led or participated in. Specify the public cloud provider(s) used (AWS/Azure/GCP), the concrete services and patterns you chose for compute, storage, networking and managed databases, your role in architecture and deployment, and measurable results (for example: latency reduction, cost delta, availability improvement, deployment frequency). Include any follow-up training or certifications that supported your work.
Sample Answer
Direct answer: The strongest version of this story names the specific cloud provider and concrete services/patterns chosen (not a vague "we moved to the cloud"), explains the candidate's actual role in architecture and execution decisions, and closes with measurable, specific results rather than a general "it went well."
Structured elaboration. Public cloud provider(s) used: name it specifically (AWS/Azure/GCP), since a vague answer here is often an early signal to an interviewer that the rest of the story may also lack specificity. Concrete services and patterns for compute, storage, networking, and managed databases: name actual services for all four, not just the ones that come to mind first (networking in particular is easy to skip since it's less visible than compute or storage) (e.g., "we moved a fleet of on-prem VMs to EC2 behind an Application Load Balancer, provisioned a new VPC with public/private subnet segmentation mirroring our existing security zones and per-tier security groups, ran a temporary Site-to-Site VPN back to the on-prem data center specifically to carry replication traffic during the migration window, migrated the database to RDS PostgreSQL via DMS (Database Migration Service) with change-data-capture (CDC)-based replication for a near-zero-downtime cutover, and moved file storage to S3") rather than generic category names, since specificity here is what lets an interviewer probe deeper and distinguish real hands-on experience from a surface-level description. Role in architecture and deployment: be honest and specific about scope (did the candidate design the migration strategy, execute a specific piece of it, lead the team, or contribute as an individual engineer on a defined workstream); overstating scope tends to unravel under a good interviewer's follow-up questions about decisions the candidate claims to have made. Measurable results: latency reduction (with actual before/after numbers if remembered, even approximate), cost delta (a concrete percentage or dollar figure, understanding this may be approximate from memory but should still be a real number, not "it was cheaper"), availability improvement (a specific uptime or incident-rate change), deployment frequency (if relevant, how release cadence changed post-migration due to new CI/CD capability). Follow-up training or certifications: mentioning relevant certifications or continued learning shows the migration wasn't a one-off task but built lasting capability, which is a positive signal beyond the migration itself.
Worked example. A strong answer: "I was the lead engineer on migrating our order-processing service from on-prem VMware to AWS. We used EC2 with an ALB for the application tier, a new VPC with private subnets for the application and database tiers and a temporary Site-to-Site VPN back to our on-prem datacenter to carry DMS replication traffic securely during the migration window, RDS PostgreSQL with DMS-based CDC replication for the database (targeting near-zero downtime), and moved file storage to S3 with a dual-write period during transition. I owned the database migration and cutover plan specifically, while a colleague led the application-tier work. Post-migration, we measured a 30% reduction in p99 latency (mostly from moving off aging on-prem hardware to modern instance types), a roughly 20% reduction in infrastructure cost after right-sizing, and we went from monthly to weekly deploys once we had the new CI/CD pipeline in place. I got my AWS Solutions Architect Associate certification during the project, partly to make sure I understood the platform deeply enough to make good calls during cutover."
Preparing one story for several framings. The same underlying migration experience gets probed from several different angles across a real interview loop, and it is worth preparing one well-detailed story that can flex to answer each: sometimes the ask is this general "walk me through a migration" framing; sometimes it is narrower, "tell me about a time you had to convince skeptical stakeholders to adopt a particular migration approach," which wants the persuasion and technical-evaluation angle foregrounded instead of the end-to-end summary; and sometimes it is "tell me about a time priorities shifted mid-migration," which wants the adaptability and communication angle foregrounded. Rehearsing the same real project along all three angles, rather than having only one fixed narration of it, means a candidate isn't caught flat-footed when the interviewer's specific phrasing doesn't match the version they rehearsed.
Trade-offs & pitfalls. A common weak version of this answer stays at the category level ("we moved to managed services and it was faster and cheaper") without naming specific services, specific numbers, or a specific role; interviewers use exactly this kind of question to distinguish candidates who did hands-on migration work from those who were adjacent to a project without deep involvement, and specificity is the main signal that separates the two.
Unlock Full Question Bank
Get access to all 44 Cloud Migration Strategy and Execution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.