AWS 289: Migration portfolio waves, business dependencies, test gates, and ownership
Why this lesson matters
AWS287 governs one wave and its cutover. Portfolio governance keeps many future waves feasible at the same time. It must detect that three applications need the same DBA, two release trains collide with quarter-end freeze, a shared identity platform moves after its dependents, the test environment cannot host concurrent rehearsals, a target account is not vended, and eight applications have technical contacts but no business acceptance owner.
A portfolio board that shows “80 percent complete” while hiding these constraints creates last-minute deferrals and unsafe pressure. Good governance makes demand, capacity, dependencies, evidence, decisions, exceptions, money and outcomes visible early enough to act.
Outcomes
By the end, you can:
- distinguish foundation, governance, portfolio, migration and supporting workstreams;
- maintain one controlled source of portfolio metadata and confidence;
- create a rolling 12-week pipeline without falsely freezing distant plans;
- model cross-wave business, technical, organizational and calendar dependencies;
- plan scarce people, test environments, bandwidth, quotas and vendor capacity;
- define gate evidence, authority, expiry, exception and escalation;
- use RACI correctly without confusing “consulted” with accountable ownership;
- track risks, issues, assumptions, dependencies and decisions separately;
- connect spend and delivered benefits to applications and waves;
- design governance meetings that decide rather than merely report; and
- rebalance a supplied 40-application portfolio from evidence.
The four core workstreams
| Workstream | Primary responsibility | Output consumed downstream |
|---|---|---|
| Foundation | landing zone, network, identity, operations and trained people | ready platform and capacity |
| Project governance | budget, schedule, communication, escalation, reporting and business outcomes | decisions, priorities and resolved constraints |
| Portfolio | metadata, prioritization, strategy, deep dives and wave plans | migration-ready application packets |
| Migration | build, replicate, test, cut over and hypercare | actual migration evidence and lessons |
Security/compliance, cloud operations, application testing, specialized workload and vendor teams often support them. Assign one project technical lead across workstreams so local optimization does not damage the end-to-end outcome.
Information moves both ways. Portfolio feeds scope/readiness to migration; migration returns actual timings, defects and new dependencies so portfolio criteria improve. Foundation capacity and governance decisions constrain both.
Governance is not project administration
Governance must answer:
- Are business outcomes, deadline, scope and budget still valid?
- Which committed waves are ready, and what evidence is missing?
- Which constraints threaten multiple waves?
- Who can decide, fund or accept each exception?
- Are teams operating within safe capacity?
- Are benefits and source decommissioning actually occurring?
- What has changed since the last approved baseline?
Status collection can be automated. Governance time should resolve conflicts, allocate resources, accept/reject risks and preserve decision evidence.
A controlled portfolio data model
Maintain canonical records with stable IDs and source/provenance:
Application
Business capability, criticality, environments, lifecycle, strategy, owner, technical lead, users, RTO/RPO, data classification, dependencies, cost and confidence.
Move group and wave
Group ID, application/component members, hard/soft dependency reasons, target account/Region, planned phases/cutover, migration pattern, owner, readiness, capacity demand and status.
Work item and gate
Evidence requirement, artifact/link/hash/version, owner, approver, due/expiry date, pass/fail/waiver, decision and comments.
Resource/capacity
Named team/role or system, available units/time, committed demand, contingency and collision.
RAID and decision
Type, statement, impact, probability/severity, owner, action, due date, escalation, state and decision history.
Do not let each dashboard redefine application names, owners and statuses. Use controlled vocabulary, validation, change history and access control. Preserve facts separately from predictions.
Metadata completeness and confidence
A row is not complete because every cell contains text. Define mandatory metadata by migration pattern and stage. For example, MGN rehost needs source identity, supported OS/disk, target account/subnet/SG/instance, owner and test/rollback fields; DMS needs engine/schema/logging/key/LOB and data owner evidence.
Measure:
mandatory field completeness = valid required fields / expected required fields
owner confirmation = records confirmed by accountable owner / in-scope records
freshness = records reviewed within allowed age / in-scope records
conflict rate = unresolved contradictory fields / compared fields
Report critical exceptions, not just percentages. One missing payment database owner can be more important than 100 complete low-risk servers.
Rolling 12-week planning horizon
Use three horizons:
- Weeks 1–4 committed: scope and dates controlled; readiness remediation nearly complete; changes require formal approval.
- Weeks 5–8 prepared: deep dives, target design and ownership underway; dates are forecast and capacity reserved.
- Weeks 9–12 candidate: prioritized pipeline with known gaps; movement is expected as evidence improves.
Keep four or five waves in preparation so migration teams do not starve, but never label candidate dates as commitments. At each weekly refresh, promote only when entry gates pass and demote/defer with a recorded reason. Recalculate downstream effects.
Portfolio dependency graph
Add more than network edges:
- application/API/database/file/message dependencies;
- shared identity, DNS, PKI, network, proxy and observability foundations;
- common database, license, appliance or vendor;
- business process order, such as customer onboarding before billing;
- shared product/application owner or support team;
- release-train/version compatibility;
- data retention/residency/regulatory review;
- facility/contract exit or hardware end-of-support deadline;
- common test data/environment; and
- funding/procurement prerequisite.
Represent direction, criticality, tolerance, evidence, owner and treatment. A shared team creates a scheduling dependency even when applications never communicate.
Cross-wave ordering rules
Examples:
- shared platform must be ready and tested before dependent applications;
- development/test environment proves the pattern before production unless business constraints justify otherwise;
- external partner certification completes before endpoint cutover;
- target account/network/security quotas pass before replication starts;
- source license/hardware deadline sets a latest finish but does not waive tests;
- paired producer/consumer versions preserve contract compatibility; and
- decommission follows the last dependent, retention and acceptance, not first migration.
Detect cycles. A circular “A must move before B, B before C, C before A” often reveals an unmodeled hard move group or need for temporary bridging/version compatibility.
Capacity planning across waves
Build a weekly capacity ledger for:
- portfolio analysts and application owners;
- migration engineers and automation;
- network, IAM, DNS, security, database and mainframe specialists;
- testing/UAT teams and business approvers;
- cloud operations/service desk/hypercare;
- vendors and change advisory board;
- test environments and representative data;
- Direct Connect/VPN/DataSync/DMS/MGN throughput;
- account vending, subnet IPs, quotas and target compute; and
- budget/procurement approval.
Use effort units or hours by phase, not headcount alone. A person at 20 percent capacity cannot attend three simultaneous cutovers. Reserve contingency for defects and rollback; 100 percent planned utilization guarantees queues.
available safe capacity = nominal capacity - BAU/on-call - leave - contingency
overload = committed demand + forecast demand - available safe capacity
Show overload by week/role. Rebalance scope or dates; do not solve it by silently assuming overtime.
Test capacity is a portfolio constraint
Testing requires environments, data, integrations, licenses, automation and business users. Track:
- environment booking and reset time;
- target/source version and configuration;
- masked representative data and refresh duration;
- external partner/window availability;
- performance-tool and license capacity;
- test type, owner, evidence and defect retest time;
- shared downstream rate limits; and
- UAT approver calendar.
Prevent two waves from corrupting one shared test environment. A green unit-test result does not satisfy integration, data, performance, recovery, security or business acceptance gates.
Gate catalog
Standardize gates while allowing pattern-specific evidence.
| Gate | Minimum evidence | Accountable approval |
|---|---|---|
| Portfolio intake | owner, capability, scope, strategy/confidence | portfolio/business owner |
| Deep-dive complete | dependencies, target, data, RTO/RPO, cost | application/architecture |
| Foundation ready | account/network/IAM/security/quotas/operations | foundation/security |
| Build ready | approved design, access, tooling, rollback assumptions | migration lead |
| Test ready | environment/data/cases/owners and isolation | testing/application |
| Cutover commit | tests, replication, change, staff, runbook, rollback | business/change authority |
| Acceptance | SLO/data/business/security/backup/hypercare | service owner/operations |
| Decommission | no callers, retention/legal/license/cost evidence | asset/data/business owners |
Each gate has versioned evidence, expiry and pass/fail/conditional status. “Conditional pass” requires explicit condition, owner, deadline, residual risk and escalation if missed.
Exceptions and waivers
Never turn a failed gate green to preserve a date. Create a waiver containing:
- exact unmet control and evidence;
- reason and alternatives considered;
- affected applications/waves and worst credible impact;
- temporary compensating controls;
- risk owner and independent approver at correct authority;
- expiry, remediation owner/date and monitoring trigger; and
- rollback/revocation condition.
Some gates are nonwaivable by policy or law. Repeated waivers indicate a foundation/process defect requiring program action, not application-by-application normalization.
RACI and real ownership
For every decision/task have exactly one Accountable owner where practical. Responsible executes, Consulted supplies input, and Informed receives communication. A name in a cell is not capacity, authority or commitment.
Validate:
- person/role accepted the assignment;
- manager allocated time;
- required access and competence are available;
- delegate/backup exists for critical windows;
- authority matches financial/risk/business decision; and
- departure/leave/on-call conflicts are reflected.
Application owner, business owner, technical owner, data owner and operations owner may differ. Do not let the migration team accept business functionality on their behalf.
RAID and decision discipline
- Risk: uncertain future event; probability/impact and mitigation.
- Issue: event already happened; containment/resolution and escalation.
- Assumption: unverified condition used for planning; validation/expiry.
- Dependency: another deliverable/team/event required; provider/consumer/date.
- Decision: chosen option with evidence, authority and consequences.
Do not mix them in one vague “risk log.” Link each to applications, waves, budget/schedule and gate. Aging and breached due dates trigger escalation. Closing requires evidence, not a comment saying “handled.”
Governance layers and cadence
| Forum | Cadence | Decision focus |
|---|---|---|
| Strategic steering | monthly/exception | business scope, budget, contracts, severe escalations/outcomes |
| Program governance | weekly | cross-workstream capacity, schedule, funding, major RAID/decisions |
| Workstream reviews | several times weekly | deliverables, blockers, quality and next commitments |
| Application-owner commit | before wave commitment | owner availability, evidence, freeze/cutover acceptance |
| Infrastructure/operations | weekly and pre-wave | platform capacity, incidents, next sprint and handoff |
| Migration business hours | scheduled/open | application-owner questions and gap resolution |
| Cutover command center | event-based | AWS287 execution decisions |
Every meeting has current input, threshold-based exceptions, decision rights, action owner/date and published minutes. Cancel status meetings that neither decide nor unblock.
Traffic-light status with objective rules
Define status from evidence:
- Green: mandatory gates pass, no breached critical dependency, capacity reserved.
- Amber: recoverable gap with owner/date inside commitment threshold and accepted risk.
- Red: mandatory gate failed, critical unknown/overload, or forecast misses latest safe date.
- Grey: not assessed or evidence expired, never treated as green.
Show trend and forecast, not only present state. An application can be currently amber but forecast red because its owner due date exceeds wave commit.
Financial governance and business case
Track per application/wave:
- one-time discovery, remediation, migration, testing and partner labor;
- target infrastructure/platform/security/observability/support;
- source-target parallel run and migration tools/data transfer;
- licenses/contracts and termination dates/penalties;
- contingency and realized defects/delay;
- source decommission savings and avoided capital expense; and
- expected business value with owner, baseline and realization date.
Separate forecast, committed, actual and benefit. A migration that creates AWS spend while source hardware/licenses remain is not financially complete. Chargeback/showback should not encourage teams to hide shared migration resources.
Portfolio metrics that do not mislead
Use a balanced set:
- applications/move groups by stage and confidence;
- committed versus migrated versus accepted versus decommissioned;
- readiness first-pass and gate defect rates;
- aging ownerless/unknown dependency/waiver items;
- forecast versus actual wave time/cost;
- resource overload and test-environment contention;
- deferral/rollback/incident causes;
- metadata freshness/conflict/completeness;
- hypercare exit and operations acceptance;
- source cost removed and benefit realized; and
- runbook automation/lessons applied.
Server count rewards easy volume and can hide value/risk. Never count a retired server and a regulated revenue application as equivalent progress.
Change control and baseline
Baseline committed waves with scope, target, dates, strategy, assumptions, capacity and gate versions. Change requests state reason, impact on dependencies/cost/capacity/tests, options, approver and effective version.
Discovery is expected to change candidate waves. It must not silently change a committed cutover. Emergency security/business changes can proceed through an explicit expedited path with evidence and retrospective review.
Diagnose portfolio failure
| Symptom | Likely governance cause | Corrective action |
|---|---|---|
| Migration team waits for work | portfolio pipeline fewer than prepared waves | plan ahead; enforce deep-dive metadata SLAs |
| Every wave defers at test | test capacity/evidence planned too late | put environment/data/UAT gates earlier |
| Same specialist blocks waves | no capacity ledger/backup/cross-training | stagger waves, allocate backup and train |
| Green until cutover week | subjective status or expired evidence | threshold/forecast status and evidence expiry |
| Budget high, savings absent | source decommission/benefit outside governance | track acceptance-to-decommission and contract closure |
| Waivers keep recurring | foundation or standard pattern defect | escalate root cause and fix shared capability |
| Dashboard totals disagree | multiple IDs/sources/status definitions | canonical data model, provenance and reconciliation |
Hands-on workshop: govern 40 applications
Create a 12-week portfolio workbook with:
- canonical application/move-group/wave metadata and confidence;
- technical/business/organizational dependency graph;
- committed/prepared/candidate horizons;
- weekly role/system/test/bandwidth/quota/budget capacity ledger;
- blackout, release, vendor, leave and change calendars;
- gate catalog, evidence versions and expiry;
- RACI acceptance and backups;
- linked RAID/decision/waiver registers;
- governance calendar, thresholds and escalation paths;
- forecast/actual cost and benefit/decommission tracking;
- balanced dashboard with objective status; and
- rebalanced waves plus change/decision history.
Inject: eight ownerless apps, shared Oracle DBA overload, account-vending delay, subnet-IP shortage, quarter-end blackout, one shared test environment, expired security evidence, funding shortfall and a repeated backup waiver. Show which waves move, what can be fixed centrally and who decides.
Knowledge check
- How does this lesson differ from AWS287?
AWS287 executes one wave; AWS289 governs the portfolio pipeline and shared constraints across waves.
- Why plan multiple waves ahead?
To keep enough assessed, ready work flowing to migration teams while preserving time to resolve gaps.
- Is a filled RACI cell proof of ownership?
No. The person needs acceptance, authority, capacity, access, skill and backup.
- Why separate risks and issues?
Risks are uncertain future events; issues already occurred and require containment/resolution.
- What is wrong with overriding a failed gate to green?
It hides risk and removes accountability; use an explicit approved, expiring waiver.
- Why track decommissioning and benefit after migration?
Target launch alone may increase cost and does not prove intended business/financial outcome.
- Why is server count a weak success metric?
It ignores application value, complexity, acceptance, incidents and removed source cost.
Lesson acceptance
Pass only if the 40-application plan reconciles metadata, dependencies, capacity, calendars, gates, owners, exceptions, money and outcomes with traceable decisions. Reject portfolios with subjective green status, names without accepted capacity, hidden waivers, conflicting IDs, overloaded shared resources, unversioned committed-wave changes, no test/UAT constraints, or completion measured only by migrated servers.