AWS 390: Release rollback, roll-forward, feature flags, and database-change compatibility
Why this lesson matters
Returning traffic to old code is only safe when old code, data, events, configuration, secrets, dependencies, and clients remain compatible. Recovery may require traffic rollback, code rollback, feature disable, compatibility mode, data compensation, or roll-forward. The fastest technical switch is not always the safest business recovery.
Recovery decision model
| Option | Best fit | Main limit |
|---|---|---|
| Traffic switch to blue/old Region | Old environment still healthy and state-compatible | Sessions/writes/external effects remain |
| Code/artifact rollback | Immutable prior release can run now | New schema/data/events may be unreadable |
| AppConfig feature disable/kill switch | Fault isolated behind remotely evaluated control | Retrieval/cache/default and hidden side effects |
| Compatibility/read-only mode | Preserve safe operations while repair occurs | Reduced business capability and queue backlog |
| Roll-forward | Cause known and safe fix faster than reversal | Requires build/test/deploy under pressure |
| Data restore/compensation | Corrupted state or irreversible operation | Data loss, reconciliation and business approval |
Predefine decision inputs: exact release/config/schema versions, exposure, first/last bad event, state writes and consumers, external side effects, data integrity, user SLI, previous artifact/capacity, repair estimate, RPO/RTO, compliance, and accountable incident commander. Stop expansion first; preserve evidence; avoid simultaneous untracked changes.
Expand-contract data evolution
Use backward- and forward-compatible stages:
- Expand schema/API/event contract with optional fields, nullable/additive structures, and old-reader compatibility.
- Deploy code that can read old/new and writes in a controlled compatible form, often with dual-read/write only when reconciled and observable.
- Backfill in idempotent bounded batches with checkpoints, throttling, validation, and rollback/compensation.
- Switch readers/writers and observe mixed versions, queues, replicas, caches, analytics, and external consumers.
- Contract/remove old structures only after every consumer is proven upgraded and the rollback window closes.
Avoid rename/drop/type narrowing as one deployment. Database rollback scripts can destroy new writes. Track schema migration ID, transaction boundaries, locking, replica lag, storage, duration, and restore point. Events require versioning, tolerant readers, idempotency keys, ordering and replay strategy; a code rollback does not remove already-published events.
Feature flags and AppConfig
Separate deployment from release with server-side flags. Define owner, purpose, default/fail behavior, target segments, dependencies, exposure metric, expiry/removal date, audit, and test matrix for both states. Flag evaluation must be fast/cached, but stale cache and control-plane outage require a safe explicit default. Never use a client-visible flag as an authorization boundary.
AppConfig can validate syntax/semantics and deploy gradually; CloudWatch alarms can roll back configuration during deployment, including documented missing-data behavior. Reverting configuration does not undo side effects. Test alarm actions, agent/cache behavior, entity consistency where used, and emergency StopDeployment/revert authority. A kill switch must be rehearsed and independently reachable during application failure.
Recovery runbook and evidence
detect and declare -> identify exact bad revision and exposure
-> stop rollout / reduce traffic / disable feature if safe
-> classify stateless versus stateful and compatibility
-> choose rollback, compatibility, compensation, or forward fix
-> execute one controlled action
-> verify release + data + user SLI through bake window
-> reconcile artifacts, flags, schema, queues, audit and cleanup
Evidence includes source/artifact/config/schema IDs, deployment and flag history, traffic weights, alarms, traces/logs, database migration/backfill checkpoints, event offsets/DLQ, external calls, approval timeline, user outcome, and residual work. Declare recovery only after delayed effects and backlog are stable.
Workshop and failure matrix
Given a checkout service release that adds a column, emits event v2, changes tax API, and enables a flag to 10 percent, design expand-contract stages and recovery choices. Calculate canary observations, define alarms and missing data, build compatibility matrix for old/new code/schema/event, and decide actions for code error, slow query, bad backfill, duplicate event, dependency regression, and incorrect tax side effect.
Analyze 20 failures: prior artifact deleted, old capacity absent, down migration loses writes, migration lock, replica lag, old reader rejects field, new writer omits old field, event replay duplicates, cache mixes schema, flag default unsafe, flag service unavailable, alarm disabled, missing data rollback, segment inconsistent, kill switch untested, external payment sent, traffic rollback split writes, roll-forward tests rushed, restore violates RPO, and dashboard green while backlog grows.
Cost and acceptance
Price parallel environments, retained artifacts/capacity, AppConfig requests/deployments, telemetry, migration/backfill compute and I/O, replica/storage growth, queues/replay, external compensation, and incident/customer impact. This lesson creates nothing.
Submit decision tree, compatibility matrix, five-stage schema/event plan, flag contract, AppConfig alarm design, six scenario decisions, 20 failures, recovery evidence, cost, and cleanup of temporary flags/old schema. Pass requires immutable recovery artifacts, reversible traffic/config controls, state-aware decisions, idempotent data operations, and measured user plus data integrity.