Lesson 390 · AWS Learning Path

AWS 390: Release rollback, roll-forward, feature flags, and database-change compatibility

· Published · 4 min read

Labelled process diagram for AWS 390: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

Returning traffic to old code is only safe when old code, data, events, configuration, secrets, dependencies, and clients remain compatible. Recovery may require traffic rollback, code rollback, feature disable, compatibility mode, data compensation, or roll-forward. The fastest technical switch is not always the safest business recovery.

Recovery decision model

OptionBest fitMain limit
Traffic switch to blue/old RegionOld environment still healthy and state-compatibleSessions/writes/external effects remain
Code/artifact rollbackImmutable prior release can run nowNew schema/data/events may be unreadable
AppConfig feature disable/kill switchFault isolated behind remotely evaluated controlRetrieval/cache/default and hidden side effects
Compatibility/read-only modePreserve safe operations while repair occursReduced business capability and queue backlog
Roll-forwardCause known and safe fix faster than reversalRequires build/test/deploy under pressure
Data restore/compensationCorrupted state or irreversible operationData loss, reconciliation and business approval

Predefine decision inputs: exact release/config/schema versions, exposure, first/last bad event, state writes and consumers, external side effects, data integrity, user SLI, previous artifact/capacity, repair estimate, RPO/RTO, compliance, and accountable incident commander. Stop expansion first; preserve evidence; avoid simultaneous untracked changes.

Expand-contract data evolution

Use backward- and forward-compatible stages:

  1. Expand schema/API/event contract with optional fields, nullable/additive structures, and old-reader compatibility.
  2. Deploy code that can read old/new and writes in a controlled compatible form, often with dual-read/write only when reconciled and observable.
  3. Backfill in idempotent bounded batches with checkpoints, throttling, validation, and rollback/compensation.
  4. Switch readers/writers and observe mixed versions, queues, replicas, caches, analytics, and external consumers.
  5. Contract/remove old structures only after every consumer is proven upgraded and the rollback window closes.

Avoid rename/drop/type narrowing as one deployment. Database rollback scripts can destroy new writes. Track schema migration ID, transaction boundaries, locking, replica lag, storage, duration, and restore point. Events require versioning, tolerant readers, idempotency keys, ordering and replay strategy; a code rollback does not remove already-published events.

Feature flags and AppConfig

Separate deployment from release with server-side flags. Define owner, purpose, default/fail behavior, target segments, dependencies, exposure metric, expiry/removal date, audit, and test matrix for both states. Flag evaluation must be fast/cached, but stale cache and control-plane outage require a safe explicit default. Never use a client-visible flag as an authorization boundary.

AppConfig can validate syntax/semantics and deploy gradually; CloudWatch alarms can roll back configuration during deployment, including documented missing-data behavior. Reverting configuration does not undo side effects. Test alarm actions, agent/cache behavior, entity consistency where used, and emergency StopDeployment/revert authority. A kill switch must be rehearsed and independently reachable during application failure.

Recovery runbook and evidence

detect and declare -> identify exact bad revision and exposure
 -> stop rollout / reduce traffic / disable feature if safe
 -> classify stateless versus stateful and compatibility
 -> choose rollback, compatibility, compensation, or forward fix
 -> execute one controlled action
 -> verify release + data + user SLI through bake window
 -> reconcile artifacts, flags, schema, queues, audit and cleanup

Evidence includes source/artifact/config/schema IDs, deployment and flag history, traffic weights, alarms, traces/logs, database migration/backfill checkpoints, event offsets/DLQ, external calls, approval timeline, user outcome, and residual work. Declare recovery only after delayed effects and backlog are stable.

Workshop and failure matrix

Given a checkout service release that adds a column, emits event v2, changes tax API, and enables a flag to 10 percent, design expand-contract stages and recovery choices. Calculate canary observations, define alarms and missing data, build compatibility matrix for old/new code/schema/event, and decide actions for code error, slow query, bad backfill, duplicate event, dependency regression, and incorrect tax side effect.

Analyze 20 failures: prior artifact deleted, old capacity absent, down migration loses writes, migration lock, replica lag, old reader rejects field, new writer omits old field, event replay duplicates, cache mixes schema, flag default unsafe, flag service unavailable, alarm disabled, missing data rollback, segment inconsistent, kill switch untested, external payment sent, traffic rollback split writes, roll-forward tests rushed, restore violates RPO, and dashboard green while backlog grows.

Cost and acceptance

Price parallel environments, retained artifacts/capacity, AppConfig requests/deployments, telemetry, migration/backfill compute and I/O, replica/storage growth, queues/replay, external compensation, and incident/customer impact. This lesson creates nothing.

Submit decision tree, compatibility matrix, five-stage schema/event plan, flag contract, AppConfig alarm design, six scenario decisions, 20 failures, recovery evidence, cost, and cleanup of temporary flags/old schema. Pass requires immutable recovery artifacts, reversible traffic/config controls, state-aware decisions, idempotent data operations, and measured user plus data integrity.

Official sources

Advertisement