AWS 326: Improving existing solutions
Why this lesson matters
Existing workloads carry users, data, revenue, contracts, undocumented behavior, operational habits, technical debt, and incident history. Replacing them from frustration can create greater risk than the original problem. Improving them requires a measured current state, causal diagnosis, explicit target outcomes, and dependency-safe change.
Continuous improvement is a core SAP-C02 capability. The architect must distinguish symptoms from causes, urgent containment from durable correction, incremental modernization from replacement, and successful deployment from proven user improvement.
Learning outcomes
By the end, you can:
- baseline user outcomes, architecture, operations, cost, risk, and debt;
- reconstruct undocumented request, data, identity, and failure paths;
- diagnose causes using correlated evidence rather than one metric;
- choose retain, remediate, replatform, refactor, replace, or retire boundaries;
- prioritize improvements by outcome, risk, dependency, and reversibility;
- design coexistence, migration, rollback, and decommissioning;
- validate improvement with before/after evidence and guardrail metrics;
- build a practical 30/60/90-day roadmap.
1. Establish authority and safety
Before discovery, identify workload owner, data owner, technical owner, on-call team, business approver, security contact, and change authority. Define read-only evidence scope, maintenance restrictions, sensitive-data handling, and emergency escalation.
Do not “clean up” unknown resources, enable expensive services, alter logging, or run load/failure tests against production without approval. Preserve evidence timestamps and redact account IDs, customer data, tokens, and internal addresses.
2. Define improvement outcomes
Replace vague goals with baselines and targets:
| Vague request | Testable improvement |
|---|---|
| Make it reliable | Raise checkout success from 97.8% to 99.9% monthly while meeting one-hour RTO |
| Make it faster | Reduce p95 customer latency from 2.4 s to below 700 ms at 3,000 requests/s |
| Reduce cost | Lower cost per completed order by 25% without worsening error, latency, or recovery |
| Improve security | Remove public database reachability and prove approved application-only access |
| Modernize | Cut lead time from ten days to one day with unchanged control evidence |
Define guardrail metrics so one improvement does not damage another quality. A cost change should preserve error rate, latency, recovery, and security. A performance change should not silently increase inconsistency or spend.
3. Reconstruct the current state
Build an evidence-backed model from several sources:
- stakeholder interviews and operator observation;
- accounts, Regions, resources, tags, ownership, and quotas;
- DNS, edge, network, identity, certificate, and key configuration;
- application components, versions, deployments, and dependencies;
- schemas, access patterns, data flows, retention, backups, and restores;
- metrics, logs, traces, synthetic checks, and business events;
- incidents, post-incident actions, support cases, and change history;
- bills, commitments, utilization, transfer, and license costs;
- vulnerabilities, findings, exceptions, audits, and technical debt.
Mark every diagram element as observed, inferred, or unknown. Compare declared configuration with runtime traffic. An unused-looking security group may support a rare recovery path; a diagrammed dependency may no longer receive traffic.
Draw normal and failure paths. Follow one customer request through DNS, edge, identity, load balancing, application, cache, database, events, downstream providers, logging, and response. Then trace dependency timeout, retry, AZ loss, deployment, restore, and operator intervention.
4. Safe read-only evidence
Examples depend on permission and enabled services:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws resourcegroupstaggingapi get-resources --resources-per-page 50 --output json
aws cloudwatch describe-alarms --state-value ALARM --output table
aws configservice get-compliance-summary-by-resource-type --output table
aws securityhub get-findings --max-results 20 --output json
aws compute-optimizer get-enrollment-status --output json
aws backup list-backup-jobs --by-state FAILED --max-results 20 --output table
The tagging API is not a complete inventory. An empty Security Hub, Config, Backup, or Compute Optimizer response may mean the service is disabled, scoped differently, or inaccessible. Record coverage before interpreting results.
For Cost Explorer, use approved dates rather than copying stale lesson dates:
aws ce get-cost-and-usage \
--time-period Start=YYYY-MM-01,End=YYYY-MM-01 \
--granularity MONTHLY \
--metrics UnblendedCost UsageQuantity \
--region us-east-1 \
--output json
Do not aggregate unlike UsageQuantity units. Choose AmortizedCost or NetAmortizedCost when commitment/discount analysis requires it and finance agrees on the metric.
5. Baseline by user journey
Infrastructure averages can hide customer pain. For each critical journey define success, latency percentiles, correctness, freshness, durability, availability period, volume, and business value. Correlate with resource saturation, queues, errors, retries, throttles, dependency latency, deployment, and changes.
Use representative windows: normal, peak, seasonal, incident, and quiet. Check missing data and aggregation. Average latency can improve while p99 worsens. A falling error count can reflect falling traffic. CPU can be low while memory, lock, connection, storage, or dependency limits dominate.
6. Diagnose causes, not symptoms
Use a timeline and hypothesis loop:
- State the symptom and affected users.
- Bound start/end, accounts, Regions, versions, and journeys.
- Correlate changes, demand, dependencies, quotas, and failures.
- Form competing hypotheses.
- State evidence that would confirm or reject each.
- Run the safest discriminating check.
- Contain customer impact where necessary.
- Correct the causal mechanism and test recurrence prevention.
Example: high API latency may arise from slow database queries, exhausted connections, retry storms, cache misses, downstream timeout, packet loss, GC, deployment regression, or throttling. Scaling instances treats only some causes and may amplify database pressure.
Separate root cause, contributing conditions, detection gaps, and recovery gaps. Avoid single-person blame; improve system controls and learning.
7. Assess six-pillar and organizational debt
Review Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability together. Add delivery and organizational constraints such as owner gaps, unsupported runtime, fragile manual changes, missing skills, vendor lock-in, and expiring contracts.
Classify debt:
- active risk causing incidents or exposure;
- delivery friction slowing safe change;
- cost waste with measurable opportunity;
- lifecycle risk from end-of-support technology;
- architecture constraint blocking requirements;
- accepted debt with owner, reason, and review date.
Debt is not automatically a rewrite mandate. Quantify business impact and change risk.
8. Choose the improvement strategy
Apply the smallest boundary that resolves the cause while moving toward a coherent target:
| Strategy | Suitable when | Main risk |
|---|---|---|
| Retain | Current design meets outcomes and change has low value | Unreviewed debt becomes permanent |
| Remediate | Configuration, capacity, code, query, or operational defect is bounded | Local fix hides systemic issue |
| Replatform | Managed platform removes material operations with limited code change | Compatibility and hidden service behavior |
| Refactor | Domain, scale, resilience, or delivery need requires code/data redesign | Distributed complexity and long coexistence |
| Replace/repurchase | Product satisfies need better than custom ownership | Data migration, integration, contract, exit |
| Retire | Capability and dependencies are no longer needed | Hidden users, legal retention, or restore need |
Different components can use different strategies. Validate dependencies before retirement.
9. Prioritize an improvement portfolio
Score business impact, security/reliability risk reduction, urgency, confidence, effort, dependency, reversibility, cost, and learning value. Mandatory risk remediation should not lose to cosmetic high-ROI work.
Sequence foundations before dependent changes: observability before blind optimization, tested backup before risky migration, identity/network baseline before broad deployment, and compatibility before database cutover. Maintain quick wins, foundational work, experiments, and strategic stages in one roadmap.
Define outcome hypothesis, owner, prerequisites, change, validation, guardrails, rollback, evidence, and review date for every item.
10. Improve through controlled stages
Use strangler routing, branch by abstraction, expand/migrate/contract schema change, shadow traffic, dual read, controlled dual write, canary, blue/green, or queue-based decoupling where appropriate.
Dual write is not automatically safe; partial success causes divergence. Define source of truth, idempotency, reconciliation, repair, and cutover. Shadow traffic must protect sensitive data and prevent side effects. Rollback after target writes may require reverse replication, reconciliation, or forward repair.
Each stage needs entry criteria, change window, owner, expected signals, abort threshold, rollback/repair, observation period, and exit evidence. Limit simultaneous variables so results remain diagnosable.
11. Modernize operations with the workload
An improved service still fails if its operating model remains manual and ambiguous. Include infrastructure as code, deployment safety, patch/lifecycle ownership, SLOs, actionable alarms, distributed tracing, runbooks, capacity reviews, security response, restore exercises, cost anomaly handling, and game days.
Remove obsolete monitors, roles, DNS, certificates, queues, replication, firewall rules, backups, and licenses only after the new path is accepted and retention obligations are met.
12. Verify improvement
Compare before and after using the same journey definition, load, percentile, time window, cost metric, and boundary. Include change-induced incidents and new operational work. Validate user outcome, not merely deployment completion.
Use leading indicators such as queue age, saturation, deployment failure, restore-test success, patch age, and alarm coverage, alongside lagging outcomes such as availability, revenue loss, incidents, and cost per order.
A result is inconclusive when traffic, season, measurement, or functionality differs materially. Record uncertainty instead of claiming success.
13. Guided workshop: unreliable legacy workload
A three-tier order application runs in one AWS account. It uses internet-facing EC2 instances changed manually, a single database writer, local sessions, synchronous payment/inventory/email calls, broad IAM roles, public administration, untested snapshots, and noisy CPU alarms. Promotions cause timeouts and duplicate retries. Monthly cost rose 35%, but cost allocation and data-transfer ownership are unclear. The business cannot tolerate a big-bang rewrite.
Produce:
- owner, authority, and evidence charter;
- measurable outcome and guardrail metrics;
- current account/resource/ownership inventory;
- request, identity, network, and data flows;
- dependency and integration register;
- user-journey baseline;
- incident/change timeline;
- competing-cause hypothesis table;
- security and six-pillar risk review;
- data, backup, restore, and retention assessment;
- cost and utilization baseline with metric definitions;
- retain/remediate/replatform/refactor/replace/retire analysis by component;
- target-state architecture and ADRs;
- 30/60/90-day roadmap;
- coexistence, data migration, and reconciliation design;
- staged deployment with abort and rollback/repair criteria;
- before/after validation dashboard specification;
- decommission checklist and residual-debt register.
14. Common failures
- rewriting before measuring the present system;
- changing compute, database, network, and code simultaneously;
- optimizing averages instead of affected journeys and percentiles;
- treating every finding as equal severity;
- assuming a managed service fixes application behavior;
- migrating data without reconciliation and target-write rollback;
- declaring success at deployment rather than observation;
- deleting the old path before business acceptance;
- retaining duplicate systems indefinitely without an exit owner.
Cost and cleanup
This lesson is T0 and creates no resources. Approved future tests require cost caps, sanitization, tags, expiration, manifests, and verified cleanup. Improvement economics must include coexistence and migration, not only target-state run cost.
Knowledge check
- Why is one low CPU graph insufficient? It omits other constraints, journeys, percentiles, and time ranges.
- When should a component be retained? When it meets outcomes and changing it has insufficient value or excessive risk.
- Why is dual write risky? Partial success creates divergence without idempotency and reconciliation.
- What proves improvement? Comparable before/after user and guardrail evidence over a representative period.
- When may the old path be removed? After accepted validation, retention/dependency checks, rollback decisions, and decommission approval.
Lesson acceptance
Submit all 18 artifacts. Claims must use dated evidence, diagnosis must test competing causes, strategy must be chosen by component, stages must preserve service and data safety, and success must be proven against original user outcomes and guardrails before decommissioning.