# P12 AWS FIS safe experiment workbook

Use with AWS236. This is a design and evidence exercise. Do not create a template
or start an experiment from this workbook.

## 1. Authority and scope

- Change/experiment ticket:
- Business and workload owner:
- Experiment commander and independent abort operator:
- Account, Region and environment:
- Why pre-production represents production sufficiently:
- Approved maintenance window and communications channel:
- Explicitly excluded accounts, Regions, resources and customer journeys:
- FIS experiment role owner:
- CloudWatch alarm owner:
- Safety-lever operator and tested access path:
- Recovery owner and rollback decision authority:

## 2. Architecture and failure history

| Journey/component | Dependency | Failure boundary | Existing redundancy/recovery | Known incident or concern | Evidence |
|---|---|---|---|---|---|
|  |  |  |  |  |  |

Do not experiment on an unknown architecture or use chaos testing to investigate
an already unhealthy service. Resolve known critical weaknesses first.

## 3. Steady state and hypothesis

Baseline observation window:

| Business/technical metric | Namespace/dimensions/query | Normal range | Warning threshold | Stop threshold | Missing-data behavior | Owner |
|---|---|---|---|---|---|---|
| successful requests |  |  |  |  |  |  |
| p95/p99 latency |  |  |  |  |  |  |
| error rate |  |  |  |  |  |  |
| queue age/backlog |  |  |  |  |  |  |
| data-integrity signal |  |  |  |  |  |  |

Hypothesis: If **[bounded fault]** affects **[exact target set]** for **[duration]**,
then **[business/technical impact]** will remain within **[threshold]**, recovery
will complete within **[RTO]**, and measured data loss will remain within **[RPO]**.

Disproof criteria:

## 4. Target-resolution proof

| Target name | Resource type | IDs/tags/filters | Selection mode | Candidate count | Maximum affected | Shared/ASG behavior | Exclusion proof |
|---|---|---|---|---:|---:|---|---|
| SandboxInstances | `aws:ec2:instance` | `FisReady=true`, `Environment=sandbox` | `COUNT(1)` |  | 1 |  |  |

- Before-preview inventory time/command:
- Preview experiment ID and `actionsMode=skip-all` evidence:
- Resolved targets independently matched to owned inventory:
- Time gap and possible drift between preview and real execution:
- `emptyTargetResolutionMode` and reason:
- Random selection uncertainty:

A preview does not prove permission to execute the fault, and its randomly sampled
target can differ from a later run. Revalidate immediately before approved start.

## 5. Action graph and recovery semantics

| Action name | Action ID | Parameters/duration | Target | Starts after | Expected side effect | FIS post action | Manual recovery/timeout |
|---|---|---|---|---|---|---|---|
| stopOneInstance | `aws:ec2:stop-instances` | `startInstancesAfterDuration=PT5M` | SandboxInstances | none | one instance stops | restart request | ASG replacement/KMS/start failure runbook |

Draw parallel and sequential actions. An omitted `startAfter` means an action can
start with the experiment. Confirm whether stopping the experiment performs the
action's documented post action; never assume rollback is instantaneous or total.

## 6. Guardrails

### Stop conditions

| Alarm ARN/redacted name | Exact metric/query | Evaluation delay | ALARM behavior | INSUFFICIENT_DATA behavior | Tested before fault? |
|---|---|---:|---|---|---|
|  |  |  |  |  |  |

### Human and account controls

- Safety lever currently disengaged evidence:
- Safety-lever engage drill (safe environment) and UTC time:
- Operator can stop selected experiment evidence:
- Maximum wall-clock timer and independent timer owner:
- IAM policy restricted by action, ARN/tag and account:
- SCP/permission-boundary/change-window controls:
- Backup/restore and capacity verified:
- On-call, security, networking and service-owner acknowledgements:

## 7. Observability and evidence

| Evidence | Destination | Encryption/access | Retention | Expected delivery delay | Owner |
|---|---|---|---|---|---|
| FIS logs schema v2 | CloudWatch Logs/S3 |  |  |  |  |
| CloudTrail | trail/lake |  |  |  |  |
| EventBridge state events | rule/target |  |  |  |  |
| dashboard/report | dashboard/S3 prefix |  |  |  |  |
| application traces/logs |  |  |  |  |  |
| business/data checks |  |  |  |  |  |

Do not use the PDF report to diagnose a failed experiment; retain detailed FIS
logs and application evidence. A completed action proves API completion, not the
resilience hypothesis.

## 8. Go/no-go and execution record

This section may be completed only during a separately authorized real exercise.

- Baseline healthy and stop alarm tested:
- Candidate/resolved target diff accepted:
- No conflicting incident, deploy or maintenance:
- Backups/capacity/recovery operators ready:
- Commander says GO at UTC:
- Experiment ID and template ID/version/export hash:
- Resolved targets:
- Timeline and action states:
- Stop condition/manual stop/safety lever events:
- Actual customer/technical impact:
- Recovery start/end and actual RTO:
- Last consistent point, reconciliation and actual RPO:

## 9. Result and learning

- Hypothesis supported, disproved or inconclusive:
- Evidence supporting that classification:
- Detection/runbook/design gaps:
- Corrective action, owner and due date:
- Residual risk and exception expiry:
- Re-test prerequisites:

## 10. Cleanup and cost

- FIS action-minutes and account count:
- Workload, CloudWatch, S3 and report charges:
- Temporary alarms/dashboards/log destinations/templates:
- Retention-approved evidence retained:
- Exact negative inventory after cleanup:
- Cleanup approver and UTC time:
