AWS 350: Bash automation, strict mode, exit codes, logging, and idempotency
Why this lesson matters
Shell scripts often connect Git, AWS CLI, package tools, deployment systems, and incident runbooks. A five-line script can change hundreds of resources. Reliable automation therefore needs explicit inputs, safe quoting, controlled failure, useful logs, bounded retries, idempotent state checks, dry-run behavior, tests, and rollback.
You already know Linux commands. This lesson focuses on turning commands into trustworthy automation without pretending that set -e makes a script safe.
Learning outcomes
You will be able to:
- select Bash deliberately with a shebang and check required tools;
- explain
errexit,nounset,pipefail, andERRtrap limitations; - parse and validate arguments without executing input as code;
- quote expansions and preserve array element boundaries;
- use functions, local variables,
readonly, and explicit return codes; - emit structured, redacted logs to the correct stream;
- implement bounded retry for genuinely transient and idempotent operations;
- reconcile desired and observed state so repeated runs converge;
- support dry-run, atomic replacement, cleanup traps, tests, and rollback.
Safety model
This lab writes only under a new temporary directory. Do not substitute /etc, a production repository, or an AWS resource. The final AWS pattern remains pseudocode until it has authorization, least privilege, tests, and cleanup.
Create the lab:
lab_dir=$(mktemp -d -t aws350.XXXXXX)
printf 'lab_dir=%s\n' "$lab_dir"
cd "$lab_dir"
mkdir -p bin state tests
Keep that terminal open. Verify pwd before every cleanup.
Strict mode is a tool, not a proof
A common baseline is:
#!/usr/bin/env bash
set -Eeuo pipefail
-erequests exit after an unhandled non-zero status, but Bash has context-dependent exceptions in tests,&&/||lists, pipelines, functions, and substitutions.-Eallows anERRtrap to be inherited by functions, command substitutions, and subshell environments in more cases.-utreats expansion of an unset parameter as an error, so optional values need forms such as${value:-}or${value:?message}.pipefailmakes a pipeline fail if any command fails, using the rightmost non-zero status.
Strict mode cannot validate business logic, make retries safe, prevent word splitting in unquoted expansions, or roll back a partial change. Handle expected failures explicitly:
if result=$(some_command 2>&1); then
printf 'success: %s\n' "$result"
else
status=$?
printf 'command failed status=%d message=%s\n' "$status" "$result" >&2
fi
Do not write some_command || true unless failure is deliberately accepted, bounded, logged, and tested.
Quoting, arrays, and input boundaries
Use "$value", "${array[@]}", and -- before untrusted positional arguments when the command supports it. Never build a command in a string and run it with eval.
files=("release one.env" "release-two.env")
for file in "${files[@]}"; do
printf 'file=%q\n' "$file"
done
Validate identifiers with allowlists:
validate_service() {
local value=${1:?service required}
[[ $value =~ ^[a-z][a-z0-9-]{1,30}$ ]]
}
The regular expression limits syntax; it does not prove authorization or existence.
Build an idempotent local reconciler
Create bin/reconcile-service with this complete implementation:
#!/usr/bin/env bash
set -Eeuo pipefail
readonly EX_USAGE=64
readonly EX_DATAERR=65
readonly EX_SOFTWARE=70
readonly EX_TEMPFAIL=75
service=""
desired_version=""
state_root=""
dry_run=0
lock_dir=""
temp_file=""
usage() {
printf 'Usage: %s --service NAME --version VERSION --state-root DIR [--dry-run]\n' "${0##*/}"
}
log() {
local level=${1:?level required}
local event=${2:?event required}
shift 2
printf 'ts=%(%Y-%m-%dT%H:%M:%SZ)T level=%q event=%q' -1 "$level" "$event" >&2
printf ' %q' "$@" >&2
printf '\n' >&2
}
die() {
local status=${1:?status required}
shift
log ERROR failure "$@"
exit "$status"
}
cleanup() {
local status=$?
[[ -z ${temp_file:-} || ! -e $temp_file ]] || rm -f -- "$temp_file"
[[ -z ${lock_dir:-} || ! -d $lock_dir ]] || rmdir -- "$lock_dir" 2>/dev/null || true
log INFO exit "status=$status"
exit "$status"
}
on_error() {
local status=$?
log ERROR command_failed "status=$status" "line=${BASH_LINENO[0]:-unknown}" "command=$BASH_COMMAND"
return "$status"
}
trap on_error ERR
trap cleanup EXIT
trap 'exit 130' INT
trap 'exit 143' TERM
while (($#)); do
case $1 in
--service) (($# >= 2)) || die "$EX_USAGE" "missing=service"; service=$2; shift 2 ;;
--version) (($# >= 2)) || die "$EX_USAGE" "missing=version"; desired_version=$2; shift 2 ;;
--state-root) (($# >= 2)) || die "$EX_USAGE" "missing=state_root"; state_root=$2; shift 2 ;;
--dry-run) dry_run=1; shift ;;
--help) usage; exit 0 ;;
*) usage >&2; die "$EX_USAGE" "unknown_argument=$1" ;;
esac
done
[[ $service =~ ^[a-z][a-z0-9-]{1,30}$ ]] || die "$EX_DATAERR" "invalid_service=$service"
[[ $desired_version =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]] || die "$EX_DATAERR" "invalid_version=$desired_version"
[[ -n $state_root && -d $state_root && ! -L $state_root ]] || die "$EX_DATAERR" "invalid_state_root=$state_root"
readonly target="$state_root/$service.env"
lock_dir="$state_root/.$service.lock"
mkdir -- "$lock_dir" 2>/dev/null || die "$EX_TEMPFAIL" "lock_busy=$lock_dir"
desired=$(printf 'service=%s\nversion=%s\n' "$service" "$desired_version")
current=""
[[ ! -f $target ]] || current=$(<"$target")
if [[ $current == "$desired" ]]; then
log INFO no_change "target=$target" "version=$desired_version"
exit 0
fi
if ((dry_run)); then
log INFO change_planned "target=$target" "version=$desired_version"
exit 0
fi
temp_file=$(mktemp "$state_root/.$service.XXXXXX") || die "$EX_SOFTWARE" "mktemp_failed"
printf '%s\n' "$desired" > "$temp_file"
chmod 0640 "$temp_file"
mv -- "$temp_file" "$target"
temp_file=""
log INFO changed "target=$target" "version=$desired_version"
Install and inspect it locally:
chmod 0755 bin/reconcile-service
bash -n bin/reconcile-service
sed -n '1,240p' bin/reconcile-service
The script compares desired and observed content, then atomically renames a completed temporary file on the same filesystem. The lock directory rejects concurrent writers. EXIT cleanup preserves the original exit status and removes owned temporary state.
Test positive, repeated, dry-run, and invalid paths
bin/reconcile-service --service orders --version 1.2.3 --state-root "$PWD/state"
first_hash=$(sha256sum state/orders.env)
bin/reconcile-service --service orders --version 1.2.3 --state-root "$PWD/state"
second_hash=$(sha256sum state/orders.env)
test "$first_hash" = "$second_hash"
bin/reconcile-service --service orders --version 2.0.0 --state-root "$PWD/state" --dry-run
grep -Fx 'version=1.2.3' state/orders.env
if bin/reconcile-service --service '../bad' --version 1.2.3 --state-root "$PWD/state"; then
printf 'unexpected success\n' >&2
exit 1
else
printf 'invalid input correctly rejected status=%d\n' "$?"
fi
Idempotency means repeated execution with the same desired state has no additional effect. It does not mean an operation is safe for concurrent callers, all failures are reversible, or remote APIs provide exactly-once execution.
Understand exit status and pipelines
Run this controlled comparison:
set +o pipefail
false | true
printf 'without_pipefail=%d\n' "$?"
set -o pipefail
false | true
printf 'with_pipefail=%d pipestatus=%s\n' "$?" "${PIPESTATUS[*]}"
Capture PIPESTATUS immediately if you need component statuses because a later command replaces it. Use command conditions for expected outcomes rather than depending on subtle errexit behavior.
Exit codes are part of the interface. This script uses sysexits-style values for usage/data/software/temporary failure, but an organization may define another documented scheme. Never return zero after a failed required operation.
Logging without leaking secrets
Logs should answer when, level, event, operation/target, request or correlation ID, attempt, duration, and outcome. Send diagnostics to stderr so stdout can remain machine-readable. Redact credentials, tokens, cookies, private payloads, presigned URLs, and sensitive command expansions.
Avoid set -x in secret-bearing automation. If trace is necessary in a safe lab, set a controlled PS4, direct BASH_XTRACEFD to a protected file, and disable tracing before reading secrets. Quoted %q output improves shell-readable escaping but is not JSON and does not sanitize secrets.
Bounded retry
Retry only operations classified as transient and safe to repeat. Do not retry authentication denial, validation errors, destructive mutations, or unknown partial outcomes blindly.
retry() {
local max=${1:?max required}
local delay=${2:?delay required}
shift 2
local attempt status
for ((attempt=1; attempt<=max; attempt++)); do
if "$@"; then
return 0
else
status=$?
fi
((attempt == max)) && return "$status"
printf 'retry attempt=%d status=%d delay=%d\n' "$attempt" "$status" "$delay" >&2
sleep "$delay"
delay=$((delay * 2))
done
}
Production retry should also apply maximum delay, total deadline, jitter, error classification, observability, and service quota/rate guidance. Preserve idempotency tokens when the API supports them.
AWS CLI automation pattern
Separate discovery, plan, apply, verify, and cleanup. Resolve identity and Region explicitly. Prefer structured --query output over parsing human tables. Example read-only preflight:
aws --version
aws sts get-caller-identity --output json
aws configure get region
aws ec2 describe-regions --query 'Regions[].RegionName' --output text
For mutations, require an approved account/Region/resource marker, use unique ownership tags, record created IDs immediately, wait with a deadline, verify postconditions independently, and clean only IDs created by this run. Do not discover cleanup targets from broad name matching.
Static and dynamic validation
Run syntax checking on every change. If available, use ShellCheck with reviewed exclusions rather than ignoring all findings:
bash -n bin/reconcile-service
command -v shellcheck >/dev/null && shellcheck bin/reconcile-service
Add automated cases for help, missing argument, invalid service, invalid version, invalid/symlink state root, first create, no-change repeat, version update, dry-run, lock contention, unwritable path, interrupted temporary file, filenames with spaces, and cleanup. Test under the Bash versions your runners actually use.
Failure diagnosis
| Symptom | Likely cause | Evidence and correction |
|---|---|---|
| Pipeline looks successful | pipefail absent or status overwritten | Minimal reproduction; inspect component statuses |
| Empty optional variable exits | nounset expansion | Define default or require value explicitly |
| Script continues unexpectedly | errexit exception/context | Handle status in if; add focused test |
| Repeated run changes again | Comparison is volatile or incomplete | Normalize desired state; compare stable fields |
| Partial file appears | Direct overwrite/interruption | Write same-filesystem temp, validate, atomic rename |
| Two runs corrupt state | No concurrency control | Lock, conditional API, or external coordinator |
| Retry duplicates work | Operation not idempotent | Idempotency key/state record/reconciliation |
| Logs expose token | tracing/raw command or response | Rotate secret, redact, restrict evidence, redesign logging |
Rollback and independent challenge
Before changing orders to 2.0.0, copy its exact prior content to a versioned file inside the lab and record its hash. Apply, verify, then restore through the reconciler using the prior version. Rollback is another forward operation with its own validation, not an assumption that all side effects reverse.
Independently extend the script with --log-format text|json, a maximum lock age policy that never deletes an active lock blindly, and an approved-state backup. Add tests for every branch. Explain why parsing with eval, unquoted $@, broad globs, predictable /tmp names, and curl | bash are unsafe.
Acceptance and cleanup
Submit the script, syntax/static results, test matrix, first-run and no-change logs, hash equality, dry-run proof, invalid-input status, lock contention result, atomic-write explanation, retry classification, rollback proof, and threat review.
Pass requires all tests, no secret output, no unquoted argument bug, bounded scope, meaningful non-zero failures, unchanged state on dry-run, convergent second run, cleanup preserving exit status, and rollback verified from content and hash.
After review, remove only the displayed lab directory using your normal safe deletion process. Never copy a broad recursive removal command from training material.