AWS 369: EKS deployment, GitOps concepts, and safe rollout evidence
Why this lesson matters
An EKS deployment is safe only when declared intent, controller reconciliation, workload identity, scheduling, probes, rollout settings, policy, observability, and recovery agree. GitOps adds a versioned desired-state workflow, but a repository commit is not proof that the cluster or users reached that state.
Kubernetes and GitOps mental model
Kubernetes is a control loop. A Deployment declares desired replicas and pod template; its controller creates ReplicaSets and Pods until observed state approaches desired state. The scheduler places Pods, kubelet runs containers and probes, Services select ready endpoints, and ingress or load balancing carries traffic.
GitOps stores desired state in Git and lets an in-cluster or platform controller pull and reconcile it. Changes normally flow through review, validation, merge, reconciliation, health assessment, and retained evidence. Manual kubectl edit changes create drift and may be overwritten; emergency change procedure must update or suspend desired state deliberately and reconcile Git afterward.
| Evidence layer | Question | Example proof |
|---|---|---|
| Source | What was approved? | Signed commit, review, manifest diff |
| Artifact | What executes? | ECR digest, scan, SBOM, signature |
| Reconciler | What was applied? | Sync revision, condition, event |
| Kubernetes | What is running? | Deployment/ReplicaSet/Pod UID and digest |
| Network | What receives traffic? | Ready endpoints, ingress/target health |
| User | Does it work safely? | SLI, synthetic transaction, release ID |
Immutable images, identity, and secrets
Pin images by digest rather than a mutable tag. Admission policy can require approved registries, digest pinning, non-root execution, read-only root filesystems, resource requests/limits, and prohibited capabilities. Treat policy denial as a release result, not a reason to bypass policy.
Use a distinct Kubernetes service account for each workload boundary. EKS Pod Identity or IAM roles for service accounts provides AWS credentials to Pods without node-wide credentials. Scope the trust and IAM policy, and verify which SDK credential provider is used. The node role, cluster creator/admin access, GitOps controller role, and workload role are different principals.
Do not commit plaintext secrets. Reference an approved secret system, encrypt repository material where appropriate, control decryption separately, rotate values, and prevent secret output in diffs, events, logs, and support bundles.
Rollout mechanics
For a Deployment with desired replicas 20, maxSurge: 25% permits up to 25 Pods and maxUnavailable: 10% permits up to 2 unavailable during rollout, subject to Kubernetes rounding rules and actual scheduling. Capacity, IPs, topology constraints, quotas, and disruption can still prevent progress.
Readiness removes an unready Pod from Service endpoints. Startup probes protect slow startup before liveness begins. Liveness restarts a stuck container. These serve different purposes. Avoid readiness/liveness checks that depend directly on a shared external dependency because one dependency incident can remove or restart every Pod. Validate critical user dependencies through separate synthetic signals and graceful degradation.
Set progressDeadlineSeconds, resource requests/limits, graceful termination, preStop only when needed, and enough terminationGracePeriodSeconds for draining. A PodDisruptionBudget limits voluntary disruption; it does not guarantee application availability or block every failure. Spread replicas across zones/nodes and ensure the topology plus budget permits maintenance.
kubectl rollout undo changes Deployment history, but GitOps may immediately reapply the bad Git revision. In GitOps, recovery usually means reverting desired state or selecting a previously approved revision, then proving reconciliation and user recovery. Database and event compatibility still require expand-contract and idempotency.
Progressive delivery and observability
Standard Deployment rolling updates control Pod counts, not metric-based canaries. Progressive controllers can manage weighted traffic and analysis, but add custom resources, identities, failure modes, and upgrade responsibility. Adopt them only when the platform team can operate them.
Label telemetry with cluster, namespace, workload, revision, and safe release identity while controlling cardinality. Correlate Kubernetes events, controller conditions, scheduler messages, container status/restarts, application logs, traces, ingress/target health, node pressure, and business SLIs. Events expire, so export required evidence.
Read-only inspection
aws eks describe-cluster --name CLUSTER --region ap-south-1
kubectl config current-context
kubectl auth can-i get deployments -n NAMESPACE
kubectl get deploy,rs,pods,svc,endpointslices -n NAMESPACE -o wide
kubectl rollout status deployment/APP -n NAMESPACE --timeout=5m
kubectl describe deployment APP -n NAMESPACE
kubectl get events -n NAMESPACE --sort-by=.metadata.creationTimestamp
kubectl get pod POD -n NAMESPACE -o jsonpath='{.spec.containers[*].image}'
Confirm account, cluster, context, namespace, RBAC identity, desired/available/updated replicas, conditions, ReplicaSet revision, actual image digest, readiness, endpoints, and events. Do not share tokens, certificates, internal endpoints, secrets, account IDs, or unredacted manifests.
Workshop and failure game day
Review a supplied Deployment and GitOps change. Pin its digest; add service account, security context, resources, probes, rollout limits, topology spread, disruption budget, and release label. Calculate peak Pods and minimum ready Pods. Write policy checks, reconcile the change, verify the exact digest and user outcome, revert Git, and prove stable recovery.
Inject: image pull denial, admission rejection, missing service-account permission, unschedulable resources, subnet IP shortage, failed startup, bad readiness, liveness loop, wrong Service selector, PDB blocking drain, topology rule conflict, node loss, Git drift, reconciler outage, and rollback with incompatible schema. For each retain source revision, sync state, controller/Kubernetes event, Pod reason, endpoint state, SLI, owner, and action.
Cost and acceptance
Price EKS control plane, worker/Fargate capacity including surge, EBS/EFS, load balancers, NAT and cross-AZ traffic, ECR, logs/metrics/traces, scanners, GitOps/progressive controllers, and retained environments. This no-create lesson requires no cleanup beyond local redacted files.
Submit the six-layer evidence chain, rollout math, identity map, policy set, probe rationale, scheduling/disruption design, Git rollback procedure, fifteen failure results, cost model, and evidence-retention plan. Pass requires digest-to-Pod proof, least privilege, meaningful readiness, sufficient capacity, Git-consistent recovery, and a measured user outcome.