Lesson 311 · AWS Learning Path

AWS 311: Responsible generative AI controls and human approval

· Published · 11 min read

Labelled process diagram for AWS 311: Approved purpose and governed inputs to Model, retrieval, guardrail, and tool controls to Human-approved or bounded output to Evaluation, monitoring, appeal, and shutdown...

Why this lesson matters

A generative model predicts useful-looking content from patterns. It can summarize, classify, draft, retrieve, call tools, and converse, but it can also fabricate facts, follow malicious instructions, reproduce sensitive data, produce harmful content, or take an authorized action for the wrong reason. Fluency is not evidence of truth.

Responsible AI is therefore a complete system property. It includes the business purpose, data rights, model and prompt, retrieval sources, tools, identity, interface, human authority, testing, monitoring, incident response, and retirement. A content filter alone cannot prove fairness, factuality, privacy, or safe action.

Human approval is also not automatically safe. A reviewer can be overloaded, underqualified, misled by confident prose, or given only an Approve button. The system must show evidence, make rejection easy, preserve accountability, and fail closed when review is unavailable for a high-impact action.

What you will be able to do

By the end, you can:

  • explain hallucination, grounding, prompt injection, indirect injection, jailbreak, data poisoning, memorization, toxicity, bias, and overreliance;
  • classify generative-AI use cases by impact, autonomy, reversibility, and data sensitivity;
  • design controls before input, during inference, after output, and before action;
  • distinguish automated validation, human review, human approval, and human execution;
  • build evidence-based evaluation datasets and red-team tests;
  • define safe retrieval and tool-use boundaries;
  • design reviewer context, authority, separation of duties, service levels, and fallback;
  • monitor quality, abuse, drift, overrides, incidents, and cost;
  • respond to leaked data, poisoned sources, unsafe output, and unauthorized actions; and
  • produce a governance dossier for a customer-support assistant.

Before you start

  • Use synthetic prompts and documents. Do not paste customer, employee, medical, legal, financial, security, source-code, credential, or regulated data into an unapproved model.
  • Do not let a lab system send messages, modify records, approve money, change access, or operate physical equipment.
  • Model/provider terms, Region, data use, retention, training, copyright, and subcontractor behavior require current legal and security review.
  • Redact system prompts, guardrail configurations, test attacks, internal sources, model IDs, account IDs, user data, and harmful generated content from shared evidence.
  • Default to a T0 local threat model and evaluation plan. Any live evaluation needs owner approval and a strict cost/token cap.

1. Start with purpose, not a chatbot

Write an AI system card:

FieldRequired answer
Intended user and taskWho uses it and what exact outcome is supported?
Prohibited useWhich decisions, data, actions, and populations are excluded?
ImpactWhat happens when it is wrong, delayed, unavailable, or abused?
AutonomyDraft, recommend, request approval, or act?
ReversibilityCan every effect be cancelled or compensated?
EvidenceWhich source and test prove an answer or action?
Human authorityWho can approve, reject, edit, escalate, and override?
LifecycleWho owns versions, monitoring, incidents, and retirement?

Reject generative AI when templates, search, deterministic rules, or a purpose-built classifier solve the task with lower uncertainty. A model is not justified merely because users prefer a conversational interface.

2. Understand the failure modes

Hallucination is plausible but unsupported or false output. Retrieval reduces some unsupported answers only if the source is correct, current, authorized, successfully retrieved, and actually used. A citation can point to a real source that does not support the claim.

Prompt injection places instructions in user input to override intended behavior. Indirect prompt injection hides instructions in retrieved pages, documents, emails, images, or tool results. A system prompt is not a security boundary because adversarial content can influence the same model context.

Jailbreaking tries to evade behavior restrictions. Data poisoning corrupts training, customization, retrieval, memory, or feedback data. Sensitive-data disclosure can arise from user input, retrieved sources, logs, model output, tool results, or memorized content.

Other risks include harmful/toxic output, stereotypes and unequal error rates, copyright or license conflict, insecure generated code, automation bias, sycophancy, denial-of-wallet token use, tool loops, unavailable models, and silent behavior changes after model or policy updates.

3. Classify impact and autonomy

Use four control classes:

  1. Low impact, no action: brainstorming over public data. Automated filters plus user notice and sampling may suffice.
  2. Moderate impact, reversible: internal draft or ticket suggestion. Require source access, quality checks, review for selected cases, and rollback.
  3. High impact, human decision: legal, financial, employment, identity, health, safety, or access support. Qualified human decision, independent evidence, appeal, and strict data controls are mandatory.
  4. Action-capable agent: the system can call tools or alter external state. Apply least privilege, allowlisted tools/arguments, approval before material effects, idempotency, budgets, and kill switches.

Never lower a classification because the model “only recommends.” Users can rely on recommendations as decisions. Assess the real downstream effect.

4. Build defense in depth

approved user and purpose
  -> input validation/minimization
  -> model + versioned system instruction
  -> authorized retrieval/tools
  -> output validation/grounding
  -> risk-based human gate
  -> bounded action
  -> audit, feedback, monitoring, incident response

Before inference

Authenticate the user; authorize the use case and source; minimize/redact data; limit size, files, and media; scan uploads; classify prompt risk; rate-limit by identity; set token/time/cost limits; and reject prohibited tasks.

During inference

Pin the model/version where possible; version prompts and parameters; delimit data from instructions; retrieve only documents the requesting identity may access; limit context; apply content/sensitive-information policies; and keep temperature/decoding suitable for the task.

After inference

Validate schema, citations, calculations, links, secrets, PII, harmful content, and business rules. Recompute deterministic fields outside the model. Require “insufficient evidence” rather than forcing an answer. Do not execute prose as code, SQL, shell, or a URL.

Before action

Map a typed request to an allowlisted operation. Re-authorize against current user identity and resource. Validate every argument. Show a human the proposed effect, target, evidence, uncertainty, and rollback. Use one-time approval, idempotency keys, transaction limits, and post-action verification.

5. Protect retrieval and memory

Retrieval-augmented generation has two authorization points: indexing and retrieval. During ingestion, verify source owner, classification, license, malware, document version, and retention. Preserve provenance, chunk-to-document mapping, and deletion. Do not let anyone upload “knowledge” into a production corpus without moderation and approval.

At query time, filter by the requesting identity and current entitlements before content enters model context. Document-level access is insufficient if one document mixes public and restricted sections. Test revoked users, changed groups, deleted documents, tenant boundaries, and cached answers.

Treat retrieved text as untrusted data. It may contain direct or hidden instructions. Keep tool credentials outside the context and require independent policy checks for actions.

Conversation history and long-term memory create a new datastore. Define who can write/read it, what is remembered, expiry, correction, deletion, tenant isolation, and whether the user can inspect it. Never remember secrets merely because they appeared in a conversation.

6. Design safe tool use and agents

Give each tool a narrow typed contract. Prefer get_order_status(order_id) over unrestricted database access. Separate read tools from write tools. Use a dedicated execution role whose permissions match the tool, not the model service.

For each action define:

  • authorized caller and target;
  • accepted arguments and validation;
  • maximum amount, recipients, records, or duration;
  • required fresh data and preconditions;
  • human approval threshold;
  • idempotency and duplicate detection;
  • success verification;
  • compensation or rollback;
  • immutable audit fields; and
  • timeout, retry, loop, and total-step limits.

Prompt injection must not expand tool permission. Even when the model proposes a valid function call, deterministic authorization decides whether it runs.

7. Make human approval meaningful

Human-in-the-loop can mean four different controls:

  • review: inspect quality, possibly after output;
  • approval: explicitly authorize a proposed effect;
  • execution: human performs the effect in another system;
  • override: human changes an automated result under recorded authority.

For a high-impact approval, show original request, affected person/resource, generated proposal, authoritative sources, unsupported claims, policy checks, model/prompt/knowledge versions, uncertainty or risk indicators, previous decisions, and the exact external effect. Do not show only polished prose.

The reviewer needs training, permission, enough time, an independent evidence path, and choices to reject, edit, request information, or escalate. Measure queue age, approval/rejection rate, edit distance, disagreement, escalation, sampled reviewer accuracy, and downstream incidents. High automatic approval can indicate rubber-stamping.

Fail closed when required approval is unavailable. Do not silently convert an approval workflow to automatic execution to meet latency. Define emergency/manual procedures outside the model.

8. Evaluate before launch

Create a versioned evaluation set with normal, edge, adversarial, multilingual, accessibility, subgroup, stale-source, conflicting-source, insufficient-evidence, and malformed-input cases. Keep a hidden holdout set to prevent tuning to the test.

Evaluate dimensions separately:

DimensionExample measure
Task successRubric score or exact structured-field correctness
GroundingClaims supported by cited authorized sources
SafetyHarmful responses blocked or safely redirected
PrivacySecret/PII extraction and tenant-leak tests
Injection resistanceDirect/indirect attacks that cause policy violation
FairnessQuality/error differences across relevant groups/languages
Tool safetyUnauthorized/invalid/duplicate action attempts prevented
OperationsLatency, availability, token use, and cost per completed task
Human controlCorrect escalation, reviewer comprehension, reversal success

Automated model-as-judge scoring is scalable but can be biased, inconsistent, or favor related models. Calibrate it against qualified human labels and never use it as the only launch proof. Record confidence intervals and failure examples, not only averages.

Red-team system prompts, retrieval, files, URLs, tools, memory, and the user interface. Include prompt extraction, encoded attacks, role-play, multilingual injection, malicious document text, poisoned knowledge, data exfiltration, tool-argument manipulation, recursive calls, and denial-of-wallet.

9. Monitor and respond

Log request ID, authenticated actor, use case, model/prompt/guardrail/knowledge/tool versions, policy decisions, retrieved-source IDs, approval identity/time, tool call, result, latency, tokens, and cost. Minimize or tokenize content and set retention/access controls; logging everything can create a larger breach.

Monitor policy blocks, sensitive-data detections, unsupported-answer rate, citation failures, human edits/rejections, escalation, tool denial, duplicate prevention, harmful feedback, latency, model errors, token spikes, and cost. Sample accepted answers for quality because users report only some errors.

Incident runbooks must cover sensitive-data disclosure, unsafe output, poisoned source, compromised tool, unauthorized action, abusive user, model regression, and provider outage. Be able to disable a model, prompt, corpus, tool, tenant, or whole feature independently. Preserve evidence, notify owners, correct downstream effects, rotate credentials, delete poisoned content, and retest before re-enable.

10. Practical governance workshop

Scenario: an assistant drafts customer refund responses, retrieves policy/order data, and may propose refunds up to a limit. It must never issue a refund by itself.

Submit:

  1. system card and rejected simpler alternatives;
  2. risk classification and prohibited-use list;
  3. data-flow/trust-boundary diagram;
  4. retrieval ingestion and query authorization;
  5. prompt/model/knowledge versioning;
  6. typed read and refund-proposal tools;
  7. deterministic policy and argument checks;
  8. reviewer screen fields and approval authority;
  9. idempotent human-executed refund path;
  10. 40-case evaluation set with pass thresholds;
  11. 15 direct/indirect injection tests;
  12. privacy, fairness, grounding, and tool-safety results;
  13. logging/retention/monitoring design;
  14. incident kill-switch and rollback drill;
  15. token/retrieval/reviewer/cost model; and
  16. architecture decision record.

Inject a malicious policy document, revoked employee access, false citation, leaked API key in input, duplicate approval click, overloaded review queue, model-version regression, and tool timeout after an ambiguous result. Show evidence, containment, recovery, and prevention.

Diagnose from evidence

SymptomEvidenceSafe response
Fluent answer is falseclaims, citations, retrieved chunks, source versionMark unsupported, correct source/retrieval/evaluation, do not merely rewrite prompt.
User sees another tenant's factactor, retrieval filter, cache/memory, source ACLDisable affected path, investigate breach, repair authorization and retest revocation.
Malicious document changes behaviorchunk content, instruction trace, tool proposalQuarantine source, block action, strengthen ingestion and deterministic tool policy.
Guard/filter blocks valid usepolicy/version, category, context, false-positive setRoute safe fallback/review and tune through governed change.
Reviewers approve everythingqueue time, UI evidence, edit/reject/escalation metricsStop high-impact automation, improve authority, staffing, training, and interface.
Duplicate external actionrequest/approval/action IDs, timeout/retry historyReconcile target state and enforce idempotency before retry.
Token cost spikesinput/history/retrieval size, loops, abusive identityStop loop, cap context/steps, rate-limit and investigate.

Cost and cleanup

Price model input/output tokens, embeddings, retrieval/vector storage, reranking, guardrails, agents/tools, Lambda/Step Functions, logs, evaluations, human review, red-team work, monitoring, and incident handling. Cost per token is not cost per correct business outcome.

For T0, retain only redacted local design evidence. For an approved pilot, disable tools first; revoke temporary roles/credentials; delete synthetic prompts, knowledge sources, indexes, memories, logs, evaluations, and test identities under retention approval; remove endpoints/functions; and verify no scheduled ingestion or billable store remains.

Knowledge check

  1. Why is grounding not proof? Sources can be wrong, stale, unauthorized, or misused.
  2. Why is a system prompt not authorization? Model instructions are probabilistic and can be influenced; policy must be deterministic.
  3. What is indirect injection? Malicious instructions arrive through retrieved or tool-provided data.
  4. When must approval fail closed? When a required qualified reviewer is unavailable for a high-impact effect.
  5. Why can human review fail? Automation bias, poor context, overload, weak authority, or rubber-stamping.
  6. Why re-authorize tool calls? User and resource permissions can differ from model/runtime access and can change.
  7. Why not log every prompt forever? Logs can duplicate sensitive content and expand breach/retention scope.
  8. What proves safe launch? Versioned multidimensional evaluation, human/process tests, red-team evidence, monitoring, rollback, and accountable approval.

Lesson acceptance

Pass when every input, source, generated claim, proposed action, approval, effect, log, incident, cost, and deletion has an owner and test. Fail if the design treats a guardrail as complete safety, lets prompts grant permission, executes free-form output, exposes cross-tenant retrieval, hides uncertainty from reviewers, lacks a kill switch, or uses human approval as an unmeasured checkbox.

Official sources

Advertisement