AWS 312: Amazon Bedrock architecture, guardrails, identity, and human approval
Why this lesson matters
Amazon Bedrock is a managed platform for building generative-AI applications with foundation models and capabilities such as Guardrails, Knowledge Bases, Agents, Flows, prompt management, evaluation, model customization, and managed inference choices. It removes model-server operations for many workloads, but the application still owns user authorization, data rights, source quality, prompt injection defenses, tool permissions, human approval, observability, and business outcomes.
A secure Bedrock request is more than InvokeModel. It starts with an authenticated application user, assumes an application role, selects an allowed model or inference profile, retrieves only authorized data, invokes with a versioned prompt and guardrail, validates the response, and may call a separately authorized tool. Each hop has a different identity and failure mode.
Guardrails are useful policy controls, not a universal security boundary. Their scope depends on the API and how they are attached. AWS documents IAM condition-key and orchestration limitations, and input tags can affect where input policy is applied even though output remains evaluated. Architects must test the exact Converse, Invoke, Knowledge Base, Agent, or application path.
What you will be able to do
By the end, you can:
- distinguish Bedrock control plane, runtime, model, inference profile, provisioned throughput, guardrail, Knowledge Base, Agent, Flow, and tool;
- select Bedrock versus SageMaker AI and purpose-built AI services;
- trace application user identity separately from AWS runtime identity;
- enforce allowed models, Regions, guardrails, retrieval sources, and tools with layered policy;
- design direct model invocation, retrieval-augmented generation, and agentic workflows;
- evaluate Guardrails content, denied-topic, word, sensitive-information, contextual-grounding, and automated-reasoning policies where supported;
- account for geographic and global cross-Region inference data movement;
- build human approval around material actions;
- monitor tokens, latency, throttling, safety interventions, retrieval, tools, and cost;
- diagnose access, model, guardrail, knowledge, agent, quota, and output failures; and
- produce a secure support-assistant architecture and threat model.
Before you start
- Complete AWS311. Its governance and evaluation controls remain mandatory even when Bedrock features are enabled.
- Use synthetic prompts and public or owned documents. Never test with production secrets or customer data.
- Confirm model/provider terms, model ID/version, Region, throughput mode, data-residency requirement, quotas, and feature compatibility.
- Do not enable an action group or tool that can write to production in a lab.
- Redact guardrail IDs, role ARNs, knowledge-base IDs, prompts, traces, retrieved text, model outputs, endpoint details, and account IDs.
- T0 is read-only inventory/design. T1 needs an approved model, token/spend limit, CloudWatch controls, and cleanup.
1. Build the Bedrock mental model
application user identity
-> application/API authorization
-> AWS workload role
-> Bedrock runtime API
+--> foundation model or inference profile
+--> guardrail evaluation
+--> Knowledge Base retrieval
+--> Agent/Flow orchestration
-> action group/tool role -> external system
-> output validation -> human approval -> bounded effect
The control plane lists or configures models, guardrails, knowledge bases, agents, prompts, and throughput. The runtime plane invokes models, converses, retrieves/generates, invokes agents, applies guardrails, or performs data automation according to feature.
Model access and commercial availability vary by provider, model, Region, and account. A model identifier is not a stable business abstraction. Keep an application-level model profile that records provider, exact model/version, supported modalities/context, Region, inference mode, price, evaluation, fallback, and retirement.
2. Select the right capability
| Need | Direction | Boundary |
|---|---|---|
| Call a supported foundation model with prompts | Bedrock model runtime | Application owns prompt/output policy and evaluation. |
| Multi-turn normalized messages and tool definitions | Converse API where model supports it | Model behavior and tool loop remain application concerns. |
| Retrieve enterprise sources for grounded response | Knowledge Bases | Ingestion, chunking, authorization, freshness, and source quality remain yours. |
| Model plans and calls actions | Agents or application orchestration | Least privilege, deterministic validation, approval, and idempotency remain yours. |
| Reusable orchestration graph | Bedrock Flows where fit | Do not hide business state, compensation, or approvals in an opaque graph. |
| Content and sensitive-data policy | Guardrails | Not factuality, complete injection prevention, or user authorization. |
| Custom model training/infrastructure control | SageMaker AI | Choose from requirements, not product preference. |
| Standard OCR/speech/translation/classification | Purpose-built AI API | Often simpler, more testable, and cheaper. |
Start with direct invocation. Add retrieval only when current source evidence is required. Add agents only when dynamic planning materially helps and tool effects can be bounded. Every layer adds latency, cost, failure modes, and attack surface.
3. Design IAM and identity propagation
The application must know the end user, tenant, role, and request purpose. Bedrock receives the AWS principal invoking the service; it does not automatically enforce your application's tenant model.
Use a workload role with only required runtime actions and resources. Separate administrators who create guardrails/agents from applications that invoke approved versions. Control:
- allowed model and inference-profile ARNs;
bedrock:InvokeModeland streaming equivalents;- guardrail identifier/version;
- Knowledge Base and Agent resources;
- KMS, S3, OpenSearch/vector store, Secrets Manager, and logging access;
- tool Lambda or service role;
- Region through organization/IAM controls; and
iam:PassRolefor Bedrock service roles.
Resource-level support differs by action. Test policies with real denied and allowed calls. A broad bedrock:* on * can expose new models/features later.
AWS documents bedrock:GuardrailIdentifier condition-key enforcement for specified invocation paths, but orchestration APIs can make internal calls that do not carry the same parameter and may receive access denied. Input tags can also affect input evaluation scope. Use IAM enforcement only with exact API-path tests, application-side guardrail attachment, output checks, and regression tests.
Never pass a caller-supplied role ARN or tool name directly to execution. Map application permissions to fixed, reviewed capabilities.
4. Invoke models safely
Prefer a normalized request layer that accepts an internal schema, validates input, chooses an approved model profile, sets bounded generation parameters, attaches the required guardrail, records a request ID, and returns a typed response.
{
"requestId": "synthetic-123",
"useCase": "support-draft-v1",
"userContext": {"tenant": "tenant-a", "role": "support-agent"},
"input": {"question": "synthetic question"},
"controls": {
"modelProfile": "approved-support-model-v3",
"guardrailVersion": "7",
"maxOutputTokens": 800,
"toolsAllowed": []
}
}
Do not let users choose arbitrary models, token counts, system prompts, guardrail IDs, or tool schemas. Set request/body/time limits and rate limits. Streaming improves perceived latency but output can arrive before the full result is known; buffer or incrementally evaluate content according to the risk.
Retries after timeout can duplicate token cost and agent/tool work. Give each request and action an idempotency strategy. Retry throttling and transient service errors with bounded exponential backoff and jitter, not every validation or safety rejection.
5. Use inference profiles and understand data movement
On-demand inference can use a model directly or supported inference profile. Provisioned throughput reserves model capacity for supported models/terms and can create commitment and utilization risk.
Cross-Region inference profiles route requests among listed Regions to improve throughput and resilience. Geographic profiles stay within a documented geography. Global profiles can route across commercial Regions worldwide. Pricing and quota behavior depend on profile and source Region.
Do not call cross-Region inference “single-Region.” Prompts and outputs can move outside the source Region. Review destination Regions, data classification, provider/model availability, SCPs, KMS/data dependencies, logging, legal restrictions, and failover tests. A guardrail profile can likewise evaluate across Regions within its supported geography.
6. Configure and test Guardrails
Guardrails can combine supported policies such as:
- denied topics;
- harmful-content categories and strengths;
- word/phrase filters;
- sensitive-information detection and masking/blocking;
- contextual grounding checks;
- automated reasoning checks where available; and
- blocked-message behavior.
Version guardrails. DRAFT changes are mutable; application promotion should bind an approved version and evaluation report. Test false positives, false negatives, multiple languages, misspellings, encoding, long context, streaming, retrieved content, model output, and each API path.
aws bedrock list-guardrails --max-results 20 --output table
aws bedrock get-guardrail --guardrail-identifier REDACTED --guardrail-version REDACTED
ApplyGuardrail can evaluate text independently of model invocation, useful for input/output or non-Bedrock content paths. It still does not authorize the user, verify facts, sanitize executable content, or guarantee all sensitive data is found.
Store intervention category and trace metadata without unnecessarily retaining harmful or sensitive text. Provide safe user messages that do not reveal filter internals.
7. Build authorized Knowledge Bases
A Knowledge Base connects a data source, ingestion/chunking/embedding process, vector or supported structured store, retrieval configuration, and model response path. Design:
- approved source owner and classification;
- malware/content/injection screening;
- document/chunk metadata and tenant/access tags;
- embedding model and dimension compatibility;
- sync, failure, deletion, and stale-document reconciliation;
- retrieval count/search type/reranking;
- requesting-user authorization before context;
- citation and insufficient-evidence behavior; and
- source and vector-store encryption/access/logging.
aws bedrock-agent list-knowledge-bases --max-results 20 --output table
aws bedrock-agent list-agents --max-results 20 --output table
Knowledge Base retrieval may use a service role with broader source access than the end user. The application must enforce tenant/document entitlements, often through metadata filters and separate stores where stronger isolation is required. Test revoked access and cached results.
Chunking affects meaning. Small chunks lose context; large chunks add irrelevant or malicious text and token cost. Tables, images, headers, footnotes, and versioned policies require format-aware tests. Reconcile source count, indexed count, failed documents, deletes, and timestamps.
8. Bound Agents, Flows, and tools
An Agent can use instructions, action groups, Knowledge Bases, memory/session context, and orchestration traces. The model may choose tools and arguments, but deterministic code must authorize and validate every proposed action.
Action-group Lambda should:
- accept only a strict schema;
- derive tenant/user from trusted context;
- reject caller-supplied privilege fields;
- re-read current resource state;
- enforce limits and policy;
- use idempotency keys;
- return bounded, non-secret output;
- log request/action/result IDs; and
- distinguish retryable from permanent errors.
Use return-control or an approval workflow for material actions. Never hide “human approval” as a tool the model can self-approve. The approver must see exact target, effect, evidence, risk, and rollback.
Limit agent turns, tool calls, time, tokens, and total cost. Detect loops and repeated identical calls. Do not expose shell, arbitrary HTTP, broad SQL, dynamic IAM, or unrestricted code execution.
9. Protect prompts, sessions, and observations
Version system prompts, prompt templates, examples, and tool schemas. Prompt management can organize variants but approval and regression testing remain required. Keep secrets out of prompts; model context is not a secret store.
Session IDs, conversation state, memory, traces, and logs can contain customer content. Use unpredictable identifiers, enforce ownership on every request, set expiration/deletion, and prevent users from attaching to another session. Avoid placing sensitive data in CloudWatch dimensions or X-Ray annotations.
Bedrock invocation logging is deliberate and optional; if enabled, scope destination, encryption, access, content fields, retention, and deletion. CloudTrail records control-plane/runtime events according to current support, but not every business fact needed for audit. Add application-level evidence.
10. Evaluate, deploy, and monitor
Use a fixed evaluation suite from AWS311 plus Bedrock-specific cases for model/profile versions, Guardrail versions, retrieval, agent traces, tool schemas, and cross-Region behavior. Compare quality, grounding, safety, tool correctness, latency, token use, and cost.
Promote through development, shadow, internal pilot, limited users, and broader release. Keep a previous model/prompt/guardrail/knowledge/tool bundle. Rollback all coupled versions, not just the model.
Monitor:
- invocation count, latency, errors, throttles, input/output tokens;
- guardrail interventions by policy/category;
- retrieval empty/stale/wrong-tenant/citation failures;
- agent turns, tool calls, denials, retries, loops, and ambiguous results;
- human approval, rejection, edit, queue age, and incident rate;
- cost by application/tenant/use case; and
- model/profile/guardrail/source/tool configuration changes.
11. Diagnose from evidence
| Symptom | Evidence | Response |
|---|---|---|
| Access denied invoking model | principal, action, model/profile ARN, Region, SCP, guardrail condition | Correct the exact identity/resource/API mismatch. |
| Model unavailable or throttled | model/Region/profile, quota, token/request rate, service health | Use approved profile/backoff/fallback; never silently change behavior. |
| Guardrail not applied | API path, request fields, identifier/version, IAM, trace | Stop affected path and enforce/test at application and IAM layers. |
| Valid input blocked | intervention category, language, policy version, false-positive set | Safe fallback/review, governed policy tuning, regression test. |
| Knowledge answer cites wrong tenant | actor, metadata filter, retrieved chunks, store/cache | Disable retrieval, investigate breach, repair isolation and revocation tests. |
| Ingestion says complete but content missing | source manifest, sync job failures, parsing/chunking, deletes | Reconcile counts and re-ingest only corrected approved sources. |
| Agent repeats action | trace, session, tool result, idempotency record, timeout | Stop loop, reconcile target, return explicit terminal state. |
| Costs surge | token/context size, retrieval count, agent turns, abuse, retries | Kill/cap path, rate-limit, optimize only after root cause. |
12. Secure support-assistant workshop
Design an internal support assistant that retrieves tenant-authorized product documentation, drafts answers, and proposes but never executes service credits.
Deliver:
- Bedrock-versus-alternative ADR;
- model/profile/Region/data-movement record;
- user-to-workload identity map and IAM policies;
- direct invocation contract;
- versioned Guardrail and API-path enforcement tests;
- Knowledge Base ingestion/authorization/deletion design;
- citation and insufficient-evidence rules;
- Agent/tool schema with least-privilege role;
- service-credit deterministic checks;
- human approval UI and execution separation;
- session/logging/retention controls;
- 50-case evaluation and injection suite;
- monitoring, quota, token, and cost dashboard;
- eight failure diagnoses;
- kill-switch/rollback game day; and
- cleanup/no-create proof.
Cost and cleanup
Cost includes input/output tokens, model/provider/profile pricing, provisioned throughput commitments, embeddings, retrieval/reranking, vector storage, ingestion, Guardrails text units/policies, Agents/Flows/tool compute, logs, evaluations, human review, transfer, and connected services.
For T0, save redacted inventories only. For a pilot, disable agents/tools first; revoke roles; delete aliases/agents/flows/prompts/guardrail versions/knowledge ingestion resources in safe dependency order; clean vector and source data under retention approval; remove logs, keys/grants, functions, test identities, and provisioned capacity; then verify no scheduled sync or billable store remains.
Knowledge check
- Why does Bedrock not enforce application tenancy? It sees the AWS caller, while the application owns end-user/tenant mapping.
- Guardrail versus authorization? Content policy versus deterministic permission.
- Why test every API path? Direct, retrieval, agent, and internal orchestration calls attach controls differently.
- What changes with cross-Region inference? Requests may be processed in documented destination Regions.
- Why can retrieval leak data? A service role can access more than the user unless per-query authorization is enforced.
- Why re-authorize action groups? Model-selected tools/arguments are untrusted proposals.
- Why is streaming riskier for output control? Content can reach users before full-result validation.
- What must rollback include? Compatible model, prompt, guardrail, knowledge, tool, schema, and application versions.
Lesson acceptance
Pass when the reviewer can trace user identity to every model, source, policy, tool, approval, Region, log, cost, and cleanup action, with negative tests. Fail if wildcard Bedrock permissions, arbitrary model selection, unfiltered cross-tenant retrieval, unapproved cross-Region movement, prompt-based authorization, mutable DRAFT production controls, unrestricted tools, or model-selected self-approval remains.