Lesson 307 · AWS Learning Path

AWS 307: Amazon SageMaker AI and managed AI services

· Published · 18 min read

Labelled process diagram for AWS 307: Business task and governed data to Managed AI API or SageMaker lifecycle to Model inference and control to Quality, safety, latency, and human-review evidence, with decision...

Why this lesson matters

Machine learning does not convert data into facts. It learns statistical behavior from examples and returns predictions with uncertainty. A model can score well on a test set and still fail because the data leaked the answer, the population changed, one group was underrepresented, the threshold was wrong, or the application treated a confidence score as guaranteed truth.

AWS offers several layers. Purpose-built AI services expose APIs for document extraction, text analysis, image analysis, speech transcription, and translation. Amazon Bedrock provides managed foundation models and generative-AI application capabilities. Amazon SageMaker AI provides broad control over preparing data, building, training, tuning, evaluating, registering, deploying, and monitoring custom predictive, classical ML, and foundation models. Amazon SageMaker Unified Studio is a wider data and AI development environment; it does not make every feature inside it a SageMaker AI feature.

The fastest API is not automatically the safest choice. A receipt field extracted by Textract, face similarity from Rekognition, medical transcript, fraud score, translation, or model-generated answer can affect money, privacy, access, or health. Architects must define the decision, acceptable errors, evidence, human authority, data rights, rollback, and operating cost before selecting the service.

This lesson builds that decision discipline from first principles. You will design four workloads, follow one custom model through its entire lifecycle, select an inference mode, test failure and drift, and produce evidence that another engineer or risk owner can challenge.

What you will be able to do

By the end, you can:

  • explain features, labels, training, validation, testing, inference, confidence, threshold, overfitting, leakage, drift, and bias in plain language;
  • distinguish SageMaker AI, SageMaker Unified Studio, Amazon Bedrock, purpose-built AI APIs, and ordinary deterministic code;
  • select a managed API when its task and limits fit, and justify SageMaker AI when custom model control is required;
  • map SageMaker data preparation, training, experiments, tuning, pipelines, registry, deployment, monitoring, and governance;
  • design representative datasets with ownership, lineage, consent, retention, split, and label-quality evidence;
  • select metrics and operating thresholds from the cost of false positives and false negatives;
  • distinguish real-time, serverless, asynchronous, and batch inference;
  • deploy with immutable model, container, code, schema, and configuration versions plus safe rollback;
  • detect data quality, model quality, bias, feature-attribution, and concept-drift problems;
  • apply IAM, VPC, S3, KMS, ECR, logging, artifact, and supply-chain controls;
  • explain Comprehend, Textract, Rekognition, Transcribe, Translate, and human-review boundaries;
  • handle AI-services content-use opt-out policy as a governance decision, not an assumed default;
  • diagnose training, endpoint, quota, latency, quality, and managed-API failures from evidence;
  • estimate training, endpoint, serverless, storage, data, API, review, and idle-resource costs; and
  • produce a reviewable AI/ML architecture decision dossier.

Before you start

  • Complete AWS049 through AWS051 for IAM, AWS160 through AWS165 for data security, and AWS300 through AWS305 for governed analytics.
  • Use synthetic, public-domain, or explicitly approved data. Do not upload customer records, faces, voices, health data, secrets, copyrighted corpora, or production documents for practice.
  • The default T0 path creates nothing. The optional T1 path requires account-owner approval, a budget alarm, a dedicated sandbox, an approved dataset, and a deletion plan.
  • Never use an AI result as the sole basis for a legal, employment, lending, health, identity, safety, or access decision without required policy, validation, and qualified human review.
  • Confirm service and feature availability in the selected Region. Languages, model versions, instance types, accelerators, customization, PII handling, and synchronous/asynchronous operations vary.
  • Redact account IDs, role ARNs, endpoint names, S3 locations, model artifacts, labels, prompts, output, and business thresholds from shared evidence.
  • A notebook is an interactive development tool, not a production deployment, evidence archive, or approval process.

Create a local evidence directory:

mkdir -p "$HOME/aws307-evidence"
cd "$HOME/aws307-evidence"

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
date -u +%FT%TZ

1. Learn the ML vocabulary before choosing a service

A feature is an input used by a model. A label is the known answer used for supervised training, such as whether a transaction was confirmed fraudulent. Training adjusts model parameters from examples. Inference applies the trained model to new input. A model artifact is learned state, not the training program or serving container.

Split data by purpose:

  • the training set fits the model;
  • the validation set selects algorithms, hyperparameters, and thresholds; and
  • the untouched test set estimates final performance.

If records from the same customer, patient, device, document, or future time cross these boundaries, the model can appear better than it is. Data leakage occurs when training includes information unavailable at prediction time or indirectly reveals the answer. Random splitting is wrong for many time-series and grouped-entity problems.

Overfitting means learning training-specific noise rather than general behavior. Underfitting means the model is too simple or insufficiently trained. Data drift means input distribution changed. Concept drift means the relationship between input and correct outcome changed. Drift is evidence to investigate, not automatic proof that retraining is safe.

For binary classification:

Actual and predicted resultMeaningBusiness question
True positiveCorrectly predicted positiveWhat useful action follows?
False positivePredicted positive but actually negativeWhat is the cost of needless review, block, or alarm?
False negativePredicted negative but actually positiveWhat is the cost of missing the event?
True negativeCorrectly predicted negativeDoes class imbalance make this dominate accuracy?

Precision asks how many predicted positives were correct. Recall asks how many actual positives were found. A threshold trades these errors. Accuracy can be misleading when 99.9 percent of events are normal. Regression work may use MAE, RMSE, quantile loss, or domain tolerances. Ranking, forecasting, detection, extraction, translation, speech, vision, and generative tasks need different evaluation methods.

2. Select the correct AWS layer

business decision and governed data
             |
             v
deterministic rule sufficient? ---- yes ---> ordinary code or rules
             |
             no
             v
standard supported perception task? -------> purpose-built AI API
             |
             no or customization required
             v
foundation-model application? -------------> Amazon Bedrock
             |
             custom lifecycle/infrastructure control
             v
Amazon SageMaker AI
NeedLikely directionReject or reconsider when
Stable exact business ruleDeterministic codeProbabilistic errors add no value.
Entities, sentiment, PII, text classificationAmazon ComprehendLanguage, task, volume, or custom-quality need does not fit.
OCR, forms, tables, queries, expense or identity-document extractionAmazon TextractA deterministic parser is sufficient or unsupported document variation dominates.
Image/video labels, moderation, faces, livenessAmazon RekognitionConsent, law, demographic performance, or human-review requirements cannot be met.
Batch or streaming speech to textAmazon TranscribeLanguage, audio quality, latency, privacy, or domain vocabulary is unsuitable.
Neural text/document translationAmazon TranslateQualified human translation is required or language/domain quality fails.
Prompting, RAG, agents, managed foundation modelsAmazon BedrockA smaller deterministic or purpose-built solution is safer and cheaper.
Custom training, tuning, containers, lifecycle, or inference controlSageMaker AIThe team cannot own data/model operations or a managed API already meets the need.
Shared governed data/analytics/AI development experienceSageMaker Unified StudioDo not confuse workspace governance with proven model quality.

SageMaker AI was formerly named Amazon SageMaker. Current documentation uses SageMaker AI for ML build, train, and deploy capabilities. Unified Studio brings data, analytics, and AI work into projects and spaces with governance integration. Bedrock and SageMaker AI can be combined: prototype with managed foundation models, use SageMaker when infrastructure and customization need deeper control, or import eligible customized models into supported managed inference. Record the exact product, feature, Region, and model version rather than writing only "use SageMaker."

SageMaker Edge Manager reached end of life on April 26, 2024. Its APIs and managed fleets are unavailable. Do not include it in new edge architecture; assess currently supported device runtime and fleet controls such as those covered in AWS306.

3. Start with the decision and baseline

Write a one-page problem contract before collecting data:

  1. user and decision being supported;
  2. current deterministic or human baseline;
  3. prediction target and time horizon;
  4. action taken from each output;
  5. false-positive and false-negative harm;
  6. prohibited inputs and uses;
  7. required latency, throughput, availability, and explainability;
  8. qualified human review and appeal path;
  9. measurable launch and stop criteria; and
  10. owner for data, model, application, risk, and retirement.

"Predict churn" is incomplete. "Every Monday, rank opted-in customers for retention review, improve recall at the approved contact-capacity precision, never auto-cancel or alter price, and allow an analyst to reject the recommendation" is testable.

Build a non-ML baseline. A simple rule or linear model may be easier to explain and cheaper to operate. A complex model must beat it on business-relevant holdout evidence, not just one aggregate metric.

4. Govern data and labels

Create a data sheet containing source owner, collection purpose, lawful basis or consent, schema, event time, geographic scope, population, missingness, known exclusions, retention, license, encryption, quality checks, and deletion route. Record where a label came from and how long after the event it becomes reliable.

Common traps include:

  • a fraud label that actually means "chargeback was filed," missing unreported fraud;
  • an outcome recorded after the prediction time leaking into features;
  • duplicate documents in train and test;
  • images from one camera in training and the same event in testing;
  • majority-reviewer opinion treated as objective truth;
  • historical decisions encoding earlier discrimination; and
  • a convenient dataset that does not represent production languages, devices, seasons, or subgroups.

SageMaker Processing can run repeatable data preparation and evaluation jobs. Data Wrangler, Feature Store, Ground Truth, and other tools can help depending on the workflow, but none automatically gives permission, representative coverage, or correct labels. Feature Store needs entity keys, event time, online/offline consistency, freshness, deletion, and lineage design. Avoid training-serving skew by using the same governed transformation logic or an independently tested equivalent.

Use immutable S3 object versions or content hashes for input manifests. Encrypt with an approved KMS key, restrict bucket and access-point policies, block public access, and log authorized access. If the workload uses a VPC, distinguish network isolation from service access: private subnets may still need S3 and ECR endpoints, DNS, CloudWatch, STS, and approved package repositories.

5. Build and train reproducibly

A SageMaker training job starts managed compute, retrieves an algorithm or container, reads input channels, trains, writes output artifacts to S3, emits logs/metrics, and releases training compute when complete. You still pay for job compute and related storage/data movement, and notebook or Studio applications can remain billable after the job.

versioned data + processing code + container digest + parameters
                         |
                         v
                  training job
                 /      |       \
          logs/metrics  model.tar.gz  lineage
                         |
                         v
             evaluation on locked test data

Choose a built-in algorithm, framework container, custom script, or custom container according to support and control needs. Pin image digests and dependencies. Scan custom images in ECR, generate a software bill of materials where required, restrict outbound network access, and never download unverified code or models during a production job.

Hyperparameter tuning launches multiple jobs to search configured ranges against an objective metric. It can spend far more than one job and overfit the validation set through repeated selection. Set maximum jobs, parallel jobs, early stopping where suitable, instance limits, budget alarms, and a separate final test.

Experiments or managed MLflow can track runs, parameters, metrics, artifacts, and lineage. Tracking makes work reproducible only if code, data, environment, seed, dependencies, and evaluation procedure are also versioned. A notebook output without these inputs is not a recoverable experiment.

Read-only inventory:

aws sagemaker list-domains --max-results 20 --output table
aws sagemaker list-training-jobs --max-results 20 --output table
aws sagemaker list-processing-jobs --max-results 20 --output table
aws sagemaker list-hyper-parameter-tuning-jobs --max-results 20 --output table
aws sagemaker list-pipelines --max-results 20 --output table

6. Evaluate the model, not one number

Lock an evaluation protocol before viewing final results. Include the baseline, primary metric, secondary guardrail metrics, confidence intervals where appropriate, subgroup slices, threshold-selection rule, stress tests, missing/invalid inputs, and an error-review sample.

For a fraud model, report precision and recall across thresholds, cost-weighted error, calibration, transaction-value bands, regions, customer tenure, device types, and delayed labels. A probability-like output is not necessarily calibrated. If the application interprets 0.8 as an 80 percent likelihood, calibration must be tested.

SageMaker Clarify can help analyze pre-training and post-training bias, feature attribution, and explainability. These reports are evidence, not a declaration that a system is fair. The protected or operational groups that matter must be defined with legal and domain experts, and proxy variables can preserve harmful effects even if a sensitive column is removed.

Model cards can record intended use, risk rating, training details, evaluation, limitations, and approvals. A model package in Model Registry binds a deployable artifact and inference specification to version and approval state. Approval must be based on traceable evidence, not merely changing a status field.

Reject launch when the test population is unrepresentative, subgroup harm exceeds policy, labels are unreliable, the threshold cannot meet operational capacity, sensitive input is unapproved, or the human-review process has no authority or response time.

7. Automate promotion with controlled pipelines

SageMaker Pipelines can connect processing, training, evaluation, condition, registration, and deployment-related steps. A typical flow is:

validate data -> process -> train -> evaluate -> quality gate
                                              |
                              fail <----------+----------> register candidate
                                                               |
                                                     independent approval
                                                               |
                                                    staged deployment

Use idempotent paths and unique execution IDs. A rerun must not silently overwrite evidence for an earlier model. Separate pipeline execution authority, model approval, and production deployment. Store evaluation output and approval identity. Infrastructure as code should define roles, network, encryption, alarms, and endpoint configuration.

A pipeline that completes successfully proves orchestration, not model suitability. Condition steps compare machine-readable criteria, but humans must approve policy exceptions, new intended uses, or material risk changes.

8. Choose one of four inference modes

ModeBest fitImportant behavior
Real-time endpointSustained interactive low-latency requestsPersistent instances, autoscaling, endpoint variants, and idle cost
Serverless inferenceIntermittent requests that tolerate cold startsMemory-based managed capacity; payload and feature limits differ
Asynchronous inferenceLarge payloads or long near-real-time processingQueued work, output storage, notifications, and scale-to-zero support
Batch TransformWhole offline datasetsTemporary job compute, S3 input/output, partitioning, and record association

Current AWS documentation lists ordinary real-time requests for low-latency work, serverless for intermittent traffic, asynchronous for payloads up to 1 GB and processing up to one hour, and batch for large offline datasets. Verify exact Region, container, feature, and quota limits when implementing.

A SageMaker model references model data and a serving image. An endpoint configuration references variants and instances. The endpoint is the callable resource. Production evidence must bind the model package, artifact checksum, container digest, schema, endpoint configuration, scaling policy, timeout, retry behavior, data-capture policy, load test, canary strategy, and rollback configuration.

Autoscaling reacts after a metric changes; it does not remove model-loading or capacity-start time. Measure end-to-end p50, p95, and p99 latency, queueing, errors, model latency, resource use, invocations per instance, and cost per accepted prediction.

9. Monitor quality and operate change

CloudWatch provides job and endpoint infrastructure evidence. Model Monitor can schedule checks for data quality, model quality, bias drift, and feature-attribution drift where configured. Data capture creates a sensitive storage boundary. Each monitor needs a baseline, constraints, schedule, destination, alert owner, and response.

Separate:

  • service health: errors, latency, throttling, saturation;
  • data health: schema, nulls, ranges, categories, drift;
  • model health: precision, recall, calibration, subgroup behavior;
  • business health: reviewer workload, benefit, and customer harm; and
  • governance health: version approval, evidence, access, and retention.

Model-quality monitoring needs delayed ground truth joined to predictions by a stable ID. If labels arrive after 30 days, today's dashboard cannot prove today's accuracy. Drift may indicate a source bug, attack, seasonality, policy change, or concept change. Diagnose first. Retraining is a new model change that still requires versioned data, evaluation, approval, canary release, and rollback.

10. Secure the complete ML path

Humans use IAM identities; SageMaker assumes execution roles. Scope roles to exact S3 prefixes, KMS keys, ECR repositories, logs, registry resources, and required actions. Restrict iam:PassRole, because passing a powerful execution role can become privilege escalation.

Use separate development and production boundaries, approved VPC paths, TLS, KMS encryption, Secrets Manager or short-lived credentials, pinned and scanned ECR images, CloudTrail, CloudWatch, data minimization, retention/deletion controls, and incident response for poisoned data, stolen models, abusive inference, or leaked output.

Network isolation does not limit what a broad role can read. Encryption does not make prohibited data approved. Explainability does not make an unsafe decision safe.

11. Understand each purpose-built AI service

Amazon Comprehend

Comprehend analyzes supported-language text for entities, key phrases, sentiment, syntax, events, targeted sentiment, and PII. It supports custom classification and custom entity recognition. Synchronous calls handle immediate supported inputs; asynchronous jobs process S3 datasets. Scores are probabilistic, and PII detection can miss content.

Amazon Textract

Textract detects text and analyzes forms, tables, selections, signatures, queries, and layout. Specialized APIs cover expenses and other supported document types. Asynchronous start/get workflows suit multipage documents. Preserve page, geometry, relationship, confidence, and model-version evidence. Custom Queries adapters still require representative labeled documents and evaluation.

Amazon Rekognition

Rekognition analyzes images and video for labels, moderation, text, faces, liveness, and other supported features. Face detection asks whether a face exists; comparison estimates similarity. Similarity is not identity proof. Thresholds must reflect false-match and missed-match harm. AWS recommends human review before comparison affects rights, privacy, or access.

Amazon Transcribe

Transcribe provides batch and streaming speech recognition. Languages and features such as custom vocabulary, speaker/channel handling, analytics, and PII processing vary. Noise, accent, cross-talk, and domain vocabulary affect quality. PII redaction is predictive and may miss data; AWS states that it is not medical de-identification. Clinical output requires trained professional review.

Amazon Translate

Translate provides supported-language text and document translation. Custom terminology guides preferred terms; parallel data influences supported batch output without training a separate model. Neither guarantees legally, clinically, or technically correct translation. Preserve source, language codes, terminology version, and reviewer evidence.

Human review

Amazon Augmented AI can route supported service or custom-model output to human workflows. Define activation thresholds, reviewer qualifications, work-team access, sensitive-data exposure, task instructions, service level, sampling, disagreement resolution, and final authority. Human review is a designed control, not a checkbox.

aws comprehend list-document-classifiers --output table
aws textract list-adapters --max-results 20 --output table
aws rekognition list-collections --max-results 20 --output table
aws transcribe list-transcription-jobs --max-results 20 --output table
aws translate list-terminologies --max-results 20 --output table

12. Govern content use and high-impact decisions

AWS Organizations AI-services opt-out policies control whether customer content for supported services may be stored and used for service improvement. Effective policy can be inherited, and an organization can opt out all current and future supported services. Review the current supported-service list, service terms, inheritance, and effective policy with security and legal owners.

Do not claim that opt-out deletes workload data or that every service has identical terms. The policy concerns service-improvement content; data needed to provide a service follows its own lifecycle. Data residency, biometrics, recording consent, child data, health information, copyright, and automated decisions require jurisdiction-specific approval.

13. Complete the four-workload workshop

Create aws307-decision-dossier.md and compare deterministic code, a purpose-built API, Bedrock where relevant, and SageMaker AI for:

  1. invoice extraction with low-confidence accounts-payable review;
  2. approved support-call transcription, redaction, and analytics;
  3. a custom fraud-prioritization model with delayed labels and analyst review; and
  4. safety-instruction translation with controlled terminology and a qualified linguist.

For each, record data class, permission, Region, operation mode, representative evaluation set, metrics, threshold, human step, failure path, monitoring, quota, unit cost, retention, and rejection rule.

The fraud design must trace:

approved snapshots -> processing -> grouped/time-aware split -> baseline
  -> bounded training/tuning -> locked evaluation -> subgroup review
  -> model package -> independent approval -> shadow/canary endpoint
  -> analyst queue -> delayed labels -> monitors -> rollback/retraining

Inject a renamed schema column, tripled positive-class rate, failed p99 objective, and subgroup false-negative breach. For each, state evidence, containment, owner, rollback decision, and durable correction.

14. Diagnose from evidence

SymptomEvidenceSafe response
Training fails before startFailure reason, role, S3/KMS/ECR, subnet capacityCorrect the smallest failed path; do not add administrator access.
Training metric missingLogs, metric name/regex, code outputEmit and verify deterministic metrics before tuning.
Validation good, production poorSplit, leakage, duplicates, time/entity slicesStop promotion and rebuild representative evaluation.
Endpoint creation failsContainer logs, model format, artifact access, health timeout, quotaCorrect package or proven capacity issue, then retest.
Peak latency is highinvocation/model latency, concurrency, scaling lag, dependenciesLoad-test, tune, scale with headroom, and protect callers.
Prediction changes unexpectedlymodel/config/image/schema/upstream versionsRestore immutable prior configuration and trace promotion.
Drift alert firesbaseline, input sample, releases, seasonality, source healthDiagnose; never auto-promote a retrained model.
Textract field is wrongblocks, geometry, confidence, image quality, adapterRoute to review and preserve source evidence.
Face match rejects userimage quality, threshold, population testsOffer alternate verification and qualified review.
Transcript leaks PIIaudio, output tokens, mode/language supportStop release; use independent controls and validation.

15. Cost, quotas, and cleanup

SageMaker cost can include Studio applications, processing, training, tuning trials, HyperPod, endpoints, serverless duration/memory, asynchronous hosting, Batch Transform, Feature Store, MLflow, storage, capture, monitoring, labeling/review, ECR, S3, KMS, logs, NAT, and transfer. Real-time endpoints cost while provisioned even when idle.

Purpose-built services charge by service-specific units such as text, pages, images, video, audio, characters, customization, or requests. Human review adds task and workforce cost. Price normal volume, retries, peak quota, low-confidence review rate, and retained evidence.

For T0, prove no creation. For an approved experiment, remove only exact owned applications, endpoints before endpoint configurations, models, monitors, pipelines, model packages, outputs, feature groups, custom AI resources, collections, transcripts, logs, images, and buckets according to retention approval. Verify readback and later billing data.

Practical submission

Submit:

  1. ML glossary and confusion matrix;
  2. problem contract and non-ML baseline;
  3. four-workload service decision matrix;
  4. data and label-quality sheet;
  5. leakage-resistant split and reproducibility manifest;
  6. bounded training/tuning and experiment lineage;
  7. metrics, threshold, calibration, subgroup, and error review;
  8. model card and approval evidence;
  9. pipeline and separation-of-duties design;
  10. inference and deployment/rollback specification;
  11. service/data/model/business monitoring;
  12. purpose-built API confidence and review rules;
  13. security, opt-out, privacy, and retention decision;
  14. four failure diagnoses;
  15. cost/quota worksheet; and
  16. cleanup or no-create proof.

Knowledge check

  1. Why can accuracy mislead? Expected direction: class imbalance can hide failure on the rare class.
  2. Why protect the test set? Expected direction: repeated tuning against it makes the estimate optimistic.
  3. When prefer a managed AI API? Expected direction: its supported task, quality, privacy, latency, and cost fit without custom lifecycle ownership.
  4. Bedrock versus SageMaker AI? Expected direction: managed foundation-model applications versus deeper custom training, lifecycle, infrastructure, and inference control.
  5. Why is registry approval insufficient alone? Expected direction: verify evidence, reviewer, intended use, and tests.
  6. Which mode fits a 500 MB, 20-minute request? Expected direction: assess asynchronous inference and verify current limits.
  7. Why not auto-promote after drift? Expected direction: drift has multiple causes and a new model still requires validation.
  8. Confidence versus correctness? Expected direction: confidence is a score whose calibration and errors require testing.
  9. Why is PII redaction not the only control? Expected direction: predictive detection can miss content and coverage varies.
  10. What changed for Edge Manager? Expected direction: it ended service in 2024 and is unavailable.

Lesson acceptance

You pass when a reviewer can trace each workload from a legitimate decision and governed source through service selection, error costs, version, inference, human authority, monitoring, incident response, cost, and deletion. The custom model must be reproducible, leakage-resistant, subgroup-tested, independently approved, safely deployable, observable after labels arrive, and rollback-ready.

Fail if the design starts training before defining the decision, trusts one accuracy number, uses unapproved data, treats confidence as truth, selects an unavailable service, deploys an unversioned notebook artifact, exposes a broad role, auto-promotes retraining, omits high-impact human review, or leaves a billable endpoint ownerless.

Official sources

Advertisement