Free live cohort on Google Meet — register your interest →

Chapter 3 · Domain 1 · 20% of the exam

The AI/ML Lifecycle and MLOps

21 min read · Chapter 3 of 17

On this page
  1. Certification Blueprint
  2. What This Chapter Covers
  3. The Pipeline
  4. The Two Feedback Loops
  5. AWS Services for Each Stage
  6. Where Models Come From
  7. Serving a Model in Production
  8. MLOps
  9. Metrics
  10. Decision Rules and Exam Signals
  11. Distractor Patterns
  12. Scenario Walkthrough
  13. Key Concepts
  14. Revision Flashcards
  15. The Four-Beat Answer
  16. Why This Helps You
  17. Chapter Checklist
  18. After the Chapter

Certification Blueprint

Field Coverage
Exam AWS Certified AI Practitioner (AIF-C01), exam guide v1.1
Domain Content Domain 1 — Fundamentals of AI and ML
Exam weight 20% of scored content
Task statement 1.3 Describe the AI/ML development lifecycle
Objectives 1.3.1 pipeline components · 1.3.2 FM sources · 1.3.3 production methods · 1.3.4 services per stage · 1.3.5 MLOps · 1.3.6 metrics

What This Chapter Covers

Chapters 01 and 02 stopped at the decision — what a system is, and whether AI is the right tool. This chapter is everything after the decision: the stages of building it, where models come from, how they are served, and what keeps them working once real traffic arrives.

Chapters 01-02 answered Chapter 03 answers
What is this system? What are the stages of building it?
Should we use AI at all? Where do models come from?
Which technique and which service? How is a model served in production?
— What keeps it working after launch?

The practice focus is unusual. Task 1.3 questions describe a symptom — error climbing in production, a model that passed evaluation and failed live, a team that cannot reproduce last month's result — and the answer depends on locating when in the lifecycle it happened. The distractors are all real pipeline activities, performed at the wrong point.

This chapter completes Domain 1.

The Pipeline

Five-stage AI/ML pipeline running left to right from data collection, exploratory analysis and pre-processing, through feature engineering, model training and hyperparameter tuning, evaluation against the success metric, and finally deployment and monitoring in production

What to remember from this diagram: the shape is a sequence, but Objective 1.3.1's verb is "describe and differentiate". Sequencing is half the objective. Separating the stages that sit next to each other is the other half, and it is the half that gets tested.

The diagram groups nine named activities into five stages so it stays readable. Learn the nine.

What each stage produces

Name the output, not the activity. The exam asks what a stage hands to the next one.

# Stage Produces
1 Data collection A stored, accessible dataset
2 Exploratory data analysis (EDA) Understanding — distributions, gaps, outliers
3 Data pre-processing Clean, consistent, usable data
4 Feature engineering Model-ready input variables
5 Model training A candidate model
6 Hyperparameter tuning A better-configured candidate model
7 Evaluation A pass/fail verdict against the success metric
8 Deployment A callable model in production
9 Monitoring Drift, quality and cost signals

Two rows repay extra attention. EDA produces understanding, not a changed dataset — that is the whole distinction from pre-processing. And tuning produces a better-configured model, not a better-trained one: training learns parameters from data, while hyperparameters are settings chosen before learning that govern how it happens.

Differentiating adjacent stages

These four pairs are what the exam swaps. Each has a one-word separator.

Pair Separator
EDA vs pre-processing EDA looks; pre-processing changes
Pre-processing vs feature engineering Pre-processing fixes what is wrong; feature engineering creates what is useful
Training vs hyperparameter tuning Training learns the parameters; tuning sets the knobs that govern learning
Evaluation vs monitoring Evaluation is before deployment, once; monitoring is after, continuously

Removing duplicate rows is pre-processing. Deriving "days since last purchase" from a timestamp is feature engineering. The last pair decides more questions than the other three combined: when a scenario describes a live production symptom and offers "evaluate the model", that option is on the wrong side of launch.

The Two Feedback Loops

Pipeline drawn as a cycle in which evaluation failing the success metric loops back to model development, and monitoring detecting data drift in production loops all the way back to data preparation

What to remember from this diagram: a pipeline drawn as a straight line is wrong. Loop 1 fires before launch; loop 2 fires years after it.

Loop From → to Trigger
1 Evaluation → model development The model fails the success bar
2 Monitoring → data preparation Data drift in production

Loop 1 runs many times during the build and nobody outside the team sees it. Loop 2 is the one that gets tested, because it is the one teams forget to build.

Drift is not a bad model

The code has not changed. The model has not been redeployed. Error is climbing anyway.

Evidence Reading
Code and model artefact unchanged Rules out a deployment or algorithm defect
Error growing over time Rules out a training bug — that would be wrong from day one
Input distribution has shifted Data drift

The fix is loop 2: return to the data and re-train. Changing the algorithm is the intuitive move and it is wrong, because nothing in the evidence points at the algorithm.

Compare with Chapter 01's overfitting. Overfitting is high training score and low live score from the start. Drift is good performance that decays. Same symptom shape, different time signature, different fix — and any clause establishing a timeline is what separates them.

AWS Services for Each Stage

Map attaching AWS services to pipeline stages: the data stage to S3, Glue, Glue DataBrew and Lake Formation; build and train to SageMaker AI, SageMaker JumpStart and Bedrock; assisted development to Amazon Q and Kiro; analysis and reporting to Amazon Quick; and serve and monitor to Bedrock, SageMaker endpoints, Model Monitor and CloudWatch

What to remember from this diagram: Objective 1.3.4 names five services explicitly — Bedrock, Amazon Q, Amazon Quick, Kiro, SageMaker AI — and expects you to know which stage each serves, not just what each does.

The scenario says… Service Stage
"Store and catalogue the raw training data" S3, Lake Formation Data collection
"Clean and normalise it without writing code" AWS Glue DataBrew Pre-processing
"Train and deploy our own model" SageMaker AI Training, deployment
"Start from a pre-built model we can adapt" SageMaker JumpStart, Bedrock Sourcing
"Call a foundation model through an API" Amazon Bedrock Deployment/serving
"Help our developers write the code" Amazon Q, Kiro Assisted development
"Build the dashboard the business reads" Amazon Quick Analysis and reporting
"Alert us when live quality degrades" SageMaker Model Monitor, CloudWatch Monitoring

Amazon Q and Kiro are assisted-development tools: they help the people building the pipeline rather than transforming the data flowing through it.

The wrong-stage distractor is this objective's whole strategy — a real service, correctly described, offered for a stage it does not serve. Amazon Quick genuinely builds dashboards; offering it for data pre-processing is not a false statement about the service, it is a false statement about the stage. Asking "which stage?" first is what makes it visible.

Where Models Come From

Decision flow starting from needing a model: if a pre-trained model already does the task it is chosen for speed, otherwise if there is enough labeled data and skill a custom model is trained, and otherwise a pre-trained model is adapted

What to remember from this diagram: "train a custom model" is the answer far less often than learners expect. It requires labeled data and the skills to train, and the scenario has to say so.

Objective 1.3.2 is worded as sources of FM models — the sourcing question on this exam is framed around foundation models.

Source Speed Control Use when
Open source pre-trained Fastest Least Something already does the task well
Adapt a pre-trained model Middle Middle Close but not exact; domain vocabulary differs
Train a custom model Slowest Most Nothing existing fits, and you have data and skills

The trade-off is a straight line: speed and cost at one end, control and specificity at the other.

Serving a Model in Production

Objective 1.3.3 names exactly two methods.

Method What it is You give up You gain
Managed API service Call a model AWS operates, such as Bedrock Control over the runtime and model internals No servers, no scaling, no patching
Self-hosted API Run the model on infrastructure you operate Nothing about control Every operational burden

The axis is operational burden versus control, and nothing else. Cost, performance and scale follow from the choice rather than driving it.

The scenario says… Method Why
"Small team, no ML operations staff" Managed API Operational burden is binding
"Get to production quickly" Managed API No infrastructure to build
"The model weights must stay on our own infrastructure" Self-hosted Control is mandated
"We need a specific model version pinned indefinitely" Self-hosted Managed services move underneath you
"Unpredictable traffic, no idle cost" Managed API Elasticity without capacity planning

"We want control" is not enough. Preference is not a constraint. The scenario must say why control is required — a regulator, a data residency rule, a contractual obligation. Without a stated reason, managed is the better answer.

MLOps

MLOps is what makes a model a maintained system rather than a successful experiment.

Three questions MLOps answers that a merely working model does not:

  • Can we reproduce last month's result?
  • Can we redeploy without a person remembering the steps?
  • Will we know when it stops working?

A model that scores well and cannot answer these is a demo.

Objective 1.3.5 names exactly seven concepts, and you should be able to recite all seven.

Concept What it means in practice
Experimentation Trying variants with results that can be compared and found again
Repeatable processes The same inputs produce the same model, without manual steps
Scalable systems Training and serving grow with data and traffic
Managing technical debt Paying down glue code, dead features and untracked datasets
Achieving production readiness Testing, rollback, monitoring and ownership before launch
Model monitoring Watching live quality, not just uptime
Model re-training A defined trigger and process for refreshing the model

Technical debt earns its place because ML systems accumulate it faster than ordinary software. Every dataset, engineered feature, experiment and model version is a dependency, and unlike code, the data underneath them keeps changing.

Metrics

Decision flow separating business questions from model questions, then routing balanced classes to accuracy and imbalanced classes to precision when false positives cost more, recall when false negatives cost more, and F1 when both matter

What to remember from this diagram: the first branch is the one people skip — is this a model question or a business question at all?

Objective 1.3.6 names accuracy, precision, recall and F1 score.

Metric Optimise when Failure it protects against
Precision A false positive is expensive False alarms
Recall A false negative is dangerous Misses
F1 Both matter; classes are imbalanced —
Accuracy Classes are balanced ⚠️ Misleading when imbalanced

On a dataset that is 99% one class, a model that always guesses that class scores 99% accuracy and has learned nothing. This is the single most testable fact in the objective.

Model metrics versus business metrics

Model metrics Business metrics
Accuracy, precision, recall, F1 Cost per user, development costs, customer feedback, ROI
Tells you is the model good? Tells you was it worth building?

Objective 1.3.6 names both families in the same bullet, deliberately. A model can be excellent and the project still a failure — a 97% accurate model that costs more to operate than the losses it prevents is a technical success and a business failure, and no confusion matrix reveals that.

Expect a question whose correct answer is a business metric while every distractor is a valid model metric. Those distractors are not false statements; they answer a question that was not asked.

Decision Rules and Exam Signals

Rule 1 — name the output, not the activity. Every pipeline stage is identified by what it hands to the next one.

Rule 2 — locate the stage before choosing the fix. Task 1.3 questions give a symptom. The distractors are real activities at the wrong point in the lifecycle.

Rule 3 — drift has a time signature. Unchanged code, plus error growing over time, plus shifted inputs. Overfitting is bad from day one; drift decays.

Rule 4 — ask "which stage?" before accepting a service. The distractor is a real service correctly described and wrongly placed.

Rule 5 — custom training needs data and skills, both stated. Otherwise it is a distractor.

Rule 6 — self-hosting needs a stated reason. Preference is not a constraint.

Rule 7 — check class balance before accepting accuracy. A stated imbalance means accuracy is wrong.

Rule 8 — ask what the question is about. Model quality, or project worth.

Distractor Patterns

Pattern What it looks like How to defuse it
Wrong-stage service Amazon Quick offered for data pre-processing Ask which stage the service serves
Drift read as a bad model "Change the architecture" for a drift symptom Code unchanged + error growing = drift
Evaluation/monitoring swap "Evaluate the model" for a live production symptom Evaluation is once, before launch
Custom model over-selected Training from scratch where a pre-trained model fits Does the scenario supply data and skills?
Accuracy on imbalanced data Accuracy offered where one class is 0.2% Ask for the class balance
Model metric for a business question Precision offered for "did it pay for itself?" Ask what the question is about
"We want control" Self-hosted chosen on preference alone The scenario must say why control is required

The first and the last separate a pass from a fail on this task statement.

Scenario Walkthrough

A retailer deployed a demand-forecasting model fourteen months ago. It passed evaluation at 94% and performed well for a year. Since a range refresh three months ago, forecast error has risen steadily. The code and the deployed model artefact are unchanged. The team asks whether it should switch to a different algorithm.

Evidence Reading Decision
Passed evaluation, performed well for a year The model was correct for its original data Not an algorithm defect
Code and artefact unchanged Rules out deployment and code faults Not a release problem
Error rising since a range refresh The input distribution changed Data drift
Error rising steadily, not from day one Rules out overfitting Confirms drift

The outcome is loop 2 — return to data preparation and re-train on data reflecting the new range. Changing the algorithm addresses a problem the scenario never described.

The load-bearing fact is "since a range refresh". The decoy is "94%", which invites an argument about model quality that has nothing to do with the failure.

Key Concepts

Term Definition
Exploratory data analysis (EDA) Examining data to understand its distributions, gaps and outliers; it produces understanding, not a changed dataset
Data pre-processing Correcting what is wrong with data — missing values, inconsistencies, duplicates
Feature engineering Creating useful model input variables from data that is already correct
Hyperparameter tuning Searching for good values of the settings that govern how learning happens, chosen rather than learned
Evaluation The pass/fail verdict against the success metric, performed once before deployment
Monitoring Continuous observation of live quality, drift and cost after deployment
Data drift Degrading production performance caused by the input distribution shifting away from the training data
Managed API service Calling a model the cloud provider operates; minimum operational burden, less control
Self-hosted API Running a model on infrastructure you operate; maximum control, every operational burden
MLOps The practices that make a model a maintained system rather than a successful experiment
Technical debt (in ML) Accumulated glue code, dead features and untracked datasets; accrues faster than in ordinary software because the data keeps changing
Model re-training Refreshing a model on newer data against a defined trigger and process
Precision The metric protecting against false alarms; optimise when a false positive is expensive
Recall The metric protecting against misses; optimise when a false negative is dangerous
F1 score The metric for when both error types matter and the classes are imbalanced
Business metric A measure of whether the project was worth building — cost per user, development costs, customer feedback, ROI

Revision Flashcards

Say the answer aloud before revealing it.

1. What does EDA produce, and how does that separate it from pre-processing? → Understanding — distributions, gaps and outliers. EDA looks; pre-processing changes.

2. What separates pre-processing from feature engineering? → Pre-processing fixes what is wrong. Feature engineering creates what is useful from data that is already correct.

3. What separates training from hyperparameter tuning? → Training learns the model's parameters from data. Tuning sets the knobs that govern how learning happens; they are chosen, not learned.

4. What separates evaluation from monitoring? → Evaluation happens once, before deployment. Monitoring happens continuously, after it.

5. Name the two feedback loops and what triggers each. → Loop 1: evaluation back to model development, triggered by failing the success metric. Loop 2: monitoring back to data preparation, triggered by data drift.

6. Give the three-part evidence signature of data drift. → Code and model unchanged, error growing over time, and the input distribution has shifted.

7. How does drift differ from overfitting? → By time signature. Overfitting is bad from day one. Drift is good performance that decays.

8. Which five services does Objective 1.3.4 name explicitly? → Amazon Bedrock, Amazon Q, Amazon Quick, Kiro, and Amazon SageMaker AI.

9. What do Amazon Q and Kiro do in a pipeline context? → Assisted development — they help the people building the pipeline, not the data flowing through it.

10. When is training a custom model the right answer? → Only when nothing pre-trained fits and the scenario supplies both labeled data and the skills to train.

11. What is the only axis separating a managed API from a self-hosted one? → Operational burden versus control. Self-hosting requires a stated reason, not a preference.

12. Name the seven MLOps concepts. → Experimentation, repeatable processes, scalable systems, managing technical debt, achieving production readiness, model monitoring, and model re-training.

The Four-Beat Answer

The core question this chapter prepares you for: "A model that worked is now failing in production. How do you diagnose and fix it?"

Four beats, checked in this order. Missing a beat is a failure state.

  1. Locate the stage — say where in the lifecycle the symptom lives. A live production symptom is monitoring, not evaluation, and naming that first rules out half the plausible fixes.
  2. Read the time signature — was it wrong from day one, or did it decay? Bad from the start points at the model or its training. Decay points at the data.
  3. Check what changed — code, model artefact, or input distribution. Unchanged code plus shifted inputs plus growing error is drift, and drift is a data problem, not an algorithm problem.
  4. Name the loop and the fix — loop 2, back to data preparation, re-train on data reflecting the new reality, and add a re-training trigger so the next occurrence is caught by monitoring rather than by a customer.

A strong answer resists proposing a new algorithm. That is the move the question is testing for.

Why This Helps You

On the job: the most common failure in deployed ML is not a bad model, it is an unmonitored one. Teams ship at 94%, celebrate, and discover eighteen months later that nobody defined a re-training trigger. Being the person who asks "what fires loop 2, and who sees it?" before launch is worth more than a point of accuracy.

In interviews: "your model's performance is degrading in production — walk me through it" is a standard question, and reaching immediately for a different algorithm is the answer that ends the conversation. The four beats above are what an interviewer is checking for.

On the exam: Task 1.3 questions are symptom-shaped, and every distractor is a real activity performed at the wrong point in the lifecycle. Locating the stage before choosing the fix is what keeps you from picking a competent answer to the wrong problem.

Chapter Checklist

  • I can sequence the pipeline and say what each stage produces
  • I can separate EDA from pre-processing, pre-processing from feature engineering, training from tuning, and evaluation from monitoring
  • I can draw both feedback loops and name what triggers each
  • I can give the three-part evidence signature of data drift
  • I can distinguish drift from overfitting by time signature
  • I can attach Bedrock, Amazon Q, Amazon Quick, Kiro and SageMaker AI to their pipeline stages
  • I can recognise a real service offered for the wrong stage
  • I can choose a model source from speed versus control, and say what custom training requires
  • I can choose managed or self-hosted, and say why a preference is not a constraint
  • I can name all seven MLOps concepts and say what MLOps adds beyond "the model works"
  • I can choose accuracy, precision, recall, F1 or a business metric from the cost of being wrong
  • I can explain why accuracy is misleading on imbalanced data, with a worked number

After the Chapter

  1. Complete student/project.md — parts 11, 12, 13 and 14 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet. These parts complete its Domain 1 coverage.
  2. Take student/quiz.md closed-book, then review the reasoning for every question you guessed, including the ones you got right.
  3. Open the official v1.1 exam guide's Domain 1 page and confirm you can attach a concept from this chapter to each of the six bullets under Task 1.3. Read the final bullet carefully and check it against any other study material you own — if a third-party source disagrees with the published guide, the published guide wins.
  4. Next chapter: Chapter 04 — Foundation Models, Tokens, Embeddings, and the FM Lifecycle (Domain 2, 24%). Domain 2 is where v1.1 changed most, and the vocabulary load is higher than anything in Domain 1: tokens, chunking, embeddings, vectors, transformer-based LLMs, multi-modal and diffusion models, the GenAI use-case families, and the FM lifecycle from data selection through feedback.

Chapter 3 quiz

13 questions on this chapter, marked instantly, with an explanation for every answer.

1. A model was deployed eleven months ago and performed well. Over the last two months its error rate has risen steadily. The code, the deployed model artefact and the inference configuration are all unchanged. Analysis shows the distribution of incoming data has shifted. What does this describe, and what is the correct response?
2. A team removes duplicate rows, fills missing postcodes and standardises inconsistent date formats. Which pipeline stage is this?
3. A different team, working on the same already-cleaned dataset, derives a new column: days elapsed between a customer's last two orders. Which pipeline stage is this?
4. Which statement correctly separates evaluation from monitoring?
5. A business intelligence team needs a dashboard summarising model outputs for executives. Which AWS service fits, and which pipeline stage does it serve?
6. A startup with three engineers and no ML operations staff wants a foundation model available behind an API as quickly as possible. There is no stated regulatory or data-residency requirement. Which serving method fits?
7. A healthcare provider's contract states that model weights may never leave infrastructure the provider operates. Which serving method is required, and why?
8. Which of the following is NOT one of the seven MLOps concepts named in Objective 1.3.5?
9. A fraud model operates where fraud is 0.2% of transactions. A vendor reports the model achieves 99.8% accuracy. What should the team conclude?
10. A screening model identifies patients who should receive a follow-up test. A missed case can be fatal; an unnecessary follow-up is inexpensive and safe. Which metric should be optimised?
11. After eighteen months, a company's leadership asks whether its ML programme has been worth the investment. Which metric answers that question?
12. A team wants a model for a task that a widely available pre-trained model already performs well. The team has no labeled data of its own and no ML training expertise. Which model source fits?
13. A team cannot reproduce the model it trained last quarter. The training data was overwritten, the notebook was edited since, and nobody recorded which hyperparameter values were used. Which MLOps concepts does this failure most directly demonstrate? (Select two.)