Chapter 12 · Domain 3 · 28% of the exam
Evaluating Foundation Models, and Proving Business Value
29 min read · Chapter 12 of 17
On this page
- Certification Blueprint
- What This Chapter Covers
- Approaches to Evaluating an FM
- The Metrics
- Evaluating Applications, Not Just Models
- Does It Meet the Business Objective?
- Decision Rules and Exam Signals
- Distractor Patterns
- Scenario Walkthrough
- Key Concepts
- Revision Flashcards
- The Five-Beat Answer
- Why This Helps You
- Chapter Checklist
- After the Chapter
Certification Blueprint
| Field | Coverage |
|---|---|
| Exam | AWS Certified AI Practitioner (AIF-C01), exam guide v1.1 |
| Domain | Content Domain 3 — Applications of Foundation Models |
| Exam weight | 28% of scored content — the largest domain on the exam |
| Task statement | 3.4 Describe methods to evaluate FM performance |
| Objectives | 3.4.1 approaches to evaluate · 3.4.2 metrics · 3.4.3 meeting business objectives · 3.4.4 evaluating FM applications · 3.4.5 business objective alignment metrics |
| Service families | Amazon Bedrock Model Evaluation |
What This Chapter Covers
Chapter 11 ended with a model. Pre-trained, fine-tuned, distilled — whichever route was taken, the output was an artefact, and the chapter deliberately gave you no way at all to know whether it was any good.
This chapter is that answer, and it completes Domain 3.
Everything in it follows from a single structural fact, and it is worth stating before any metric name appears:
Every metric in this chapter scores a proxy, never the thing you actually care about. ROUGE counts overlap with a reference summary — not usefulness. A benchmark scores a curated dataset — not your traffic. Even a human reviewer scores the sample they were shown. The skill this task statement tests is knowing which proxy you are holding and what it structurally cannot see.
That one fact explains the whole chapter. It is why several different metrics exist rather than one good one. It is why an FM application needs its own evaluation separate from the model inside it. And it is why Objectives 3.4.3 and 3.4.5 exist at all — because a number that measures a proxy can look excellent while the deployment it describes has changed nothing.
Hold that sentence. The rest of the chapter is its detail.
What to remember from this diagram: there are three layers, not one, and each is blind to the one below it in importance. A model score cannot see your users. An application score cannot see whether the task was worth doing. Only the bottom layer answers the question that authorised the budget. Most exam questions in this task statement are really asking which layer the scenario is about.
Approaches to Evaluating an FM
Objective 3.4.1 names three approaches: human-in-the-loop evaluation, benchmark datasets, and Amazon Bedrock Model Evaluation. They are not three competing products — the first two are methods, and the third is the AWS service that runs both.
Human-in-the-loop evaluation
People read model outputs and judge them.
| Property | Detail |
|---|---|
| What it gives you | The only genuine ground truth for subjective quality — tone, helpfulness, safety, nuance, whether an answer is actually useful |
| What it costs | Slow and expensive, and it does not scale with traffic |
| What it is blind to | Whatever was not in the sample. Reviewers see a slice, and a rare failure mode can sit entirely outside it |
Use it when the quality question is a judgement, not a comparison against a known correct answer. "Is this refusal appropriately worded?" has no reference string to match against.
Benchmark datasets
A standardised set of tasks with known good answers, run against a model to produce a comparable score.
| Property | Detail |
|---|---|
| What it gives you | Repeatability and comparability. Two models measured the same way can be ranked |
| What it costs | Little, once it exists — this is the cheap, automatable end |
| What it is blind to | Your workload. A benchmark is somebody else's data. Strong benchmark performance is evidence about the benchmark |
⚠️ The most examinable limitation on this page. A benchmark tells you how a model does on the benchmark. It does not tell you how it will do on your documents, your customers' phrasing, or your edge cases. A scenario in which a model "scores well on standard benchmarks but underperforms in production" is not describing a contradiction — it is describing exactly what benchmarks are for and what they are not.
Amazon Bedrock Model Evaluation
The named AWS service. It matters that it runs both kinds of evaluation:
| Mode | What happens |
|---|---|
| Automatic evaluation | Bedrock runs the model against datasets — built-in or your own — and computes metric scores without a human in the path |
| Human evaluation | Bedrock manages a workflow in which people score outputs, using either your own work team or an AWS-managed one |
The point for the exam is that choosing Bedrock Model Evaluation does not choose a method. It is where both methods are run and compared. A distractor that treats it as purely automatic — or purely human — has misread the service.
What to remember from this diagram: the branch point is does a reference answer exist. That single question separates the automatable half of this objective from the half that needs judgement, and it is the question the exam asks in scenario form.
The Metrics
Objective 3.4.2 names four: ROUGE, BLEU, BERTScore, and LLM-as-a-judge. The exam expands the first two, and the expansions are worth knowing because they encode what each one measures:
- ROUGE — Recall-Oriented Understudy for Gisting Evaluation
- BLEU — Bilingual Evaluation Understudy
Read those two names slowly. "Recall-oriented" and "Bilingual" tell you almost everything.
ROUGE
Measures overlap between the generated text and a reference, oriented toward recall: of the material in the reference, how much did the candidate capture?
Built for summarization. A good summary is one that kept the important content, so recall — did you keep it? — is the natural orientation.
BLEU
Measures overlap too, but oriented toward precision: of the material in the candidate, how much of it appears in the reference?
Built for translation. The clue is in the name — Bilingual Evaluation Understudy. A good translation should not introduce content the source did not contain, so precision — is what you produced supported? — is the natural orientation.
ROUGE and BLEU are the confusable pair in this objective, and the exam knows it. Both count n-gram overlap against a reference. The separation to memorise: ROUGE is recall-oriented and its home task is summarization; BLEU is precision-oriented and its home task is translation. If you can only hold one thing, hold the task: summaries → ROUGE, translation → BLEU.
BERTScore
Instead of counting matching words, BERTScore compares embeddings — it scores semantic similarity.
This matters because ROUGE and BLEU share a structural weakness: a paraphrase that is completely correct but uses different words scores badly. "The flight was cancelled" and "They called off the flight" have almost no n-gram overlap and identical meaning. BERTScore is the named metric that addresses exactly this.
Exam signal: a scenario that stresses "correct but worded differently", "paraphrase", or "semantically equivalent" is pointing at BERTScore.
LLM-as-a-judge
A foundation model scores another model's output against stated criteria.
| Property | Detail |
|---|---|
| What it gives you | Human-style judgement at machine scale, on open-ended tasks where no reference answer exists |
| What it costs | Inference cost per judgement, and design effort on the criteria |
| What it is blind to | Its own limitations. The judge is a foundation model, with the same capacity for bias, inconsistency and error as the model under test |
This is the v1.1-era metric most likely to appear in a scenario about open-ended output — assistants, creative generation, multi-turn conversation — where there is nothing to diff against.
What to remember from this diagram: you route by task shape, not by which metric sounds most sophisticated. And every path ends in the same warning — the one this chapter opened with.
What all four are blind to
None of these metrics knows whether the answer was useful, true, or worth producing.
- A summary can score highly on ROUGE and omit the one clause that mattered legally.
- A translation can score highly on BLEU and be unusable in register.
- BERTScore rewards semantic similarity to a reference — including similarity to a reference that was itself wrong.
- An LLM judge can confidently prefer a fluent, incorrect answer over an awkward, correct one.
Veracity is the hardest gap. A confident hallucination is fluent, well-formed, and often close in embedding space to the truth. That is why hallucination detection is treated as its own discipline in Domain 5 (Chapter 16) rather than as a metric here.
Evaluating Applications, Not Just Models
Objective 3.4.4 is the objective most often skipped in study material, and it carries a distinct idea: an FM application is not an FM.
By the time a user sees an answer, the model was one component among several. A RAG system retrieved first. An agent chose tools and sequenced calls. A workflow chained several steps. Any of those can fail while the model performs perfectly on the text it was given.
What to remember from this diagram: the dotted arrows all point backwards and none of them is labelled. An end-to-end score tells you something is wrong and cannot tell you what. That is the entire argument for component-level evaluation.
What to measure in a RAG system
| Component | The question | What a failure looks like |
|---|---|---|
| Retrieval | Did the right chunks come back? | The answer is confidently wrong because the correct passage was never retrieved |
| Generation | Is the answer faithful to what was retrieved? | The right passage was retrieved and the model contradicted or embellished it |
These two demand opposite fixes, and the exam tests exactly that. A retrieval failure is fixed by chunking, embeddings, or the index — changing the model does nothing. A faithfulness failure is fixed by prompting or a different model — improving retrieval does nothing. A team measuring only end-to-end accuracy cannot tell which one they have.
What to measure in agents and workflows
| Component | The question |
|---|---|
| Tool selection | Did it choose the right tool for the step? |
| Tool arguments | Did it call that tool with correct parameters? |
| Orchestration | Did it take the right steps, in a sensible order, and terminate? |
| End-to-end task success | Did the whole thing accomplish what the user asked? |
An agent can select a correct tool and pass it wrong arguments. It can make every individual call correctly and loop without terminating. Step-level correctness and task-level success are different measurements, and a system can pass one while failing the other.
Does It Meet the Business Objective?
Objectives 3.4.3 and 3.4.5 are the chapter's pivot, and they are where the largest number of marks are quietly lost.
Objective 3.4.3 names productivity, user engagement, and task engineering as the categories in which an FM is judged against a business objective. Objective 3.4.5 names three concrete alignment metrics: task completion rate, user satisfaction, and cost per interaction.
What to remember from this diagram: both columns contain real numbers. That is what makes the distractor work. The left column is true and irrelevant to the question being asked.
The three named alignment metrics
| Metric | What it measures | What it catches that the others miss |
|---|---|---|
| Task completion rate | The proportion of user attempts that actually reached the intended outcome | A system that produces beautiful answers nobody can act on. It is behavioural, not aesthetic |
| User satisfaction | Whether the people using it find it valuable — surveys, ratings, thumbs, repeat use | A system that technically completes tasks while being unpleasant, slow, or untrustworthy to use |
| Cost per interaction | What one interaction costs to serve | The success that is not worth having. A model that completes more tasks at four times the cost may have made things worse |
Cost per interaction is the one most often omitted, and the exam rewards naming it. It is the only alignment metric that can turn an apparent success into a failure. Chapter 05's token-based pricing is what makes it move, and Chapter 09's cost ladder is what it feeds back into.
Productivity, user engagement, and task engineering
Objective 3.4.3's three named categories are broader lenses:
- Productivity — is the work getting done faster or with less effort? Handling time, throughput, volume deflected from a human queue.
- User engagement — are people choosing to use it? Adoption, repeat use, abandonment. A tool that works and that nobody opens has not met the objective.
- Task engineering — whether the task itself was well framed for a model. Sometimes the honest evaluation finding is that the task was the wrong shape, and no model would have met the objective.
Task engineering is the subtle one. It admits the possibility that the failure is not the model's. A scenario where every technical metric is fine and the outcome is still poor may be describing a task that was never suited to an FM — which is the same judgement Chapter 02 taught under Objective 1.2.2.
Decision Rules and Exam Signals
Rule 1 — every metric scores a proxy. Ask what the number is standing in for, and what it cannot see. This is the chapter in one line.
Rule 2 — route by whether a reference answer exists. It exists → benchmark or overlap metric. It does not → human-in-the-loop or LLM-as-a-judge.
Rule 3 — summaries are ROUGE, translation is BLEU. Recall-oriented versus precision-oriented; if the orientation slips, the task association will still carry you.
Rule 4 — "correct but different words" means BERTScore. Overlap metrics punish paraphrase; embeddings do not.
Rule 5 — open-ended with no reference means LLM-as-a-judge, and the judge inherits the limitations of a foundation model.
Rule 6 — Bedrock Model Evaluation runs both automatic and human workflows. Choosing the service is not choosing the method.
Rule 7 — benchmark performance is evidence about the benchmark. "Scores well, underperforms in production" is the expected behaviour, not a paradox.
Rule 8 — an application is not a model. Evaluate retrieval, generation, tool use and orchestration as separate components, because an end-to-end score cannot localise a fault.
Rule 9 — retrieval failure and generation failure demand opposite fixes. Was the right passage retrieved? If no, the fix is the index. If yes, the fix is the prompt or the model.
Rule 10 — a technical score is never business evidence. Task completion rate, user satisfaction and cost per interaction answer the business question; ROUGE never does.
Rule 11 — name cost per interaction. It is the alignment metric most often left out and the only one that can convert a success into a failure.
Rule 12 — traditional ML metrics are a different objective. Accuracy, precision, recall and F1 belong to Objective 1.3.6 and Chapter 03. They need a labelled test set and one right answer, which generated text does not have.
Distractor Patterns
| Pattern | What it looks like | How to defuse it |
|---|---|---|
| Benchmark score as business proof | A higher benchmark result offered as evidence the deployment succeeded | Different layer. Ask which of the three layers the stem is asking about |
| BLEU for summarization | The precision-oriented translation metric applied to a summary task | Match the task shape: summaries are ROUGE |
| ROUGE for translation | The mirror error | Bilingual Evaluation Understudy — the name carries the task |
| Overlap metric for paraphrase | ROUGE or BLEU offered where the stem stresses different wording, same meaning | Overlap metrics punish paraphrase by construction. BERTScore scores meaning |
| Accuracy / precision / recall / F1 | Classical ML metrics offered for generated text | No confusion matrix exists for a paragraph. Those are Objective 1.3.6, Chapter 03 |
| "Bedrock Model Evaluation is automatic only" | The service described as excluding human review | It manages both automatic and human evaluation workflows |
| More human review to fix a scale problem | Human-in-the-loop offered where the stem stresses volume | It does not scale with traffic; that is what LLM-as-a-judge addresses |
| End-to-end score to localise a RAG fault | Overall accuracy offered as the way to find which stage failed | It cannot. Component-level measurement is the named approach |
| Swap the model to fix retrieval | A better FM offered where the correct passage was never retrieved | The model never saw it. The fix is chunking, embeddings or the index |
| Improve retrieval to fix unfaithfulness | Index tuning offered where the right passage was retrieved and contradicted | Retrieval already worked. The failure is generation |
| User satisfaction alone as success | Satisfaction offered while the stem stresses spend | Cost per interaction is the metric that catches an expensive success |
| A model can't be evaluated without a labelled dataset | Absence of labels presented as blocking evaluation | Human-in-the-loop and LLM-as-a-judge exist precisely for the unlabelled case |
The second, third and fourth rows are the most reliable ways to lose marks in Objective 3.4.2, because all three metrics compare a candidate against a reference and the differences are structural rather than obvious.
Scenario Walkthrough
A logistics company deploys an assistant that answers questions about shipping policy from an internal document set. Before launch, the team benchmarked three foundation models and chose the one with the strongest published scores. After launch: answers read well, but customers frequently receive confident statements that contradict the policy documents. Investigation shows that in most failing cases, the passage containing the correct policy was never among the retrieved chunks. Agent handover volume is unchanged from before launch. Finance notes the assistant costs roughly three times per conversation what the team modelled.
| Requirement | Reading | Decision |
|---|---|---|
| Chose the model on published benchmark scores | Evidence about the benchmark, not about this workload | Benchmarks rank models; they do not predict production. Not a wrong step — an incomplete one |
| Confident answers contradicting the documents | Sounds like a model or veracity problem | Do not stop here — the next line reclassifies it |
| Correct passage never retrieved | The model never saw the policy it contradicted | Retrieval failure, not generation. Fix the index, chunking or embeddings — a better FM changes nothing |
| Handover volume unchanged | The business outcome that justified the build has not moved | Task completion rate is flat — the objective is unmet regardless of any technical score |
| Three times the modelled cost per conversation | The spend side of the alignment question | Cost per interaction — the metric that turns "it works" into "it is not worth it" |
Five requirements, and the third one is the test. A candidate who has decided "this is a hallucination question" at line two will reach for grounding, output validation, or a stronger model — all of them Domain 5 answers to a Domain 3 problem. The stem then supplies the disqualifying fact: the passage was never retrieved. The model cannot contradict a document it was never shown.
The fourth and fifth rows are the pivot the whole chapter exists for. Every technical judgement could have been correct and the deployment would still have failed the business objective, because nothing moved in the only two numbers the business was watching.
Key Concepts
| Term | Definition |
|---|---|
| Human-in-the-loop evaluation | People reviewing and scoring model outputs; the authoritative source for subjective quality, limited by cost, speed and sample coverage |
| Benchmark dataset | A standardised task set with known good answers, used to score and compare models repeatably; evidence about the benchmark, not about your workload |
| Amazon Bedrock Model Evaluation | The AWS service for evaluating models, running both automatic metric-based evaluation and managed human evaluation workflows |
| ROUGE | Recall-Oriented Understudy for Gisting Evaluation — recall-oriented overlap against a reference; the metric family for summarization |
| BLEU | Bilingual Evaluation Understudy — precision-oriented overlap against a reference; the metric family for translation |
| BERTScore | A metric comparing embeddings rather than matching words, scoring semantic similarity; the answer when correct output is worded differently from the reference |
| LLM-as-a-judge | Using a foundation model to score another model's output against stated criteria; scales judgement to open-ended tasks, and inherits a model's biases |
| Component-level evaluation | Measuring retrieval, generation, tool selection and orchestration separately rather than only end-to-end, so a failure can be localised |
| Retrieval failure | The correct source passage was never returned; fixed at the index, chunking or embedding layer, never by changing the model |
| Faithfulness failure | The correct passage was retrieved and the generated answer contradicted or embellished it; fixed at the prompt or model layer |
| Task completion rate | The proportion of user attempts that reached the intended outcome; a behavioural business metric |
| User satisfaction | Whether users find the system valuable — ratings, surveys, repeat use, abandonment |
| Cost per interaction | What one interaction costs to serve; the alignment metric that can turn an apparent success into a failure |
| Task engineering | Whether the task itself was well framed for a foundation model; the honest finding that a task may be the wrong shape for any model |
Revision Flashcards
Say the answer aloud before revealing it.
1. State the one structural fact this whole chapter follows from. → Every metric here scores a proxy, never the thing you actually care about. ROUGE counts overlap with a reference, not usefulness; a benchmark scores a curated dataset, not your traffic; a human reviewer scores the sample they were shown. The skill being tested is knowing which proxy you are holding and what it structurally cannot see.
2. Name the three approaches in Objective 3.4.1 and the relationship between them. → Human-in-the-loop evaluation, benchmark datasets, and Amazon Bedrock Model Evaluation. The first two are methods; the third is the AWS service that runs both — automatic metric-based evaluation and managed human evaluation workflows. Choosing the service does not choose the method.
3. What does a strong benchmark score actually tell you? → That the model performs well on the benchmark. A benchmark is somebody else's curated data; it has never seen your documents, your users' phrasing, or your edge cases. A scenario describing a model that benchmarks well and underperforms in production is not a paradox — it is exactly what benchmarks do and do not measure.
4. ROUGE or BLEU for a summarization task, and why? → ROUGE. It is Recall-Oriented Understudy for Gisting Evaluation, and recall is the natural orientation for a summary: of the material in the reference, how much did the candidate keep? BLEU is precision-oriented and its home task is translation — the Bilingual in the name carries it.
5. Output is correct but phrased entirely differently from the reference. Which metric? → BERTScore. ROUGE and BLEU both count n-gram overlap, so a correct paraphrase scores badly by construction — "the flight was cancelled" and "they called off the flight" share almost no words and all of their meaning. BERTScore compares embeddings, scoring semantic similarity instead of matching words.
6. When is LLM-as-a-judge the right approach, and what is its weakness? → When the task is open-ended and no reference answer exists, and human review will not scale to the volume. Its weakness is structural: the judge is itself a foundation model, carrying the same capacity for bias, inconsistency and confident error as the model it is scoring.
7. Why is veracity the hardest thing for these metrics to catch? → Because a confident hallucination is fluent, well-formed, and often close in embedding space to the truth. Overlap and similarity metrics reward exactly those properties. This is why hallucination detection is treated as its own discipline in Domain 5 rather than as a metric in this objective.
8. Why does an FM application need evaluation separate from the FM? → Because by the time a user sees an answer, the model was one component among several — retrieval, generation, tool selection, orchestration. Any of those can fail while the model performs perfectly on the text it was handed. An end-to-end score tells you something is wrong and cannot tell you what.
9. A RAG system answers confidently and wrongly. What is the first thing to check? → Whether the correct passage was retrieved at all. If it was not, this is a retrieval failure and the fix lives in chunking, embeddings or the index — a better foundation model changes nothing, because it never saw the passage. If it was retrieved and contradicted, this is a faithfulness failure and the fix is the prompt or the model. The two demand opposite responses.
10. Name the three business objective alignment metrics in Objective 3.4.5. → Task completion rate, user satisfaction, and cost per interaction. Completion rate is behavioural — did attempts reach the intended outcome. Satisfaction is whether users find it valuable. Cost per interaction is what one interaction costs to serve.
11. Which alignment metric is most often omitted, and why does it matter? → Cost per interaction. It is the only one that can turn an apparent success into a failure: a system that completes more tasks at several times the cost may have made things worse. Token-based pricing is what makes it move, and it feeds directly back into the customization cost ladder.
12. What does "task engineering" admit that the other categories do not? → That the failure may not be the model's. Sometimes the honest evaluation finding is that the task was never well shaped for a foundation model, and no model would have met the objective — the same judgement Chapter 02 taught about recognising when AI is not the appropriate solution.
The Five-Beat Answer
The core question this chapter prepares you for: "How would you know whether a foundation model is good enough — and whether it was worth deploying?"
Five beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.
- Approach — say how you would generate a judgement at all: benchmark datasets where reference answers exist, human-in-the-loop where quality is subjective, LLM-as-a-judge where judgement must scale, all runnable through Amazon Bedrock Model Evaluation.
- Metric — pick from the task shape. ROUGE for summarization, BLEU for translation, BERTScore when correct output is worded differently, LLM-as-a-judge for open-ended work. Say what your chosen metric is blind to, because that is the sentence that shows you understand it.
- Application — state that the model is not the system. Evaluate retrieval, generation, tool use and orchestration separately, and explain that an end-to-end score cannot localise a fault.
- Business — move to the layer that authorised the spend: task completion rate, user satisfaction, cost per interaction. Say explicitly that a benchmark score is not evidence here.
- Judgement — close on the honest possibility: the model may be fine and the task badly framed. Naming task engineering shows you can distinguish a model failure from a problem-selection failure.
A strong answer moves from technical to business evidence and says what each measurement cannot see. A weak answer names ROUGE and stops.
Why This Helps You
On the job: the most expensive pattern in this area is a team that ships on benchmark scores and then cannot explain, six months later, whether the system is working. The second is a RAG deployment where every reported failure is met with "let's try a better model" because nobody ever measured retrieval separately — a fix that cannot work, applied repeatedly, at increasing cost.
In interviews: "how would you evaluate this?" is a standard senior screening question, and most candidates answer with a metric. The strong answer names the layer first, admits what the metric cannot see, and reaches business alignment without being prompted. Being able to say "the model may be fine and the task badly framed" marks out someone who has actually run one of these programmes.
On the exam: Domain 3 is 28% of scored content and this chapter closes it. The highest-value habits are routing a metric from the task shape, localising an application failure to a component, and refusing to accept a technical number as an answer to a business question.
Chapter Checklist
- I can state why every metric in this chapter measures a proxy, and give an example of what one cannot see
- I can name the three approaches in Objective 3.4.1 and say which of them are methods and which is a service
- I can explain what Amazon Bedrock Model Evaluation runs, and why "automatic only" is wrong
- I can say what a strong benchmark score is and is not evidence of
- I can expand ROUGE and BLEU, and attach each to its home task
- I can say why an overlap metric punishes a correct paraphrase, and name the metric that does not
- I can say when LLM-as-a-judge is appropriate and what it inherits from being a model
- I can explain why veracity is the hardest property for these metrics to detect
- I can argue why an FM application needs its own evaluation separate from the model
- I can tell a retrieval failure from a faithfulness failure, and give the opposite fix each demands
- I can name the four things worth measuring in an agent or workflow
- I can name the three business objective alignment metrics and what each catches that the others miss
- I can explain why cost per interaction is the one most often omitted
- I can distinguish these metrics from accuracy, precision, recall and F1, and say which objective those belong to
After the Chapter
- Complete
student/project.md— parts 43-45 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one. - Take
student/quiz.mdclosed-book, then review the reasoning for every question you guessed, including the ones you got right. Pay particular attention to questions 8 and 9 — they are deliberate mirror images, and missing both means the retrieval-versus-generation split has not landed. - Open the official v1.1 exam guide's Domain 3 page and confirm you can attach a concept from this chapter to each of the five bullets under Task Statement 3.4. Note that the fifth bullet is about business objective alignment metrics specifically — read it alongside the third, because the exam treats them as a pair.
- Domain 3 is now complete. Before moving on, check that you can place all four of its task statements: design considerations (Ch 08-09), prompt engineering (Ch 10), training and fine-tuning (Ch 11), and evaluation (this chapter). It is 28% of scored content and the largest single block of marks on the exam.
- Next: Chapter 13 — Responsible AI: Features, Guardrails, and Legal Risk (Domain 4, Task 4.1, objectives 1-4). This chapter measured whether a model is good. Chapter 13 asks a different question — whether it is acceptable — and a model can pass everything here while failing that one.
Chapter 12 quiz
13 questions on this chapter, marked instantly, with an explanation for every answer.