Free live cohort on Google Meet — register your interest →

Chapter 12 · Domain 3 · 28% of the exam

Evaluating Foundation Models, and Proving Business Value

29 min read · Chapter 12 of 17

On this page
  1. Certification Blueprint
  2. What This Chapter Covers
  3. Approaches to Evaluating an FM
  4. The Metrics
  5. Evaluating Applications, Not Just Models
  6. Does It Meet the Business Objective?
  7. Decision Rules and Exam Signals
  8. Distractor Patterns
  9. Scenario Walkthrough
  10. Key Concepts
  11. Revision Flashcards
  12. The Five-Beat Answer
  13. Why This Helps You
  14. Chapter Checklist
  15. After the Chapter

Certification Blueprint

Field Coverage
Exam AWS Certified AI Practitioner (AIF-C01), exam guide v1.1
Domain Content Domain 3 — Applications of Foundation Models
Exam weight 28% of scored content — the largest domain on the exam
Task statement 3.4 Describe methods to evaluate FM performance
Objectives 3.4.1 approaches to evaluate · 3.4.2 metrics · 3.4.3 meeting business objectives · 3.4.4 evaluating FM applications · 3.4.5 business objective alignment metrics
Service families Amazon Bedrock Model Evaluation

What This Chapter Covers

Chapter 11 ended with a model. Pre-trained, fine-tuned, distilled — whichever route was taken, the output was an artefact, and the chapter deliberately gave you no way at all to know whether it was any good.

This chapter is that answer, and it completes Domain 3.

Everything in it follows from a single structural fact, and it is worth stating before any metric name appears:

Every metric in this chapter scores a proxy, never the thing you actually care about. ROUGE counts overlap with a reference summary — not usefulness. A benchmark scores a curated dataset — not your traffic. Even a human reviewer scores the sample they were shown. The skill this task statement tests is knowing which proxy you are holding and what it structurally cannot see.

That one fact explains the whole chapter. It is why several different metrics exist rather than one good one. It is why an FM application needs its own evaluation separate from the model inside it. And it is why Objectives 3.4.3 and 3.4.5 exist at all — because a number that measures a proxy can look excellent while the deployment it describes has changed nothing.

Hold that sentence. The rest of the chapter is its detail.

Three stacked evaluation layers connected by arrows: layer one model-level scoring output quality with ROUGE, BLEU, BERTScore, LLM-as-a-judge and benchmarks; layer two application-level covering retrieval, generation, tool calls and orchestration; layer three business-level covering task completion rate, user satisfaction and cost per interaction, with dotted notes showing each layer's blind spot and the business layer seeing the result that justified the spend

What to remember from this diagram: there are three layers, not one, and each is blind to the one below it in importance. A model score cannot see your users. An application score cannot see whether the task was worth doing. Only the bottom layer answers the question that authorised the budget. Most exam questions in this task statement are really asking which layer the scenario is about.

Approaches to Evaluating an FM

Objective 3.4.1 names three approaches: human-in-the-loop evaluation, benchmark datasets, and Amazon Bedrock Model Evaluation. They are not three competing products — the first two are methods, and the third is the AWS service that runs both.

Human-in-the-loop evaluation

People read model outputs and judge them.

Property Detail
What it gives you The only genuine ground truth for subjective quality — tone, helpfulness, safety, nuance, whether an answer is actually useful
What it costs Slow and expensive, and it does not scale with traffic
What it is blind to Whatever was not in the sample. Reviewers see a slice, and a rare failure mode can sit entirely outside it

Use it when the quality question is a judgement, not a comparison against a known correct answer. "Is this refusal appropriately worded?" has no reference string to match against.

Benchmark datasets

A standardised set of tasks with known good answers, run against a model to produce a comparable score.

Property Detail
What it gives you Repeatability and comparability. Two models measured the same way can be ranked
What it costs Little, once it exists — this is the cheap, automatable end
What it is blind to Your workload. A benchmark is somebody else's data. Strong benchmark performance is evidence about the benchmark

⚠️ The most examinable limitation on this page. A benchmark tells you how a model does on the benchmark. It does not tell you how it will do on your documents, your customers' phrasing, or your edge cases. A scenario in which a model "scores well on standard benchmarks but underperforms in production" is not describing a contradiction — it is describing exactly what benchmarks are for and what they are not.

Amazon Bedrock Model Evaluation

The named AWS service. It matters that it runs both kinds of evaluation:

Mode What happens
Automatic evaluation Bedrock runs the model against datasets — built-in or your own — and computes metric scores without a human in the path
Human evaluation Bedrock manages a workflow in which people score outputs, using either your own work team or an AWS-managed one

The point for the exam is that choosing Bedrock Model Evaluation does not choose a method. It is where both methods are run and compared. A distractor that treats it as purely automatic — or purely human — has misread the service.

A decision flow beginning at evaluating a foundation model, asking whether a correct reference answer exists; a yes branch routes to benchmark datasets for standardised repeatable comparison, a no branch asks whether the judgement is subjective or about tone, safety and nuance, routing to human-in-the-loop as the authoritative ground truth or to LLM-as-a-judge where scale is needed; all three converge on Amazon Bedrock Model Evaluation as the managed place that runs automatic and human workflows

What to remember from this diagram: the branch point is does a reference answer exist. That single question separates the automatable half of this objective from the half that needs judgement, and it is the question the exam asks in scenario form.

The Metrics

Objective 3.4.2 names four: ROUGE, BLEU, BERTScore, and LLM-as-a-judge. The exam expands the first two, and the expansions are worth knowing because they encode what each one measures:

  • ROUGE — Recall-Oriented Understudy for Gisting Evaluation
  • BLEU — Bilingual Evaluation Understudy

Read those two names slowly. "Recall-oriented" and "Bilingual" tell you almost everything.

ROUGE

Measures overlap between the generated text and a reference, oriented toward recall: of the material in the reference, how much did the candidate capture?

Built for summarization. A good summary is one that kept the important content, so recall — did you keep it? — is the natural orientation.

BLEU

Measures overlap too, but oriented toward precision: of the material in the candidate, how much of it appears in the reference?

Built for translation. The clue is in the name — Bilingual Evaluation Understudy. A good translation should not introduce content the source did not contain, so precision — is what you produced supported? — is the natural orientation.

ROUGE and BLEU are the confusable pair in this objective, and the exam knows it. Both count n-gram overlap against a reference. The separation to memorise: ROUGE is recall-oriented and its home task is summarization; BLEU is precision-oriented and its home task is translation. If you can only hold one thing, hold the task: summaries → ROUGE, translation → BLEU.

BERTScore

Instead of counting matching words, BERTScore compares embeddings — it scores semantic similarity.

This matters because ROUGE and BLEU share a structural weakness: a paraphrase that is completely correct but uses different words scores badly. "The flight was cancelled" and "They called off the flight" have almost no n-gram overlap and identical meaning. BERTScore is the named metric that addresses exactly this.

Exam signal: a scenario that stresses "correct but worded differently", "paraphrase", or "semantically equivalent" is pointing at BERTScore.

LLM-as-a-judge

A foundation model scores another model's output against stated criteria.

Property Detail
What it gives you Human-style judgement at machine scale, on open-ended tasks where no reference answer exists
What it costs Inference cost per judgement, and design effort on the criteria
What it is blind to Its own limitations. The judge is a foundation model, with the same capacity for bias, inconsistency and error as the model under test

This is the v1.1-era metric most likely to appear in a scenario about open-ended output — assistants, creative generation, multi-turn conversation — where there is nothing to diff against.

A decision flow from the question what shape is the task, branching to summarization leading to ROUGE described as recall-oriented overlap, translation leading to BLEU described as precision-oriented overlap, paraphrase or semantic match leading to BERTScore described as embedding similarity scoring meaning rather than matching words, and open-ended with no single right answer leading to LLM-as-a-judge; all four converge on a warning that every metric scores a proxy and none knows whether the answer was useful

What to remember from this diagram: you route by task shape, not by which metric sounds most sophisticated. And every path ends in the same warning — the one this chapter opened with.

What all four are blind to

None of these metrics knows whether the answer was useful, true, or worth producing.

  • A summary can score highly on ROUGE and omit the one clause that mattered legally.
  • A translation can score highly on BLEU and be unusable in register.
  • BERTScore rewards semantic similarity to a reference — including similarity to a reference that was itself wrong.
  • An LLM judge can confidently prefer a fluent, incorrect answer over an awkward, correct one.

Veracity is the hardest gap. A confident hallucination is fluent, well-formed, and often close in embedding space to the truth. That is why hallucination detection is treated as its own discipline in Domain 5 (Chapter 16) rather than as a metric here.

Evaluating Applications, Not Just Models

Objective 3.4.4 is the objective most often skipped in study material, and it carries a distinct idea: an FM application is not an FM.

By the time a user sees an answer, the model was one component among several. A RAG system retrieved first. An agent chose tools and sequenced calls. A workflow chained several steps. Any of those can fail while the model performs perfectly on the text it was given.

A component pipeline from user question through retrieval asking whether the right chunks were found, to generation asking whether the answer is faithful to what was retrieved, to tool call asking whether the right tool and arguments were used, to orchestration asking whether the steps ran in the right order and terminated, ending at a final answer; a wrong end-to-end score points back by dotted arrows at all four stages, with a note that it cannot say which stage failed and that retrieval failure and generation failure demand opposite fixes

What to remember from this diagram: the dotted arrows all point backwards and none of them is labelled. An end-to-end score tells you something is wrong and cannot tell you what. That is the entire argument for component-level evaluation.

What to measure in a RAG system

Component The question What a failure looks like
Retrieval Did the right chunks come back? The answer is confidently wrong because the correct passage was never retrieved
Generation Is the answer faithful to what was retrieved? The right passage was retrieved and the model contradicted or embellished it

These two demand opposite fixes, and the exam tests exactly that. A retrieval failure is fixed by chunking, embeddings, or the index — changing the model does nothing. A faithfulness failure is fixed by prompting or a different model — improving retrieval does nothing. A team measuring only end-to-end accuracy cannot tell which one they have.

What to measure in agents and workflows

Component The question
Tool selection Did it choose the right tool for the step?
Tool arguments Did it call that tool with correct parameters?
Orchestration Did it take the right steps, in a sensible order, and terminate?
End-to-end task success Did the whole thing accomplish what the user asked?

An agent can select a correct tool and pass it wrong arguments. It can make every individual call correctly and loop without terminating. Step-level correctness and task-level success are different measurements, and a system can pass one while failing the other.

Does It Meet the Business Objective?

Objectives 3.4.3 and 3.4.5 are the chapter's pivot, and they are where the largest number of marks are quietly lost.

Objective 3.4.3 names productivity, user engagement, and task engineering as the categories in which an FM is judged against a business objective. Objective 3.4.5 names three concrete alignment metrics: task completion rate, user satisfaction, and cost per interaction.

Two evidence columns feeding one question. The technical column shows a high benchmark score, improved ROUGE, and beating the previous model. The business column shows flat task completion rate, falling user satisfaction, and rising cost per interaction. Both feed the question did the model earn its place, answered no, with the note that a benchmark measures a curated dataset and has never seen your traffic, your users, or your bill

What to remember from this diagram: both columns contain real numbers. That is what makes the distractor work. The left column is true and irrelevant to the question being asked.

The three named alignment metrics

Metric What it measures What it catches that the others miss
Task completion rate The proportion of user attempts that actually reached the intended outcome A system that produces beautiful answers nobody can act on. It is behavioural, not aesthetic
User satisfaction Whether the people using it find it valuable — surveys, ratings, thumbs, repeat use A system that technically completes tasks while being unpleasant, slow, or untrustworthy to use
Cost per interaction What one interaction costs to serve The success that is not worth having. A model that completes more tasks at four times the cost may have made things worse

Cost per interaction is the one most often omitted, and the exam rewards naming it. It is the only alignment metric that can turn an apparent success into a failure. Chapter 05's token-based pricing is what makes it move, and Chapter 09's cost ladder is what it feeds back into.

Productivity, user engagement, and task engineering

Objective 3.4.3's three named categories are broader lenses:

  • Productivity — is the work getting done faster or with less effort? Handling time, throughput, volume deflected from a human queue.
  • User engagement — are people choosing to use it? Adoption, repeat use, abandonment. A tool that works and that nobody opens has not met the objective.
  • Task engineering — whether the task itself was well framed for a model. Sometimes the honest evaluation finding is that the task was the wrong shape, and no model would have met the objective.

Task engineering is the subtle one. It admits the possibility that the failure is not the model's. A scenario where every technical metric is fine and the outcome is still poor may be describing a task that was never suited to an FM — which is the same judgement Chapter 02 taught under Objective 1.2.2.

Decision Rules and Exam Signals

Rule 1 — every metric scores a proxy. Ask what the number is standing in for, and what it cannot see. This is the chapter in one line.

Rule 2 — route by whether a reference answer exists. It exists → benchmark or overlap metric. It does not → human-in-the-loop or LLM-as-a-judge.

Rule 3 — summaries are ROUGE, translation is BLEU. Recall-oriented versus precision-oriented; if the orientation slips, the task association will still carry you.

Rule 4 — "correct but different words" means BERTScore. Overlap metrics punish paraphrase; embeddings do not.

Rule 5 — open-ended with no reference means LLM-as-a-judge, and the judge inherits the limitations of a foundation model.

Rule 6 — Bedrock Model Evaluation runs both automatic and human workflows. Choosing the service is not choosing the method.

Rule 7 — benchmark performance is evidence about the benchmark. "Scores well, underperforms in production" is the expected behaviour, not a paradox.

Rule 8 — an application is not a model. Evaluate retrieval, generation, tool use and orchestration as separate components, because an end-to-end score cannot localise a fault.

Rule 9 — retrieval failure and generation failure demand opposite fixes. Was the right passage retrieved? If no, the fix is the index. If yes, the fix is the prompt or the model.

Rule 10 — a technical score is never business evidence. Task completion rate, user satisfaction and cost per interaction answer the business question; ROUGE never does.

Rule 11 — name cost per interaction. It is the alignment metric most often left out and the only one that can convert a success into a failure.

Rule 12 — traditional ML metrics are a different objective. Accuracy, precision, recall and F1 belong to Objective 1.3.6 and Chapter 03. They need a labelled test set and one right answer, which generated text does not have.

Distractor Patterns

Pattern What it looks like How to defuse it
Benchmark score as business proof A higher benchmark result offered as evidence the deployment succeeded Different layer. Ask which of the three layers the stem is asking about
BLEU for summarization The precision-oriented translation metric applied to a summary task Match the task shape: summaries are ROUGE
ROUGE for translation The mirror error Bilingual Evaluation Understudy — the name carries the task
Overlap metric for paraphrase ROUGE or BLEU offered where the stem stresses different wording, same meaning Overlap metrics punish paraphrase by construction. BERTScore scores meaning
Accuracy / precision / recall / F1 Classical ML metrics offered for generated text No confusion matrix exists for a paragraph. Those are Objective 1.3.6, Chapter 03
"Bedrock Model Evaluation is automatic only" The service described as excluding human review It manages both automatic and human evaluation workflows
More human review to fix a scale problem Human-in-the-loop offered where the stem stresses volume It does not scale with traffic; that is what LLM-as-a-judge addresses
End-to-end score to localise a RAG fault Overall accuracy offered as the way to find which stage failed It cannot. Component-level measurement is the named approach
Swap the model to fix retrieval A better FM offered where the correct passage was never retrieved The model never saw it. The fix is chunking, embeddings or the index
Improve retrieval to fix unfaithfulness Index tuning offered where the right passage was retrieved and contradicted Retrieval already worked. The failure is generation
User satisfaction alone as success Satisfaction offered while the stem stresses spend Cost per interaction is the metric that catches an expensive success
A model can't be evaluated without a labelled dataset Absence of labels presented as blocking evaluation Human-in-the-loop and LLM-as-a-judge exist precisely for the unlabelled case

The second, third and fourth rows are the most reliable ways to lose marks in Objective 3.4.2, because all three metrics compare a candidate against a reference and the differences are structural rather than obvious.

Scenario Walkthrough

A logistics company deploys an assistant that answers questions about shipping policy from an internal document set. Before launch, the team benchmarked three foundation models and chose the one with the strongest published scores. After launch: answers read well, but customers frequently receive confident statements that contradict the policy documents. Investigation shows that in most failing cases, the passage containing the correct policy was never among the retrieved chunks. Agent handover volume is unchanged from before launch. Finance notes the assistant costs roughly three times per conversation what the team modelled.

Requirement Reading Decision
Chose the model on published benchmark scores Evidence about the benchmark, not about this workload Benchmarks rank models; they do not predict production. Not a wrong step — an incomplete one
Confident answers contradicting the documents Sounds like a model or veracity problem Do not stop here — the next line reclassifies it
Correct passage never retrieved The model never saw the policy it contradicted Retrieval failure, not generation. Fix the index, chunking or embeddings — a better FM changes nothing
Handover volume unchanged The business outcome that justified the build has not moved Task completion rate is flat — the objective is unmet regardless of any technical score
Three times the modelled cost per conversation The spend side of the alignment question Cost per interaction — the metric that turns "it works" into "it is not worth it"

Five requirements, and the third one is the test. A candidate who has decided "this is a hallucination question" at line two will reach for grounding, output validation, or a stronger model — all of them Domain 5 answers to a Domain 3 problem. The stem then supplies the disqualifying fact: the passage was never retrieved. The model cannot contradict a document it was never shown.

The fourth and fifth rows are the pivot the whole chapter exists for. Every technical judgement could have been correct and the deployment would still have failed the business objective, because nothing moved in the only two numbers the business was watching.

Key Concepts

Term Definition
Human-in-the-loop evaluation People reviewing and scoring model outputs; the authoritative source for subjective quality, limited by cost, speed and sample coverage
Benchmark dataset A standardised task set with known good answers, used to score and compare models repeatably; evidence about the benchmark, not about your workload
Amazon Bedrock Model Evaluation The AWS service for evaluating models, running both automatic metric-based evaluation and managed human evaluation workflows
ROUGE Recall-Oriented Understudy for Gisting Evaluation — recall-oriented overlap against a reference; the metric family for summarization
BLEU Bilingual Evaluation Understudy — precision-oriented overlap against a reference; the metric family for translation
BERTScore A metric comparing embeddings rather than matching words, scoring semantic similarity; the answer when correct output is worded differently from the reference
LLM-as-a-judge Using a foundation model to score another model's output against stated criteria; scales judgement to open-ended tasks, and inherits a model's biases
Component-level evaluation Measuring retrieval, generation, tool selection and orchestration separately rather than only end-to-end, so a failure can be localised
Retrieval failure The correct source passage was never returned; fixed at the index, chunking or embedding layer, never by changing the model
Faithfulness failure The correct passage was retrieved and the generated answer contradicted or embellished it; fixed at the prompt or model layer
Task completion rate The proportion of user attempts that reached the intended outcome; a behavioural business metric
User satisfaction Whether users find the system valuable — ratings, surveys, repeat use, abandonment
Cost per interaction What one interaction costs to serve; the alignment metric that can turn an apparent success into a failure
Task engineering Whether the task itself was well framed for a foundation model; the honest finding that a task may be the wrong shape for any model

Revision Flashcards

Say the answer aloud before revealing it.

1. State the one structural fact this whole chapter follows from. → Every metric here scores a proxy, never the thing you actually care about. ROUGE counts overlap with a reference, not usefulness; a benchmark scores a curated dataset, not your traffic; a human reviewer scores the sample they were shown. The skill being tested is knowing which proxy you are holding and what it structurally cannot see.

2. Name the three approaches in Objective 3.4.1 and the relationship between them. → Human-in-the-loop evaluation, benchmark datasets, and Amazon Bedrock Model Evaluation. The first two are methods; the third is the AWS service that runs both — automatic metric-based evaluation and managed human evaluation workflows. Choosing the service does not choose the method.

3. What does a strong benchmark score actually tell you? → That the model performs well on the benchmark. A benchmark is somebody else's curated data; it has never seen your documents, your users' phrasing, or your edge cases. A scenario describing a model that benchmarks well and underperforms in production is not a paradox — it is exactly what benchmarks do and do not measure.

4. ROUGE or BLEU for a summarization task, and why? → ROUGE. It is Recall-Oriented Understudy for Gisting Evaluation, and recall is the natural orientation for a summary: of the material in the reference, how much did the candidate keep? BLEU is precision-oriented and its home task is translation — the Bilingual in the name carries it.

5. Output is correct but phrased entirely differently from the reference. Which metric? → BERTScore. ROUGE and BLEU both count n-gram overlap, so a correct paraphrase scores badly by construction — "the flight was cancelled" and "they called off the flight" share almost no words and all of their meaning. BERTScore compares embeddings, scoring semantic similarity instead of matching words.

6. When is LLM-as-a-judge the right approach, and what is its weakness? → When the task is open-ended and no reference answer exists, and human review will not scale to the volume. Its weakness is structural: the judge is itself a foundation model, carrying the same capacity for bias, inconsistency and confident error as the model it is scoring.

7. Why is veracity the hardest thing for these metrics to catch? → Because a confident hallucination is fluent, well-formed, and often close in embedding space to the truth. Overlap and similarity metrics reward exactly those properties. This is why hallucination detection is treated as its own discipline in Domain 5 rather than as a metric in this objective.

8. Why does an FM application need evaluation separate from the FM? → Because by the time a user sees an answer, the model was one component among several — retrieval, generation, tool selection, orchestration. Any of those can fail while the model performs perfectly on the text it was handed. An end-to-end score tells you something is wrong and cannot tell you what.

9. A RAG system answers confidently and wrongly. What is the first thing to check? → Whether the correct passage was retrieved at all. If it was not, this is a retrieval failure and the fix lives in chunking, embeddings or the index — a better foundation model changes nothing, because it never saw the passage. If it was retrieved and contradicted, this is a faithfulness failure and the fix is the prompt or the model. The two demand opposite responses.

10. Name the three business objective alignment metrics in Objective 3.4.5. → Task completion rate, user satisfaction, and cost per interaction. Completion rate is behavioural — did attempts reach the intended outcome. Satisfaction is whether users find it valuable. Cost per interaction is what one interaction costs to serve.

11. Which alignment metric is most often omitted, and why does it matter? → Cost per interaction. It is the only one that can turn an apparent success into a failure: a system that completes more tasks at several times the cost may have made things worse. Token-based pricing is what makes it move, and it feeds directly back into the customization cost ladder.

12. What does "task engineering" admit that the other categories do not? → That the failure may not be the model's. Sometimes the honest evaluation finding is that the task was never well shaped for a foundation model, and no model would have met the objective — the same judgement Chapter 02 taught about recognising when AI is not the appropriate solution.

The Five-Beat Answer

The core question this chapter prepares you for: "How would you know whether a foundation model is good enough — and whether it was worth deploying?"

Five beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.

  1. Approach — say how you would generate a judgement at all: benchmark datasets where reference answers exist, human-in-the-loop where quality is subjective, LLM-as-a-judge where judgement must scale, all runnable through Amazon Bedrock Model Evaluation.
  2. Metric — pick from the task shape. ROUGE for summarization, BLEU for translation, BERTScore when correct output is worded differently, LLM-as-a-judge for open-ended work. Say what your chosen metric is blind to, because that is the sentence that shows you understand it.
  3. Application — state that the model is not the system. Evaluate retrieval, generation, tool use and orchestration separately, and explain that an end-to-end score cannot localise a fault.
  4. Business — move to the layer that authorised the spend: task completion rate, user satisfaction, cost per interaction. Say explicitly that a benchmark score is not evidence here.
  5. Judgement — close on the honest possibility: the model may be fine and the task badly framed. Naming task engineering shows you can distinguish a model failure from a problem-selection failure.

A strong answer moves from technical to business evidence and says what each measurement cannot see. A weak answer names ROUGE and stops.

Why This Helps You

On the job: the most expensive pattern in this area is a team that ships on benchmark scores and then cannot explain, six months later, whether the system is working. The second is a RAG deployment where every reported failure is met with "let's try a better model" because nobody ever measured retrieval separately — a fix that cannot work, applied repeatedly, at increasing cost.

In interviews: "how would you evaluate this?" is a standard senior screening question, and most candidates answer with a metric. The strong answer names the layer first, admits what the metric cannot see, and reaches business alignment without being prompted. Being able to say "the model may be fine and the task badly framed" marks out someone who has actually run one of these programmes.

On the exam: Domain 3 is 28% of scored content and this chapter closes it. The highest-value habits are routing a metric from the task shape, localising an application failure to a component, and refusing to accept a technical number as an answer to a business question.

Chapter Checklist

  • I can state why every metric in this chapter measures a proxy, and give an example of what one cannot see
  • I can name the three approaches in Objective 3.4.1 and say which of them are methods and which is a service
  • I can explain what Amazon Bedrock Model Evaluation runs, and why "automatic only" is wrong
  • I can say what a strong benchmark score is and is not evidence of
  • I can expand ROUGE and BLEU, and attach each to its home task
  • I can say why an overlap metric punishes a correct paraphrase, and name the metric that does not
  • I can say when LLM-as-a-judge is appropriate and what it inherits from being a model
  • I can explain why veracity is the hardest property for these metrics to detect
  • I can argue why an FM application needs its own evaluation separate from the model
  • I can tell a retrieval failure from a faithfulness failure, and give the opposite fix each demands
  • I can name the four things worth measuring in an agent or workflow
  • I can name the three business objective alignment metrics and what each catches that the others miss
  • I can explain why cost per interaction is the one most often omitted
  • I can distinguish these metrics from accuracy, precision, recall and F1, and say which objective those belong to

After the Chapter

  1. Complete student/project.md — parts 43-45 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one.
  2. Take student/quiz.md closed-book, then review the reasoning for every question you guessed, including the ones you got right. Pay particular attention to questions 8 and 9 — they are deliberate mirror images, and missing both means the retrieval-versus-generation split has not landed.
  3. Open the official v1.1 exam guide's Domain 3 page and confirm you can attach a concept from this chapter to each of the five bullets under Task Statement 3.4. Note that the fifth bullet is about business objective alignment metrics specifically — read it alongside the third, because the exam treats them as a pair.
  4. Domain 3 is now complete. Before moving on, check that you can place all four of its task statements: design considerations (Ch 08-09), prompt engineering (Ch 10), training and fine-tuning (Ch 11), and evaluation (this chapter). It is 28% of scored content and the largest single block of marks on the exam.
  5. Next: Chapter 13 — Responsible AI: Features, Guardrails, and Legal Risk (Domain 4, Task 4.1, objectives 1-4). This chapter measured whether a model is good. Chapter 13 asks a different question — whether it is acceptable — and a model can pass everything here while failing that one.

Chapter 12 quiz

13 questions on this chapter, marked instantly, with an explanation for every answer.

1. A team must judge whether an assistant's refusals are worded appropriately. No correct reference answer exists — the question is one of tone and judgement. Which approach does Objective 3.4.1 name for this?
2. Which statement about Amazon Bedrock Model Evaluation is correct?
3. A model scores near the top of published benchmarks. Deployed against the company's own support transcripts, it performs noticeably worse. What does this indicate?
4. A team scores generated meeting summaries against reference summaries written by staff. Which metric family was built for this task shape?
5. A team is scoring machine-translated product descriptions against professional human translations. Which metric was built for this task, and what is its orientation?
6. A model's answers are judged correct by reviewers but score poorly on the team's automated metric, because the wording differs substantially from the reference text. Which metric addresses this?
7. An assistant produces open-ended conversational replies. There is no reference answer, and the volume is far beyond what reviewers can read. Which approach fits, and what is its main weakness?
8. A RAG assistant answers a policy question confidently and incorrectly. Investigation shows the document chunk containing the correct policy was never returned by the retriever. What is the appropriate fix?
9. A different RAG assistant returns an answer that contradicts the source passage. Here the correct chunk **was** retrieved and supplied to the model. Which component has failed?
10. An agent selects the correct tool at each step and passes valid arguments every time, yet users report that many sessions never produce a final result. Which measurement would expose this?
11. **Multiple response — select TWO.** Which two of the following are named in Objective 3.4.5 as **business objective alignment** metrics?
12. A new model completes noticeably more user tasks than the one it replaced. Finance reports that each conversation now costs several times more to serve. Which alignment metric names this problem?
13. Why are accuracy, precision, recall and F1 not the metrics named in Objective 3.4.2 for scoring generated text?