Chapter 14 · Domain 4 · 14% of the exam
Bias, Variance, Datasets, and the Tools That Detect Them
32 min read · Chapter 14 of 17
On this page
- Certification Blueprint
- What This Chapter Covers
- Characteristics of Datasets
- Effects of Bias and Variance
- Tools to Detect and Monitor
- Decision Rules and Exam Signals
- Distractor Patterns
- Scenario Walkthrough
- Key Concepts
- Revision Flashcards
- The Five-Beat Answer
- Why This Helps You
- Chapter Checklist
- After the Chapter
Certification Blueprint
| Field | Coverage |
|---|---|
| Exam | AWS Certified AI Practitioner (AIF-C01), exam guide v1.1 |
| Domain | Content Domain 4 — Guidelines for Responsible AI |
| Exam weight | 14% of scored content |
| Task statement | 4.1 Explain the development of AI systems that are responsible |
| Objectives | 4.1.5 characteristics of datasets · 4.1.6 effects of bias and variance · 4.1.7 tools to detect and monitor bias, trustworthiness and truthfulness |
What This Chapter Covers
Chapter 13 gave you responsible AI as a set of commitments. Fairness, inclusivity, robustness, safety, veracity — properties a system should have, and the legal exposure that follows when it does not. Every one of them stated as a principle.
Principles do not fail loudly. That is the problem this chapter solves.
There is one observation underneath all three objectives here, and it is the reason this material is examinable at all:
A model can be accurate overall and unacceptable in particular.
An aggregate metric is an average, and an average is a summary that discards exactly the information a fairness question is asking about. Nothing in a 94% accuracy figure is false. It simply cannot answer the question "for whom?" — and every scenario in this task statement is a version of that question.
What to remember from this diagram: the aggregate is not an error. Nobody miscalculated. The 94% is arithmetically correct and completely uninformative about the 61% hiding inside it. This is why "the model is accurate" is never a sufficient answer to a fairness question, and why the exam builds so many scenarios that open with a healthy headline number.
The three objectives are one causal chain:
| Objective | Its place in the chain |
|---|---|
| 4.1.5 Dataset characteristics | The cause — what property of the data produces the disparity |
| 4.1.6 Effects of bias and variance | The effect — what you observe when it happens |
| 4.1.7 Detection tools | The instrument — what makes the effect visible |
Study them in that order and each one explains the next. Study them as three unrelated lists and you will be memorising eleven service names with nothing to attach them to.
Characteristics of Datasets
Objective 4.1.5 names four properties: inclusivity, diversity, curated data sources, balanced datasets. They are commonly treated as four words for "good data." They are not — they fail differently, they are detected differently, and the exam separates them.
What to remember from this diagram: the questions are in dependency order, and that ordering is the exam-relevant part. Balance is meaningless if a group is absent — you cannot be disproportionate about zero. Check presence before proportion.
| Characteristic | The question it asks | The failure when it is missing |
|---|---|---|
| Inclusivity | Are the groups the system will serve present at all? | A group absent from training data cannot be served well, and cannot even be measured — it has no rows to compute a metric over |
| Diversity | Does the data span the real range of cases, conditions and contexts? | The model handles the typical case and fails the unusual one; robustness collapses at the edges |
| Balance | Are the groups present in workable proportion? | A group present at 2% is present and swamped; the model minimises total error by optimising the other 98% |
| Curated sources | Is the origin known, selected and documented? | Provenance is unknown, so skew cannot be checked at all — this is the failure that hides the other three |
Inclusivity and balance are different failures
This is the distinction the exam tests most often, and it is easy to blur.
- Inclusivity is a presence question. Is the group in the data?
- Balance is a proportion question. Given that it is in the data, is there enough of it to matter to the training objective?
A dataset containing 40,000 records from one demographic group and 300 from another is inclusive and unbalanced. The second group is present, so it is not an inclusivity failure — but 300 rows against 40,000 will barely move a loss function, so the model optimises the majority almost entirely.
The practical consequence differs too. An inclusivity failure cannot be measured, because there is nothing to measure. An imbalance failure can be measured — you can compute the minority group's accuracy — which is precisely why it is the one subgroup analysis is built to find.
Curated does not mean neutral
"Curated data sources" reads like a quality guarantee and it is not one. Curation means the origin was selected and documented rather than accumulated by accident. A carefully curated dataset drawn entirely from one country is well-curated and thoroughly skewed.
What curation buys you is the ability to answer the question at all. With documented provenance, you can look at a dataset and say what it is a sample of. Without it, you cannot check for skew, which is worse than having skew you have identified — an unmeasurable dataset fails silently.
This is why curation appears in a list about fairness rather than in a list about data engineering. It is not a cleanliness property. It is the precondition for every other check on the list.
Where bias actually enters
Bias is rarely introduced in one place, and the exam likes scenarios where the obvious suspect is not the source.
What to remember from this diagram: the dashed arrow going backwards is the one that turns a one-off problem into a compounding one. A model that serves a group poorly generates poor engagement data for that group, which becomes training data that represents them even more weakly. Skew that is never corrected is not static — it deepens on every retraining cycle.
Note also the fourth stage. Deployment can introduce disparity in a model whose training data was fine, simply because the population that arrives differs from the population that was sampled. That is a scenario the exam uses, and it is the one where the answer is a monitoring tool rather than a data fix.
Effects of Bias and Variance
Objective 4.1.6 lists four things in one bullet: effects on demographic groups, inaccuracy, overfitting, underfitting. That grouping is the objective's most confusing feature, and it is deliberate rather than sloppy.
⚠️ The word "bias" does two different jobs in this objective, and the guide uses both.
Sense Means Shows up as Statistical bias The model is too simple to capture the real pattern Underfitting — wrong on training data and new data Societal bias Outcomes differ systematically across demographic groups Group-level harm — accurate overall, inaccurate for some These are not two names for one thing. A model can have low statistical bias — fitting its training data beautifully — while producing severe societal bias, because it fit training data that was itself skewed. Reading "bias" in the wrong sense is the most reliable way to lose marks in this objective, and a question's other words tell you which is meant: overfitting/underfitting and training accuracy signal the statistical sense; demographic groups and fairness signal the societal one.
Read the symptom, not the definition
The exam does not ask "what is overfitting?" It describes a model's behaviour and asks what is happening. Two numbers decide it.
What to remember from this diagram: the last branch is the one candidates never reach. Strong training accuracy and strong test accuracy is where most people stop and declare the model healthy. This objective exists because that is exactly the state in which a demographic disparity survives undetected — both headline numbers are good, and neither of them was ever asked "for whom?"
| Condition | Training accuracy | Unseen-data accuracy | What it means |
|---|---|---|---|
| Underfitting (high bias) | Poor | Poor | Model too simple; it never learned the pattern |
| Overfitting (high variance) | Excellent | Poor | Model memorised the training set instead of generalising |
| Healthy in aggregate | Good | Good | Says nothing about distribution across groups |
| Group disparity | Good | Good overall | One group's accuracy is far below the average that hides it |
The single most useful reading habit for this objective: when a question gives you two accuracy figures, it is asking about overfitting or underfitting. When it gives you one figure and mentions a group, it is asking about fairness. When it gives you one figure and no group, it is usually testing whether you notice the question cannot be answered from it.
Overfitting and underfitting are not fairness synonyms
Worth stating flatly, because the shared bullet invites the confusion.
- A model that underfits is bad for everyone, roughly equally. That is not primarily a fairness problem; it is a capability problem.
- A model that overfits has memorised its training set. If that training set was skewed, the overfitting entrenches the skew — but overfitting is not itself the disparity.
- A model that works well on average and fails one group may be neither overfitting nor underfitting. It can be a perfectly well-fitted model that learned a real pattern in unrepresentative data.
That last row is the one worth holding. A fair-looking fit and an unfair outcome are compatible, which is why fitting diagnostics do not substitute for subgroup analysis.
Inaccuracy is the shared observable
Why does the guide put all four in one bullet? Because from the outside they present identically — the model gave a wrong answer. Underfitting, overfitting and demographic disparity all surface as inaccuracy, and the diagnostic work is deciding which one produced it. The bullet groups them by symptom, and this chapter's job is to separate them by cause.
Tools to Detect and Monitor
Objective 4.1.7 names six things. They are not six alternatives to choose between — they answer different questions at different points in time, and the timing is what selects them.
What to remember from this diagram: the first question is when, not what. Clarify, Model Monitor and A2I all "detect problems with a model," and a scenario that describes the problem without describing the timing has not given you enough to choose. Find the time signal — before launch, continuously, on each prediction — and the tool follows.
| Tool | What it does | The scenario that names it |
|---|---|---|
| Analyzing label quality | Checks whether the labels themselves are correct and consistently applied | Bias entered through the labelling process — different reviewers judging alike cases differently |
| Human audits | People review model behaviour and decisions directly | Judgement is needed that no metric encodes; also the check on the checkers |
| Subgroup analysis | Computes metrics per group rather than overall | The aggregate looks healthy and a disparity is suspected — the technique, not a product |
| Amazon SageMaker Clarify | Measures bias in data before training and in the model after training; also produces feature attributions | A point-in-time measurement is needed, before deployment |
| SageMaker Model Monitor | Watches a deployed model continuously for drift in data, quality and bias | The model was acceptable at launch and you need to know it stays that way |
| Amazon Augmented AI (A2I) | Routes individual predictions to human reviewers, typically low-confidence ones | Human judgement is needed at inference time, at scale, on specific cases |
The three that get confused: Clarify, Model Monitor, A2I
All three involve checking a model. They sit at three different points in time, and that is the whole distinction.
| Clarify | Model Monitor | A2I | |
|---|---|---|---|
| When it runs | Before training, and after training — point in time | After deployment — continuously | At inference — per prediction |
| What it examines | A dataset, or a trained model | A deployed endpoint over time | One individual prediction |
| What it produces | Bias metrics and feature attributions | Alerts on drift from a baseline | A human decision |
| Answers the question | "Is this data or model biased now?" | "Has it changed since launch?" | "Is this one right?" |
The tell in a question stem is a time word. "Before we deploy" or "assess the training data" is Clarify. "Since launch" or "over the past quarter" is Model Monitor. "Each application" or "flag for review" is A2I.
⚠️ SageMaker Clarify appears in two different objectives, and this trips people. It is named in Objective 4.1.7 as a bias-detection tool, and again in Objective 4.2.2 as a tool for identifying transparent and explainable models. Both are correct. Clarify does two jobs: it computes bias metrics (this chapter) and feature attributions explaining which inputs drove a prediction (Chapter 15). A question naming Clarify is telling you the service; the surrounding words tell you which job.
Subgroup analysis is a technique, not a product
Of the six, subgroup analysis is the only one that is not a service or an activity you schedule. It is a way of computing any metric: instead of one number for the population, one number per group.
It matters disproportionately because it is what makes the disparity visible in the first place. Every service on this list either performs subgroup analysis or acts on its output. If you take one habit from this objective, it is: report per group, never in aggregate.
Human audits are not a lesser tool
There is a temptation to read the automated tools as the real answer and human audits as the old-fashioned fallback. The exam does not treat them that way, and neither should you.
A metric can only measure disparity it was configured to look for. Choosing the groups to compare is itself a human judgement, and a model can be provably fair across every attribute somebody thought to test while failing on one nobody did. Human audits are the check on that blind spot — they are the tool for questions that were never encoded, which is why they cannot be replaced by more metrics.
Decision Rules and Exam Signals
Rule 1 — an aggregate metric cannot answer a fairness question. If the scenario gives one overall number and asks about groups, the answer involves splitting it.
Rule 2 — "report per group" is the habit. Subgroup analysis underlies every tool in 4.1.7.
Rule 3 — check presence before proportion. Inclusivity asks whether a group is there; balance asks whether there is enough of it. You cannot be disproportionate about zero.
Rule 4 — curated means documented, not neutral. A well-curated dataset can be thoroughly skewed; curation is what lets you find out.
Rule 5 — read which "bias" is meant. Overfitting, underfitting, training accuracy signal the statistical sense. Demographic groups, fairness signal the societal one.
Rule 6 — two accuracy figures mean a fitting question. Poor/poor is underfitting; excellent/poor is overfitting.
Rule 7 — good/good says nothing about fairness. A well-fitted model on skewed data is the exact case this domain exists for.
Rule 8 — the tool is selected by when, not what. Clarify is point-in-time, Model Monitor is continuous, A2I is per-prediction.
Rule 9 — "since launch" is always Model Monitor. Any change-over-time phrasing rules out the one-off measurement tools.
Rule 10 — more data does not fix coverage. More of the same sources repeats the same skew; the fix is data from the groups that are missing.
Rule 11 — removing a demographic attribute does not remove bias. Correlated features carry it; and removing the attribute destroys your ability to measure the disparity.
Rule 12 — human audits catch what no metric was configured to look for. They are not a fallback.
Distractor Patterns
| Pattern | What it looks like | How to defuse it |
|---|---|---|
| Aggregate accuracy offered as fairness evidence | "The model is 94% accurate, so it performs well for all users" | An average discards the distribution; it cannot answer "for whom" |
| More data offered for a coverage gap | "Collect more training data" for a model failing one group | More of the same sources repeats the skew; the gap is coverage, not volume |
| Dropping the demographic column | "Remove the attribute so the model cannot discriminate" | Correlated features carry it anyway — and now the disparity cannot be measured |
| Clarify for continuous monitoring | Offered when the scenario says "since launch" | Clarify is point-in-time; drift over time is Model Monitor |
| Model Monitor before deployment | Offered to assess a training dataset | It monitors a deployed endpoint; there is nothing to monitor yet |
| A2I as a bias metric | Offered to measure disparity across a population | A2I routes individual predictions to humans; it does not compute population statistics |
| Overfitting named for a group disparity | A model failing rural users called "overfitted" | Overfitting is train-versus-test generalisation, not group coverage |
| Underfitting offered for a strong-but-unfair model | Model scores well overall, one group poorly | Underfitting means poor on both training and unseen data |
| Balance offered when the group is absent | "Rebalance the dataset" for a group with no records | Rebalancing needs rows to reweight; absence is an inclusivity failure |
| Curation treated as a fairness guarantee | "Sources were carefully curated, so the data is unbiased" | Curation documents provenance; it does not correct skew |
| Human audit dismissed as unscalable | Offered as inferior to automated metrics | It is the only tool for disparities nobody configured a metric for |
The first two are the highest-frequency pair in this task statement, and both are answered by the same instinct: ask what the number is an average of, and ask what the new data would be more of.
Scenario Walkthrough
A national lender deploys a model that pre-screens loan applications. At launch it was assessed at 93% accuracy and signed off. Six months on, community groups report that applicants from two rural provinces are rejected far more often than comparable urban applicants. The team checks and finds overall accuracy is still 93%. Training data was assembled from fifteen years of historical decisions, drawn from the lender's branch network, which is concentrated in cities; rural applicants are present but make up under 3% of records. Historical approvals were labelled by branch managers applying their own judgement, with no shared rubric. The team proposes removing the province field from the model's inputs and collecting more applications from its existing branches.
| Observation | Reading | Decision |
|---|---|---|
| 93% overall, unchanged, with a reported group disparity | The aggregate is the disguise, not the evidence | Subgroup analysis — recompute accuracy per province |
| Rural applicants under 3% of records | Present but swamped — a balance failure, not inclusivity | Rebalancing is possible because they are present |
| Fifteen years of branch-network data | Skew is in the sampling frame itself | A curation/provenance finding: it is a sample of city branches |
| Labels applied by managers with no shared rubric | Bias entered at the labelling stage, not only collection | Analyze label quality — the training target may itself be prejudiced |
| Proposal: drop the province field | Correlated features carry it; and disparity becomes unmeasurable | Reject — this hides the problem rather than fixing it |
| Proposal: collect more from existing branches | More of the same skew | Reject — the gap is coverage, not volume |
| Disparity appeared over six months | Change since launch | Model Monitor for the ongoing check |
Seven observations, and the two proposals are both wrong — which is the point of the scenario.
The label row is the one that separates strong candidates. Everything else in the stem points at the sample, and a candidate who diagnoses "unrepresentative training data" has found something real and stopped one step early. Fifteen years of decisions labelled by individual managers with no shared rubric means the training target itself encodes their judgement. Rebalancing the sample does not repair a label that was prejudiced when it was written; the model would learn the same rule from a better-proportioned dataset.
And note what the province proposal really costs. It is not merely ineffective — removing the attribute removes the column you need to compute the disparity, so the model becomes unfair in a way that can no longer be detected. Measurability is not a side benefit of keeping a sensitive attribute; it is often the reason to keep it.
Key Concepts
| Term | Definition |
|---|---|
| Inclusivity | Whether the groups a system will serve are present in the dataset at all; a presence property, distinct from proportion |
| Diversity | Whether the data spans the real range of cases, conditions and contexts the system will meet |
| Balanced dataset | One in which groups appear in workable proportion, so no group is so small the training objective effectively ignores it |
| Curated data source | A source whose origin is known, selected and documented; enables checking for skew but does not by itself prevent it |
| Statistical bias | Error from a model too simple to capture the underlying pattern; observed as underfitting |
| Societal bias | Systematically different outcomes across demographic groups; can coexist with an excellent statistical fit |
| Variance | Sensitivity to the particular training set; observed as overfitting |
| Underfitting | Poor accuracy on training data and unseen data; the model never learned the pattern |
| Overfitting | Strong accuracy on training data and poor accuracy on unseen data; the model memorised rather than generalised |
| Subgroup analysis | Computing a metric per group instead of over the whole population; the technique that makes disparity visible |
| Label quality analysis | Checking whether labels are correct and consistently applied, since a prejudiced label teaches a prejudiced rule |
| Human audit | Direct human review of model behaviour; the only check for disparities no metric was configured to detect |
| Amazon SageMaker Clarify | Measures bias in data before training and in a model after training, and produces feature attributions; point-in-time |
| SageMaker Model Monitor | Continuously watches a deployed model for drift in data, model quality and bias against a baseline |
| Amazon Augmented AI (Amazon A2I) | Routes individual predictions to human reviewers at inference time, typically those below a confidence threshold |
| Proxy variable | A feature correlated with a sensitive attribute that carries its information even after the attribute is removed |
Revision Flashcards
Say the answer aloud before revealing it.
1. Why can a model be 94% accurate and still fail a fairness assessment? → Because an aggregate metric is an average, and an average discards the distribution it summarises. A model at 94% overall can be 97% for a majority group and 61% for a minority group; both figures are consistent with the headline, and nothing in the headline moves when the disparity worsens. The aggregate is not wrong — it simply cannot answer "for whom", which is what a fairness question asks.
2. Name the four dataset characteristics in Objective 4.1.5 and the question each asks. → Inclusivity — are the groups the system will serve present at all? Diversity — does the data span the real range of cases and conditions? Balance — are those groups present in workable proportion? Curated sources — is the origin known, selected and documented? They are checked in that order, because proportion is meaningless for a group that is absent.
3. Distinguish inclusivity from balance. → Inclusivity is a presence question; balance is a proportion question. A dataset with 40,000 records from one group and 300 from another is inclusive and unbalanced. The consequence differs: an inclusivity failure cannot be measured at all because there are no rows to compute over, whereas an imbalance can be measured, which is what makes it findable by subgroup analysis.
4. Does "curated data sources" mean the data is unbiased? → No. Curation means the origin was selected and documented rather than accumulated by accident. A carefully curated dataset drawn entirely from one country is well-curated and thoroughly skewed. What curation buys is the ability to say what the dataset is a sample of — it is the precondition for checking skew, not a guarantee against it.
5. The word "bias" carries two meanings in Objective 4.1.6. What are they? → Statistical bias is error from a model too simple to capture the pattern, and it shows up as underfitting — poor on training data and on new data. Societal bias is systematically different outcomes across demographic groups. They are independent: a model can fit its training data beautifully, so low statistical bias, while producing severe societal bias because that data was skewed.
6. A model scores 99% on training data and 71% on data it has never seen. What is happening? → Overfitting, which is high variance. The model memorised the training set rather than learning a pattern that generalises. The signature is the gap between the two figures, not the value of either one — strong training accuracy with poor unseen-data accuracy is the definition.
7. A model scores poorly on both training data and unseen data. What is happening? → Underfitting, which is high statistical bias. The model is too simple to capture the underlying pattern, so it never learned it in the first place. Note that this is a capability failure affecting everyone roughly equally, which is why it is not primarily a fairness condition despite sharing the bullet with them.
8. Can a model be neither overfitting nor underfitting and still be unfair? → Yes, and this is the case the domain exists for. A model can fit well on training data and generalise well to unseen data — good on both headline figures — while failing one demographic group badly, because it learned a genuine pattern present in unrepresentative data. Fitting diagnostics do not substitute for subgroup analysis.
9. What separates SageMaker Clarify from SageMaker Model Monitor? → Time. Clarify is a point-in-time measurement: it assesses bias in data before training and in a model after training. Model Monitor watches a deployed endpoint continuously, alerting on drift in data, model quality and bias against a baseline. "Before we deploy" selects Clarify; "since launch" or "over the past quarter" selects Model Monitor.
10. What does Amazon A2I do, and what is it not? → It routes individual predictions to human reviewers at inference time, typically those falling below a confidence threshold. It is not a bias metric and does not compute population statistics — it produces a human decision on one case. A scenario asking to measure disparity across a population is not describing A2I.
11. Why is subgroup analysis singled out when it is not a product? → Because it is the technique every tool on the list either performs or acts upon. It means computing any metric per group rather than over the whole population, and it is the step that makes a disparity visible at all. The habit it encodes — report per group, never in aggregate — is what the objective is really testing.
12. A team removes the demographic attribute so the model "cannot discriminate." What is wrong? → Two things. Correlated features act as proxies and carry the same information, so the disparity usually survives. And removing the attribute removes the column needed to compute per-group metrics, so the disparity can no longer be detected. The result is a model that is unfair in a way nobody can now measure — worse than the starting position.
13. Why are human audits not made redundant by automated bias metrics? → Because a metric can only measure a disparity somebody configured it to look for. Choosing which groups to compare is itself a human judgement, so a model can be provably fair on every attribute that was tested and fail on one nobody thought of. Human audits are the check on that blind spot, which is a different job from computing more metrics.
The Five-Beat Answer
The core question this chapter prepares you for: "How would you know whether this model is fair?"
Five beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.
- Refuse the aggregate — say explicitly that an overall accuracy figure cannot answer the question, and that the first step is recomputing it per group. Starting here rather than with a tool name is what separates a designed answer from a remembered one.
- Name the groups, and say who chose them — subgroup analysis requires deciding which groups to compare, and that choice is a human judgement that can miss one. Saying so out loud is the beat that earns a follow-up question rather than suffering one.
- Trace where bias could have entered — collection, labelling, training, deployment, feedback. Say which stage the evidence points at, and note that labelling is the one most often missed when everything in the scenario points at the sample.
- Match the instrument to the timing — Clarify for a point-in-time assessment before or after training, Model Monitor for continuous drift after launch, A2I for human review of individual predictions, human audits for what no metric was configured to catch.
- Say what you would not do — do not drop the sensitive attribute, and do not collect more of the same data. Naming the two plausible wrong moves demonstrates you understand why they are wrong, and both appear in real proposals.
A strong answer starts by rejecting the headline number. A weak answer names a service in the first sentence.
Why This Helps You
On the job: the most common way a fairness problem survives is not that somebody ignored it. It is that the dashboard said 93% and nobody asked what the 93% was an average of. Reporting per group by default costs almost nothing to set up and is the single change that makes these failures visible while they are still cheap to fix.
In interviews: "how would you check this model for bias?" separates candidates who name a service from those who describe a method. The strong answer refuses the aggregate first, distinguishes the two senses of "bias", and mentions that removing a sensitive attribute destroys measurability — that last point is an unusually reliable marker of someone who has actually done this.
On the exam: Domain 4 is 14% of scored content across two task statements, and Task 4.1 is the larger of them with seven objectives. These three carry the domain's most scenario-heavy questions, because dataset properties and detection tools are what a scenario can actually describe. Candidates who memorise the six tool names without the timing distinction lose the questions that name two of them in the same stem.
Chapter Checklist
- I can explain why an aggregate accuracy figure cannot answer a fairness question
- I can name the four dataset characteristics and the question each one asks
- I can distinguish inclusivity from balance and say why the order matters
- I can explain why curated does not mean neutral
- I can name the five stages where bias enters and what the feedback loop does
- I can separate statistical bias from societal bias and say which signals each
- I can identify overfitting and underfitting from a pair of accuracy figures
- I can explain how a well-fitted model can still be unfair
- I can name all six detection tools from Objective 4.1.7
- I can separate Clarify, Model Monitor and A2I by when each one runs
- I can explain why Clarify appears in two different objectives
- I can say why removing a sensitive attribute makes things worse, not better
- I can explain why human audits are not replaceable by more metrics
After the Chapter
- Complete
student/project.md— parts 49-51 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one. - Take
student/quiz.mdclosed-book, then review the reasoning for every question you guessed, including the ones you got right. Pay particular attention to questions 3 and 4 — they are deliberate mirror images, and missing both means the train-versus-unseen reading has not landed. - Open the official v1.1 exam guide's Domain 4 page and confirm you can attach a concept from this chapter to each of the last three bullets under Task Statement 4.1. Note that the bullet on effects of bias and variance names demographic effects and overfitting in the same line, and satisfy yourself that you can say why both belong there.
- Next: Chapter 15 — Transparency and Explainability (Domain 4, Task 4.2). This chapter measured whether a model's outcomes are fair. Chapter 15 asks whether anyone can tell why it decided what it decided: transparent versus explainable models, SageMaker Model Cards, Clarify in its second role producing feature attributions, Amazon Bedrock Model Evaluations, open-source models and licensing, the genuine trade-off between safety and transparency, and human-centered design for explainable AI.
Chapter 14 quiz
13 questions on this chapter, marked instantly, with an explanation for every answer.