Chapter 7 · Domain 2 · 24% of the exam
Choosing GenAI, and the AWS GenAI Stack
26 min read · Chapter 7 of 17
On this page
- Certification Blueprint
- What This Chapter Covers
- The Advantages of GenAI
- The Disadvantages — Three Failures Before Four Names
- Selecting a Model: the Eight Factors
- Business Value and Metrics
- The AWS GenAI Stack
- Why Build on AWS GenAI Services
- What the Infrastructure Itself Contributes
- Cost Trade-offs
- Decision Rules and Exam Signals
- Distractor Patterns
- Scenario Walkthrough
- Key Concepts
- Revision Flashcards
- The Four-Beat Answer
- Why This Helps You
- Chapter Checklist
- After the Chapter
Certification Blueprint
| Field | Coverage |
|---|---|
| Exam | AWS Certified AI Practitioner (AIF-C01), exam guide v1.1 |
| Domain | Content Domain 2 — Fundamentals of GenAI |
| Exam weight | 24% of scored content |
| Task statements | 2.2 Capabilities and limitations of GenAI · 2.3 AWS infrastructure and technologies for GenAI |
| Objectives | 2.2.1-2.2.4 and 2.3.1-2.3.4 — all eight |
What This Chapter Covers
This chapter closes Domain 2, and it carries eight of the domain's fourteen objectives — the widest brief in the domain.
The character of the material changes here. Chapters 04, 05 and 06 were definitions: what a foundation model is, what a token costs, what an agent does. This chapter is decisions, and almost every exam question in Tasks 2.2 and 2.3 is built on a conflict — two things the scenario wants that cannot both be had.
| Chapters 04-06 gave you | This chapter asks |
|---|---|
| What a foundation model, token and embedding are | Should this problem use one at all? |
| How token pricing and context work | What does each serving choice cost? |
| What an agent, MCP and orchestration are | Which AWS service runs it? |
| Definitions | Decisions, under conflict |
Questions here rarely ask whether generative AI is good. They ask which limit disqualifies it, or which factor wins when two of them pull in opposite directions.
The Advantages of GenAI
Objective 2.2.1 names four. They are not a list of virtues — each is a specific capability a scenario can be reaching for.
| Advantage | What it actually means |
|---|---|
| Adaptability | One model serves many tasks without being retrained for each |
| Responsiveness | It answers now, on unseen input, rather than after a build cycle |
| Conversational capabilities | It holds a multi-turn exchange in which each turn depends on the last |
| Ability to generate content | It produces new artefacts, rather than classifying or scoring existing ones |
Two ways to hold it:
- A general contractor with a broad crew rather than a single specialist tradesperson. Any job gets started today; none of them is done by someone who has done only that job for twenty years.
- A fluent colleague in a language you do not speak. They will answer anything you ask immediately, and they will never say "I do not know that word." Keep this one — it carries the limits as well as the advantages.
The examinable habit is naming which advantage a scenario invokes:
| The scenario says | Advantage being invoked |
|---|---|
| "Handles enquiries about products we add every week" | Adaptability |
| "Must answer questions nobody anticipated" | Responsiveness |
| "The follow-up question depends on the previous answer" | Conversational |
| "Draft the first version of the report" | Generation |
"GenAI is powerful" is not a reason. "The catalogue changes weekly, and adaptability removes a retraining cycle per product" is.
The Disadvantages — Three Failures Before Four Names
Look at these before reading the definitions underneath them.
| What was observed | What it looked like to the business |
|---|---|
| An assistant cited a refund policy clause, with a section number, that has never existed in the policy | Confident, specific, correctly formatted — and invented |
| The same question, asked twice in one afternoon, produced two different totals | Neither answer was flagged as uncertain; both were stated plainly |
| A rejected loan application could not be explained to the applicant, because nobody could say which input drove the decision | Commercially defensible and legally indefensible |
None of these is a bug that a patch fixes. Each is a property of the technology that a design must account for. Now the names — Objective 2.2.2 gives four, and they are four different failures with four different mitigations:
| Disadvantage | Precisely | The failure above |
|---|---|---|
| Hallucination | Content that is fluent, plausible and not grounded in any source | The invented policy clause |
| Nondeterminism | The same input may produce different output on different runs | The two different totals |
| Interpretability | Why a given output was produced cannot be readily traced | The unexplainable rejection |
| Inaccuracy | The output is simply wrong against a known correct answer | Any of them, once checked |
Hallucination and inaccuracy are not synonyms, and this is the most common conflation in the objective. An inaccurate answer is wrong. A hallucinated answer is unfounded — it can be accidentally correct and still be a hallucination, because nothing grounded it. The mitigations differ, which is why the exam keeps them apart: inaccuracy is addressed by evaluating against a benchmark, hallucination by grounding the output in retrieved source material and citing it.
Nondeterminism cannot be configured away. Inference parameters influence how much variation you get — Chapter 08 covers them — but the property is inherent. A requirement for byte-identical repeatable output is a reason to reject generative AI for that part of the system, not a setting to hunt for.
What to remember from this diagram: two of the three gates are exits. This is Chapter 02's suitability screen one level down — there the question was whether to use AI at all, here it is whether to use generative AI given these four specific limits. The middle exit matters most: "not open-ended" sends you back to traditional ML or a managed AI service, which is a Domain 1 answer appearing correctly inside a Domain 2 question.
Selecting a Model: the Eight Factors
Objective 2.2.3 names eight. Learners memorise them as a flat list and then cannot use them, because a flat list gives no way to resolve a conflict — and conflict is what the questions are made of.
| Factor | The question it asks |
|---|---|
| Model types | Does the modality match what goes in and comes out? |
| Constraints | What is technically or contractually forbidden here? |
| Compliance | What is legally or regulatorily required? |
| Performance requirements | What quality bar must the output clear? |
| Latency | How fast must the answer arrive? |
| Capabilities | Can it do the specific things this use case needs? |
| Cost | What does it cost at the volume we actually expect? |
| Model complexity | How much model is the job worth? |
Model complexity is the least obvious of the eight: it asks how much model the job is worth. A large general model pointed at a narrow, well-defined task is a real and common wrong answer.
What to remember from this diagram: gates one and two eliminate. Only gate three ranks. This is the single most examinable sentence in Task 2.2, because distractors are built exactly against it — an option that is cheaper, faster, and violates a stated constraint.
That structure resolves conflicts without needing intuition:
| Conflict in the scenario | Which wins | Why |
|---|---|---|
| Compliance requires regional data residency · the best model is elsewhere | Compliance | A hard constraint cannot be traded, however good the model |
| Latency budget is 300 ms · the larger model is more accurate | Latency | Stated performance envelope; an answer that arrives late is not an answer |
| Cost is tight · a smaller model meets the quality bar | Cost | Once the bar is met, further quality is not free value |
| Cost is tight · no model within budget meets the quality bar | Neither — re-scope | This is no longer a model-selection problem |
The third row is counterintuitive for engineers, who read "more accurate" as strictly better. The fourth is the one candidates avoid: sometimes no option is acceptable, and a question that offers that reading is testing whether you will force a choice anyway.
Business Value and Metrics
Objective 2.2.4 names seven metrics. The examinable skill is not reciting them — it is knowing which end of the chain each one sits on.
| Metric | Which end |
|---|---|
| Accuracy | Model |
| Cross-domain performance | Model — does it hold up outside the data it was tuned on? |
| Efficiency | The operational bridge — time, volume, rework |
| ROI | Business |
| Conversion rate | Business |
| Average revenue per user | Business |
| Customer lifetime value | Business |
Cross-domain performance is worth learning properly, because the name does not give it away: it asks whether the model holds up outside the data it was tuned on. It is a model metric, and it is the one that predicts whether a business metric will survive contact with real usage.
What to remember from this diagram: the weak answer stops at the first box. "It is 94% accurate" is not a business case — it is the first link of one. Four of the seven named metrics are business metrics, because a sponsor does not buy accuracy; a sponsor buys the outcome accuracy caused.
The AWS GenAI Stack
Objective 2.3.1 names seven services. Several are recent additions to the exam guide — study material written against the earlier version will not contain Kiro, Strands Agents or Amazon Bedrock AgentCore at all. Amazon Quick is not one of the additions: it was already in scope, and it is Amazon Q that joined the in-scope list alongside those three.
| Service | What it is for |
|---|---|
| Amazon Bedrock | Reach managed foundation models from multiple providers behind one API |
| Amazon SageMaker AI | Build, train and deploy models with full control of the process |
| SageMaker JumpStart | Start from pre-trained models and solution templates rather than from nothing |
| Amazon Bedrock AgentCore | Run agents in production — deployment, tool access, observability, security at scale |
| Strands Agents | An open-source SDK for building agents, in Python and TypeScript |
| Kiro | Agentic development — turning prompts into executable specs |
| Amazon Quick | An AI companion for work — research, business insights and automation |
Learn the need each one answers, not the order of the list. An exam question supplies a need, never a service name — nobody is asked "what is Kiro", they are asked what fits a described situation.
What to remember from this diagram: two services in the same layer are alternatives; two in different layers are usually used together. That one sentence answers a great many "which service" questions. In particular, Strands Agents is what you build an agent with and Amazon Bedrock AgentCore is what you run it on — they are not competitors, and a question offering both as competing choices is testing exactly that.
Amazon Quick is not Amazon Q
Both are in scope for this exam. They are different services, and only one is named in Objective 2.3.1.
| Amazon Quick | Amazon Q | |
|---|---|---|
| Named in Objective 2.3.1? | Yes | No |
| Listed under | Analytics | Developer Tools |
| In one line | An AI companion for work — research, business insights, automation | A generative AI assistant in the developer-tools family |
Three more name pairs that cost marks:
| Do not confuse | With |
|---|---|
| Amazon Bedrock — reach and use managed FMs | Amazon Bedrock AgentCore — run agents in production |
| Amazon SageMaker AI — build and train with full control | SageMaker JumpStart — start from pre-trained models and templates |
| Strands Agents — the SDK you build an agent with | Amazon Bedrock AgentCore — the platform you run it on |
The general habit is worth more than the four specific pairs: when two options differ by one word, the question is usually about that word.
Why Build on AWS GenAI Services
Objective 2.3.2 names six advantages. Each is best stated as what it removes — that is what turns a list into an answer.
| Advantage | What it removes |
|---|---|
| Accessibility | Needing to source, host and operate a model yourself |
| Lower barrier to entry | Needing a specialist team before the first working prototype |
| Efficiency | Rebuilding shared plumbing for every application |
| Cost-effectiveness | Paying for idle capacity you provisioned in advance |
| Speed to market | The lead time between deciding and shipping |
| Ability to meet business objectives | The gap between a demonstration and something operable |
A managed service is not "easier". It is a different set of things that are now somebody else's job, and the exam wants you to name which things.
What the Infrastructure Itself Contributes
Objective 2.3.3 names four benefits of the infrastructure, as distinct from the services built on it. The third column is the examinable part.
| Benefit | What the platform provides | What is still yours |
|---|---|---|
| Security | Isolation, encryption, identity and access control | Deciding who should have access, and to what |
| Compliance | Audited programs and evidence you can inherit | Proving your own use case meets its own obligations |
| Responsibility | Tooling to evaluate and constrain model behaviour | Deciding what behaviour is acceptable here |
| Safety | Guardrail mechanisms that can filter and block | Defining what must be filtered or blocked |
The shared responsibility model does not disappear because a service is managed — it moves. A managed service does not make an application compliant; it gives you evidence you can inherit for the parts AWS operates. Domain 5 examines this properly.
Cost Trade-offs
Objective 2.3.4 names eight considerations. Every one is a trade — the cheaper choice always costs something else, and a candidate who cannot name what it costs has not understood it.
| Trade-off | What you gain | What you give up |
|---|---|---|
| Token-based pricing | No commitment; pay only for what is used | Spend rises with usage, and is hard to cap |
| Provisioned throughput | Predictable responsiveness and capacity | Paid whether used or not |
| Custom models | Behaviour prompting cannot reach | Training and hosting cost, and a maintenance obligation |
| Redundancy and availability | Survives a failure | Duplicate capacity, paid continuously |
| Regional coverage | Data residency and lower latency near users | Not every model is offered in every Region |
| Performance | A larger or faster tier | Cost rises faster than the quality gain |
Token-based pricing appeared in Chapter 05 as a mechanism. Here it is one option among several — the objective asks you to trade it against the alternatives, not to explain how it works.
Regional coverage catches people out, because it interacts directly with a data-residency constraint from gate one: not every model is offered in every Region.
What to remember from this diagram: traffic shape picks the serving mode; requirement picks the model. They are separate decisions, and questions frequently blend them to see whether you will. Note the "stop here" node — reaching for a custom model before prompting has been tried is the expensive wrong answer, and it appears in distractors constantly because it sounds like the thorough option.
Decision Rules and Exam Signals
Rule 1 — name the advantage, not the technology. Say which of the four the scenario reaches for and what it removes.
Rule 2 — hallucination is unfounded; inaccuracy is wrong. Different failures, different fixes.
Rule 3 — nondeterminism is inherent. If byte-identical output is required, that is an exit, not a configuration task.
Rule 4 — gates one and two eliminate, gate three ranks. Compliance, constraints and model type are never traded for cost.
Rule 5 — carry a model metric through to a business metric. Accuracy → efficiency → ROI or conversion or revenue per user or lifetime value.
Rule 6 — same layer means alternatives, different layers mean used together. This resolves most "which service" questions.
Rule 7 — traffic shape picks the serving mode. "Unpredictable" and "near zero overnight" point at on-demand; "steady, high, all day" points at provisioned throughput.
Rule 8 — the cheapest path that meets the requirement wins. Prompting before retrieval, retrieval before fine-tuning.
Distractor Patterns
| Pattern | What it looks like | How to defuse it |
|---|---|---|
| Advantage stated in general | "GenAI is flexible and powerful" | Name which of the four, and what the scenario needed it for |
| Hallucination as inaccuracy | Treating the two as one failure | Unfounded vs wrong — different mitigations |
| Nondeterminism as a bug | Offering to "fix" it with configuration | It is inherent; design around it or reject GenAI |
| Trading a hard constraint | Compliance sacrificed for cost or quality | Gates one and two eliminate; only gate three ranks |
| Model metric as business value | "94% accurate" offered as the business case | Carry it through efficiency to a business metric |
| Amazon Q for Amazon Quick | The similar name selected | Different services; only Amazon Quick is in Objective 2.3.1 |
| Bedrock for AgentCore | The familiar name selected | Bedrock reaches models; AgentCore runs agents |
| Custom model reached for first | Fine-tuning offered before prompting is tried | The cheapest path that meets the requirement wins |
The fourth is the single most reliable way to lose a Task 2.2 question, because the offending option is usually the most attractive one on cost and speed.
Scenario Walkthrough
A national insurer wants an assistant that answers policyholder questions in conversation. Policy wordings change quarterly. Regulation requires that any statement made to a policyholder be traceable to the policy document it came from, and that customer data stay within the country. Traffic is steady and high all day. The sponsor has asked what the business case is.
| Requirement | Reading | Decision |
|---|---|---|
| Multi-turn policyholder questions | Conversational capability, and content generated per question | GenAI fits on advantage |
| Every statement traceable to a source | Hallucination is disqualifying unless grounded | GenAI only with grounding and citation |
| Customer data stays in-country | Compliance — a hard constraint | Gate one; regional coverage decides the shortlist |
| Wordings change quarterly | Adaptability, and an argument against a custom model | Prompting and retrieval before fine-tuning |
| Steady, high, all-day traffic | Predictable capacity is worth reserving | Provisioned throughput over on-demand |
| "What is the business case?" | Accuracy is not the answer | Carry through efficiency to ROI and customer lifetime value |
Six requirements, six different objectives. The one nearly always missed is the fourth: quarterly change is an argument against a custom model, not for one. Content that changes faster than a training cycle is a retrieval problem — fine-tuning changes behaviour, not facts, and a custom model would be permanently behind while carrying a maintenance obligation nobody scoped.
Key Concepts
| Term | Definition |
|---|---|
| Adaptability | One model serving many tasks without being retrained for each — the advantage that removes a per-task build cycle |
| Responsiveness | Answering now, on unseen input, rather than after a build cycle |
| Conversational capability | Holding a multi-turn exchange in which each turn depends on the previous one |
| Hallucination | Fluent, plausible content that is not grounded in any source; it may be accidentally correct and is still a hallucination |
| Nondeterminism | The property that identical input may produce different output on different runs; inherent, not a defect |
| Interpretability | The degree to which the reason for a given output can be traced |
| Inaccuracy | Output that is wrong against a known correct answer |
| Cross-domain performance | A model metric asking whether performance holds up outside the data the model was tuned on |
| Provisioned throughput | Reserved model capacity bought for predictable responsiveness, paid whether used or not |
| Amazon Bedrock | The service for reaching managed foundation models from multiple providers behind one API |
| Amazon Bedrock AgentCore | The platform for running agents in production — deployment, tool access, observability and security at scale |
| Strands Agents | An open-source SDK, in Python and TypeScript, for building agents |
| Kiro | An agentic development tool that turns prompts into executable specifications |
| Amazon Quick | An AI companion for work — research, business insights and automation; the service named in Objective 2.3.1, distinct from Amazon Q |
| SageMaker JumpStart | The on-ramp of pre-trained models and solution templates, as distinct from building from nothing in Amazon SageMaker AI |
Revision Flashcards
Say the answer aloud before revealing it.
1. Name the four advantages of GenAI, and say what makes naming one better than praising GenAI. → Adaptability, responsiveness, conversational capabilities, and the ability to generate content. Naming one identifies what the scenario actually needed and what that advantage removes — "GenAI is powerful" identifies nothing and answers no question.
2. What is the difference between a hallucination and an inaccuracy? → An inaccurate answer is wrong against a known correct answer. A hallucinated answer is unfounded — nothing grounded it. A hallucination can be accidentally correct and still be a hallucination. Inaccuracy is addressed by evaluation against a benchmark; hallucination by grounding and citation.
3. Can nondeterminism be switched off? → No. Inference parameters influence how much variation you get, but the property is inherent. A requirement for byte-identical repeatable output is a reason to reject generative AI for that part of the system, not a setting to find.
4. Name the eight model-selection factors, and the three groups they sort into. → Model types, constraints, compliance (hard constraints); performance requirements, latency, capabilities (performance envelope); cost and model complexity (economics). The first two groups eliminate; only the third ranks.
5. Compliance and cost conflict. Which wins, and what is the better way to say why? → Compliance — but the structural answer is stronger: compliance is a hard constraint, so it eliminates rather than ranks. An option that saves money by violating a stated constraint has not made a trade-off, it has answered a different question.
6. What does cross-domain performance measure? → Whether a model's performance holds up outside the data it was tuned on. It is a model metric, and it is the one that predicts whether a business metric will survive real usage.
7. Why is "it is 94% accurate" not a business case? → Accuracy is a model metric and only the first link of the chain. A business case carries it through an operational effect — usually efficiency in time, volume or rework — to a business metric such as ROI, conversion rate, average revenue per user or customer lifetime value.
8. Name the seven services in Objective 2.3.1. → Amazon Bedrock, Amazon SageMaker AI, SageMaker JumpStart, Amazon Quick, Kiro, Strands Agents, and Amazon Bedrock AgentCore.
9. How do Amazon Quick and Amazon Q differ, and which is examinable under Objective 2.3.1? → They are two different services and both are in scope for the exam. Amazon Quick — an AI companion for work covering research, business insights and automation — is the one named in Objective 2.3.1. Amazon Q is not named there and sits under Developer Tools in the in-scope list.
10. Strands Agents or Amazon Bedrock AgentCore — when would you pick one over the other? → The framing is the trap: they are not alternatives. Strands Agents is the open-source SDK you build an agent with; Amazon Bedrock AgentCore is the platform you run one on. A question offering both as competing choices is testing whether you know that.
11. Traffic falls to near zero overnight and spikes during promotions. On-demand or provisioned throughput? → On-demand, token-based pricing. Reserved capacity is paid whether used or not, so it is wasted on traffic that drops to near zero. Provisioned throughput is for steady, high, predictable volume where the reservation is actually used.
12. A knowledge base changes weekly. Why is fine-tuning the wrong answer? → Because fine-tuning changes behaviour, not facts, and a training cycle is slower than the rate of change — the model is permanently behind and carries a maintenance obligation. Frequently changing content is a retrieval problem: ground the model in the current source at request time.
The Four-Beat Answer
The core question this chapter prepares you for: "How would you decide whether to use generative AI for this, and what would you build it on?"
Four beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.
- Fit — name which of the four advantages the requirement actually reaches for, then name the limit that could disqualify it. Say explicitly that nondeterminism is inherent and that a requirement for identical repeatable output is an exit rather than a configuration problem.
- Selection — sort the factors into hard constraints, performance envelope and economics, and state that only the economics may be traded. Compliance, constraints and model type eliminate.
- The stack — name the service by the need it answers, not by familiarity: Bedrock to reach managed models, SageMaker AI for full control, JumpStart to start from pre-trained, AgentCore to run agents in production, Strands Agents to build one, Kiro for spec-driven development, Amazon Quick as an AI companion for work.
- Economics — pick the serving mode from the traffic shape, and name what the choice costs. Then carry the value claim through efficiency to a named business metric rather than stopping at accuracy.
A strong answer names an exit and a trade. A weak answer lists capabilities.
Why This Helps You
On the job: the most expensive mistakes in generative AI projects are made at this stage, not in implementation — a custom model commissioned for a fast-changing-content problem, or reserved capacity bought for traffic that does not justify it. Both are decisions that look thorough and cost money for the life of the system.
In interviews: "when would you not use generative AI?" is a standard senior screening question, and it separates people who have shipped from people who have demonstrated. Naming nondeterminism as inherent, and traceability as a grounding requirement rather than a model choice, is what a strong answer sounds like.
On the exam: these eight objectives are the decision-heavy half of a domain worth 24% of scored content. The single highest-value habit is the gate structure — hard constraints eliminate, economics rank — because the most attractive distractor in a Task 2.2 question is almost always the one that buys cost or speed by breaking a stated constraint.
Chapter Checklist
- I can name all four advantages, and say which one a given scenario is reaching for
- I can separate hallucination, nondeterminism, interpretability and inaccuracy as four distinct failures
- I can explain why nondeterminism cannot be configured away
- I can sort the eight selection factors into hard constraints, performance envelope and economics
- I can say which factors may be traded and which may not
- I can carry a model metric through efficiency to a named business metric
- I can name the seven services in Objective 2.3.1 and the need each one answers
- I can distinguish Amazon Quick from Amazon Q, and Amazon Bedrock from Amazon Bedrock AgentCore
- I can state what AWS infrastructure provides and what remains the builder's responsibility
- I can choose a serving mode from the traffic shape, and name what that choice costs
After the Chapter
- Complete
student/project.md— parts 27-30 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one. - Take
student/quiz.mdclosed-book, then review the reasoning for every question you guessed, including the ones you got right. - Open the official v1.1 exam guide's Domain 2 page and confirm you can attach a concept from this chapter to each of the four bullets under Task Statement 2.2 and each of the four under Task Statement 2.3. Then open the In-Scope AWS Services page and find Amazon Quick and Amazon Q in their two different categories.
- Domain 2 is now complete. Next: Chapter 08 — Designing FM Applications: Selection, Inference Parameters, and Agents (Domain 3, Task 3.1). Domain 3 is the largest domain on the exam at 28%, and it asks how rather than whether. Chapter 08 opens it with the criteria that pick one foundation model over another — including the inference parameters this chapter deferred.
Chapter 7 quiz
13 questions on this chapter, marked instantly, with an explanation for every answer.