Chapter 9 · Domain 3 · 28% of the exam
RAG, Vector Stores, and the Customization Cost Ladder
27 min read · Chapter 9 of 17
On this page
- Certification Blueprint
- What This Chapter Covers
- What Is Retrieval Augmented Generation?
- RAG in Business Terms
- Amazon Bedrock Knowledge Bases
- Where Embeddings Live: the Four Named Services
- The Customization Cost Ladder
- Fresh Facts, or Wrong Behaviour?
- Decision Rules and Exam Signals
- Distractor Patterns
- Scenario Walkthrough
- Key Concepts
- Revision Flashcards
- The Five-Beat Answer
- Why This Helps You
- Chapter Checklist
- After the Chapter
Certification Blueprint
| Field | Coverage |
|---|---|
| Exam | AWS Certified AI Practitioner (AIF-C01), exam guide v1.1 |
| Domain | Content Domain 3 — Applications of Foundation Models |
| Exam weight | 28% of scored content — the largest domain on the exam |
| Task statement | 3.1 Design considerations for applications that use foundation models |
| Objectives | 3.1.3 RAG and its business applications · 3.1.4 AWS services that store embeddings in vector databases · 3.1.5 cost tradeoffs of FM customization |
What This Chapter Covers
A foundation model was trained on an enormous amount of general text. It has never seen your refund policy, your product catalogue, or the price change that went live on Tuesday. Every question in this chapter follows from that one fact.
There are only two families of answer, and telling them apart is the skill:
| Give it the material | Change the model | |
|---|---|---|
| What moves | The prompt — new content arrives with each request | The weights — the model itself is different afterwards |
| When it knows something new | Immediately, as soon as the source changes | After a training cycle completes |
| What it costs | Tokens per request, plus retrieval infrastructure | Training, hosting, and a permanent maintenance obligation |
| What it fixes | Missing facts | Wrong behaviour |
Almost every wrong answer in Task 3.1 is the second column applied to a first-column problem. It is attractive because it sounds more thorough — and "more thorough" is exactly what the exam punishes here.
What Is Retrieval Augmented Generation?
RAG is fetching relevant material from your own sources at the moment a question is asked, and putting it into the prompt so the model answers from it rather than from memory.
Two ways to hold it:
- An open-book exam. The candidate has not learned more; they have been handed the textbook and told which page. The same person, with the page open, now answers correctly.
- A doctor with your file on the desk. Their training did not change when they picked up your notes. What changed is what is in front of them while they answer.
Both analogies carry the load-bearing point: the person is unchanged. Nothing was learned and nothing was forgotten. That is what makes retrieval immediate and cheap, and it is the property the exam tests hardest.
What to remember from this diagram: every step happens per request, and none of them is a training step. That is why a document published a minute ago is answerable a minute later. Trace the flow once in your own words — question in, vector out, nearest passages back, prompt assembled, answer generated with citations — because a question that garbles any one of those steps is testing whether you can.
What RAG does not do
This definition is examined by its negative space more often than by its positive form.
| RAG does not | Because |
|---|---|
| Change the model's weights | Retrieved text enters the prompt, not the parameters |
| Teach the model anything permanently | The next request starts from the same unchanged model |
| Improve the model's writing style or tone | Those are behaviours; retrieval supplies content |
| Make the model better at reasoning | It supplies material to reason over, not the capability |
| Remove the need to evaluate | A wrong retrieval produces a confidently wrong grounded answer |
A scenario that says "so the model learns our documentation" is describing training, not retrieval. The word learns is very often planted in the distractor.
The last row deserves its own sentence, because it survives into real systems: a citation proves where the text came from, never that it was the right text. If a withdrawn policy is still in the index, an answer citing it looks exactly like a correct one.
RAG in Business Terms
Objective 3.1.3 asks for business applications, not architecture. The pattern repeats: an authoritative internal corpus, changing faster than any training cycle, where a wrong answer is expensive.
| Business application | What retrieval provides |
|---|---|
| Customer support over policies and terms | Current wording, and a citation the agent can show |
| Internal knowledge search across wikis and tickets | An answer instead of ten links to read |
| Product and catalogue questions | This week's specification, not last quarter's |
| Regulated advice, claims and underwriting | Traceability — every statement points at its source |
| Employee onboarding and HR self-service | One answer that matches the current handbook |
Citation is the business feature. In several of these an ungrounded answer is not merely worse, it is unusable — nobody may act on a claim that cannot be traced back to an approved source.
Amazon Bedrock Knowledge Bases
Objective 3.1.3 names this service explicitly. It is managed RAG: the retrieval pipeline offered as a service rather than assembled as an application.
| You provide | The service handles |
|---|---|
| Documents in a data source | Splitting them into passages |
| A choice of embedding model | Embedding every passage |
| A vector store, or a managed default | Indexing, and keeping the index current |
| A question at request time | Retrieval, prompt assembly, and citations |
The exam signal is the phrase "without building a retrieval pipeline." A scenario that wants grounded answers over its own documents and does not want to operate the plumbing is naming this service.
What to remember from this diagram: this half runs when documents change, not per request. Confusing it with the retrieval flow above is the most common structural error in this objective — one is ingestion, the other is retrieval, and only the question is embedded per request.
Where Embeddings Live: the Four Named Services
Objective 3.1.4 names four AWS services. They are not four competing products to rank. Each answers a different question about where your data already sits and what shape the query takes.
| Service | What it is | The scenario that selects it |
|---|---|---|
| Amazon OpenSearch Service | Search and analytics engine with vector search | Search is the primary workload; large corpora; hybrid keyword-plus-vector retrieval |
| Amazon Aurora | Managed relational database, PostgreSQL-compatible, with vector support | The data is already relational and vectors should sit beside it, at Aurora's scale |
| Amazon RDS for PostgreSQL | Managed PostgreSQL with vector support | An existing PostgreSQL estate; add vectors without adopting a new data store |
| Amazon Neptune | Graph database, with vector search alongside graph traversal | Relationships between entities are part of the answer, not just similarity |
What to remember from this diagram: the first question is about the data, not the product. A scenario that describes an existing estate has usually already chosen for you. No exam question asks which vector store is best, because none of the four is a bad product — the wrong answers in this objective are true statements about the wrong service for this scenario, which is a much harder distractor to spot than a false one.
Amazon Aurora, and why it is here
Aurora entered the exam's in-scope service list at v1.1. Study material written against the earlier version of the guide will not contain it, and it appears in this objective.
| What it is | A managed relational database service, PostgreSQL-compatible, with vector storage and similarity search |
| Why it appears in an AI objective | Vectors can live in the same database as the rows they describe |
| What that removes | A second data store to synchronise, secure, back up and pay for |
| The scenario that names it | Relational data already in Aurora, and a requirement to search it semantically |
| The pairing to keep straight | Aurora and RDS for PostgreSQL are both PostgreSQL-compatible; the existing footprint decides |
Vectors beside the rows they describe is the whole idea. A product's embedding sits in the same database as its price and stock level, and one query reaches both. If you have met Aurora only inside this objective it is easy to assume it must be a purpose-built vector product — it is not, and that assumption breaks the selection signal, which is "the data is already in Aurora."
The Customization Cost Ladder
Customization is any approach that makes a general foundation model behave usefully for your specific situation — from adding a sentence to a prompt, all the way to training a model.
Two ways to hold it:
- Renting a hall for an event. Bring your own decorations; hire a stylist; refit the room; build a venue. Each step costs more, commits you further, and is occasionally the only thing that works.
- Getting a suit. Wear it as sold; have the sleeves taken up; have it re-cut; have one made from scratch. Nobody commissions bespoke tailoring because the sleeves are long.
Objective 3.1.5 does not ask you to perform any of these. It asks you to explain the cost tradeoffs, which is a question about choosing.
| Rung | Approach | What actually changes |
|---|---|---|
| 1 | In-context learning | Nothing but the prompt — instructions and examples supplied per request |
| 2 | RAG | Still nothing in the model; retrieved source material joins the prompt |
| 3 | Fine-tuning | The model's weights, adapted on your examples |
| 4 | Distillation | A new, smaller model trained to reproduce a larger one's behaviour |
| 5 | Pre-training | A model built from scratch on a large corpus |
Rungs 1 and 2 leave the model alone. Rungs 3, 4 and 5 produce a model you now own — and owning a model is a standing obligation to retrain, re-evaluate and host, not a one-off purchase.
What to remember from this diagram: read the labels on the arrows, not the boxes. Each label is a condition that must be true before the climb is justified. If the condition is not established in the scenario, the answer is the rung you are already standing on. Climbing is justified by a condition, not by the failure of the rung below.
What each rung costs, and what it buys
| Approach | What it costs | What it buys | The condition that justifies it |
|---|---|---|---|
| In-context learning | Tokens per request; nothing up front | Immediate change, zero commitment | Always try first — it is the baseline |
| RAG | Retrieval infrastructure, embedding, larger prompts | Current facts and citations, with no training | The model lacks facts, and they change |
| Fine-tuning | A training job, custom hosting, a maintenance obligation | Behaviour, format and domain style prompting cannot reach | The behaviour is wrong and prompting has genuinely failed |
| Distillation | Training a student model against a teacher | Lower per-request cost and latency at scale | Volume makes the large model's unit cost the problem |
| Pre-training | Enormous data, compute and expertise | A model that exists on your terms | No suitable model exists — very rarely the exam's answer |
Distillation changes the shape of the cost
The other four rungs simply get more expensive as you climb. Distillation is the one that spends up front to make each request cheaper afterwards, which is why it is the only rung justified by volume rather than by capability.
| Before distillation | After distillation | |
|---|---|---|
| Model used at request time | Large, capable, expensive per call | Smaller, faster, cheaper per call |
| Up-front cost | None | A training cycle against the larger model |
| Where it pays back | — | Only at volume — the saving is per request |
| What can go wrong | — | The smaller model loses capability the larger one had |
Ask what the scenario is complaining about. "The model cannot do X" is not a distillation problem. "The model does X well and we cannot afford it at this volume" is exactly one.
Fresh Facts, or Wrong Behaviour?
What to remember from this diagram: this single question separates most Task 3.1 answers. Facts are retrieved; behaviour is tuned — and if neither is genuinely established by the scenario, the answer is a better prompt. The two failures look very similar in a question stem, which is why the exam can build a pair of near-identical scenarios with opposite answers.
When RAG is not the answer
Retrieval is this task statement's usual answer, which makes the cases where it is wrong worth naming explicitly. A candidate who has learned "always RAG" has simply learned a different wrong rule.
| The scenario says | Why retrieval does not fix it | What does |
|---|---|---|
| "Responses do not follow our house style" | Style is behaviour; there is no document to retrieve | Prompting first, then fine-tuning |
| "It must answer in a specialist notation it handles badly" | The capability is missing, not the material | Fine-tuning, or a different model |
| "Per-request cost is too high at our volume" | Retrieval increases prompt size and cost | Distillation, or a smaller model |
| "The facts are stable and already well known" | There is nothing to keep current | In-context learning |
| "Every claim must cite an approved source" | — this is the retrieval case | RAG, and grounding is mandatory |
Decision Rules and Exam Signals
Rule 1 — retrieval fills the prompt; training changes the weights. If an option says the model "learns" your documents, it is describing the other family.
Rule 2 — documents are embedded on change, questions on request. Two pipelines, two triggers.
Rule 3 — facts are retrieved, behaviour is tuned. Diagnose before choosing.
Rule 4 — the existing estate picks the vector store. Relationships mean Neptune; already on PostgreSQL means Aurora or RDS for PostgreSQL; search-first at scale means OpenSearch Service.
Rule 5 — the cheapest rung that meets the requirement wins. In-context learning, then RAG, then fine-tuning, then distillation, then pre-training.
Rule 6 — climb on a condition, not on frustration. Each arrow has a stated condition; if the scenario has not established it, do not climb.
Rule 7 — distillation is about volume, not capability. It preserves behaviour at lower unit cost and adds nothing new.
Rule 8 — grounded is not the same as correct. Citation proves provenance; evaluation proves correctness.
Distractor Patterns
| Pattern | What it looks like | How to defuse it |
|---|---|---|
| "So the model learns our documents" | RAG described as if it trained the model | Retrieval fills the prompt; weights never change |
| Fine-tuning for fresh facts | Retraining offered for content that changes weekly | Facts are retrieved; a training cycle is slower than the change |
| Ranking the vector stores | "Which is the best vector database?" | None is; the existing estate and query shape decide |
| Climbing past the working rung | Fine-tuning offered before prompting was tried | The cheapest approach that meets the requirement wins |
| Distillation for capability | Offered to fix something the model cannot do | Distillation reduces unit cost; it does not add capability |
| Pre-training as the thorough answer | Building from scratch for a domain problem | Almost never correct; name what it would actually require |
| Ingestion confused with retrieval | Embedding described as happening per question | Documents are embedded on change; questions per request |
| Grounded therefore correct | Citation treated as proof | A stale or wrong passage yields a confidently wrong grounded answer |
| True statement, wrong service | An accurate description of OpenSearch attached to a graph requirement | All four vector stores are good; match the stated data shape |
The fourth and sixth are the two most reliable ways to lose marks in Objective 3.1.5, because the offending option is the one that sounds most rigorous.
Scenario Walkthrough
A hospital group wants an assistant that answers clinician questions about its own treatment protocols. Protocols are revised continuously by a clinical governance committee. Every answer shown to a clinician must be traceable to the approved protocol document it came from. The protocol text is already stored relationally in Aurora alongside approval status and revision metadata. Answers currently read as too informal for clinical use.
| Requirement | Reading | Decision |
|---|---|---|
| Answers about the group's own protocols | Facts the model has never seen | Retrieval, not training |
| Protocols revised continuously | Faster than any training cycle | Confirms RAG; rules out fine-tuning for currency |
| Every answer traceable to its source | Grounding and citation are mandatory | RAG with citations — the business feature, not a nicety |
| Text already relational in Aurora | The estate has chosen the store | Amazon Aurora as the vector store |
| Answers read as too informal | Behaviour, not facts | Prompting first; fine-tuning only if that genuinely fails |
Five requirements, and the last one is the test. Four of them point at retrieval; the fifth is a different kind of problem sitting inside the same scenario. A candidate who has decided "this is a RAG question" will answer the tone problem with retrieval too — and retrieving the style guide delivers the very instruction that prompting already failed with.
Key Concepts
| Term | Definition |
|---|---|
| Retrieval Augmented Generation (RAG) | Fetching relevant material from your own sources at request time and placing it in the prompt, so the model answers from it rather than from memory; the model's weights are never modified |
| Grounding | Constraining an answer to retrieved source material, so that every claim traces back to a document rather than to the model's memory |
| Citation | The reference returned alongside a grounded answer identifying which source passage it came from; proves provenance, not correctness |
| Ingestion | The pipeline that runs when documents change — splitting into passages, embedding them, and indexing the vectors |
| Retrieval | The per-request step that embeds the question, finds the nearest stored passages, and returns them for the prompt |
| Amazon Bedrock Knowledge Bases | Managed RAG — the retrieval pipeline offered as a service, handling splitting, embedding, indexing, retrieval and citation |
| Amazon OpenSearch Service | A search and analytics engine with vector search; selected when search is the primary workload at large scale |
| Amazon Aurora | A PostgreSQL-compatible managed relational database that can store vectors beside the relational rows they describe; new to this exam's scope at v1.1 |
| Amazon RDS for PostgreSQL | Managed PostgreSQL with vector support; selected when an existing PostgreSQL estate should gain vectors without a new data store |
| Amazon Neptune | A graph database offering vector search alongside graph traversal; selected when relationships between entities are part of the answer |
| In-context learning | Supplying instructions and examples in the prompt itself; rung 1, the cheapest form of customization, with nothing paid up front |
| Fine-tuning | Adapting a model's weights on your own examples; rung 3, justified when behaviour rather than facts is wrong and prompting has failed |
| Model distillation | Training a smaller student model to reproduce a larger teacher model's behaviour; rung 4, justified by per-request cost at volume rather than by capability |
| Pre-training | Building a model from scratch on a large corpus; rung 5, justified only when no existing model fits the domain at all |
Revision Flashcards
Say the answer aloud before revealing it.
1. Define RAG in one sentence, and name the thing it leaves unchanged. → Fetching relevant material from your own sources at the moment a question is asked and placing it in the prompt, so the model answers from it rather than from memory. What it leaves unchanged is the model itself — the weights are never modified, which is why a document published a minute ago is answerable a minute later.
2. A colleague says the RAG system means the model "has learned" the company handbook. What is wrong? → Nothing was learned. Retrieved text enters the prompt for that one request and is gone afterwards; the model is byte-identical before and after. There is no threshold of requests at which retrieval turns into training — a system that has served ten million requests has a model in exactly the state it started in.
3. When are source documents embedded, and when are questions embedded? → Documents are embedded during ingestion, which runs when the documents change. Questions are embedded per request, during retrieval. Two pipelines with two different triggers, and confusing them is the most common structural error in this objective.
4. What does a citation prove, and what does it not prove? → It proves provenance — which source passage the answer came from. It does not prove correctness. If a withdrawn policy is still in the index, an answer citing it looks exactly like a correct one, which is why evaluating a RAG application is examined separately.
5. Name the four AWS services in Objective 3.1.4 and the signal that selects each. → Amazon OpenSearch Service when search is the primary workload at large scale; Amazon Aurora when the data is already relational in Aurora and vectors should sit beside it; Amazon RDS for PostgreSQL when there is an existing PostgreSQL estate; Amazon Neptune when relationships between entities are part of the answer rather than just similarity.
6. Why is Amazon Aurora in an AI objective, and what is it actually? → It is a managed PostgreSQL-compatible relational database, not a purpose-built vector product. It appears because vectors can live in the same database as the rows they describe, removing a second data store to synchronise and secure. It entered this exam's scope at v1.1, so older study material will not mention it.
7. List the five customization approaches from cheapest to most expensive. → In-context learning, RAG, fine-tuning, distillation, pre-training. The first two leave the model untouched; the last three produce a model you then own and must retrain, re-evaluate and host.
8. What is the difference between a facts problem and a behaviour problem? → A facts problem is the model not having information — it never saw your policy, your catalogue, your protocol. That is retrieved. A behaviour problem is the model getting format, tone or domain style wrong. That is tuned. They look similar in a question stem and have opposite answers.
9. Why is distillation justified by volume rather than by capability? → Because it spends up front to make each request cheaper afterwards, and that saving only pays back at volume. It compresses behaviour a larger model already has — it adds nothing new. "The model cannot do X" is not a distillation problem; "the model does X well and we cannot afford it at this volume" is.
10. A catalogue changes weekly. Why is weekly fine-tuning the wrong answer? → Because it answers a facts problem with training. Even if every job finished on time, the model would answer from a snapshot averaging half a week old while stock changes hourly, and each cycle buys a permanent maintenance obligation. A scheduling fix cannot repair a category error.
11. Answers are factually right but do not sound like the organisation. Prompting has failed. What now, and why not RAG? → Fine-tune on examples written in the house style, because style is behaviour rather than fact. Retrieving the style guide would deliver a description of the style — the same instruction prompting has already proved insufficient with.
12. When is pre-training the reasonable choice? → When no existing model fits the domain at all. It is justified by the absence of an alternative, not by the difficulty of the problem or by a fine-tuned model missing its quality bar. A question offering it is usually testing whether thoroughness will be mistaken for correctness.
The Five-Beat Answer
The core question this chapter prepares you for: "How would you make a foundation model answer questions about our business, and how far would you go?"
Five beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.
- Diagnose — say first whether the gap is facts the model never saw or behaviour it gets wrong. Naming this before proposing anything is what separates a designed answer from a remembered one.
- Retrieve — for a facts problem, describe RAG: fetch from your sources at request time, place it in the prompt, answer from it, cite it. Say explicitly that the model is unchanged, because that is what makes it immediate.
- Store — name where the embeddings live and why: OpenSearch Service for search-first at scale, Aurora or RDS for PostgreSQL when the data is already relational there, Neptune when relationships decide the answer. Say that the existing estate usually chooses.
- Ladder — order the five approaches by cost and give the condition on each climb. Name distillation's as volume rather than capability; it is the one most often stated wrongly.
- Stop — state the cheapest rung that meets the requirement and say what the next one would have cost. Then add that grounding is not evaluation: a citation proves provenance, not correctness.
A strong answer diagnoses before it designs, and names a rung it deliberately did not climb. A weak answer proposes fine-tuning.
Why This Helps You
On the job: the single most expensive avoidable mistake in this area is commissioning a custom model for a problem retrieval would have solved. It looks thorough at the point of approval, and it converts a content problem into a permanent obligation to retrain, re-evaluate and host — while staying, by construction, permanently out of date. The diagnosis in beat one is what prevents it.
In interviews: "when would you fine-tune instead of using RAG?" is a standard screening question, and most candidates answer it as a preference. The strong answer is structural — facts versus behaviour, plus the cost the climb commits you to — and mentioning that a citation proves provenance rather than correctness marks out someone who has actually run one of these systems.
On the exam: Domain 3 is 28% of scored content, the largest on the exam, and these three objectives are the design-decision core of it. The highest-value habit is refusing to climb: when two options both work, the cheaper one is the answer, and the more thorough-sounding option is the trap.
Chapter Checklist
- I can define RAG without describing training, and say what it leaves unchanged
- I can walk the request flow from question to grounded, cited answer
- I can separate ingestion (on document change) from retrieval (per request)
- I can name business applications of RAG and say why citation is the business feature
- I can say what Amazon Bedrock Knowledge Bases removes from a build
- I can name the four vector stores and the scenario signal that selects each
- I can explain why Amazon Aurora appears in this objective, and how it differs from RDS for PostgreSQL
- I can order the five customization approaches by cost and state each climb condition
- I can explain why distillation is justified by volume rather than capability
- I can diagnose whether a scenario describes missing facts or wrong behaviour
- I can say why a grounded answer is not automatically a correct one
After the Chapter
- Complete
student/project.md— parts 34-36 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one. - Take
student/quiz.mdclosed-book, then review the reasoning for every question you guessed, including the ones you got right. Pay particular attention to questions 7 and 9 — they are deliberate mirror images, and missing both means the facts-versus-behaviour diagnosis has not landed. - Open the official v1.1 exam guide's Domain 3 page and confirm you can attach a concept from this chapter to the third, fourth and fifth bullets under Task Statement 3.1. Then find Amazon Aurora on the In-Scope AWS Services page and note which category it sits in.
- Next: Chapter 10 — Prompt Engineering (Domain 3, Task 3.2). This chapter named the bottom rung of the ladder but did not teach it. Chapter 10 is that rung in full: the constructs of a prompt, the techniques that make in-context learning work, and the risks — exposure, poisoning, hijacking and jailbreaking — that arrive as soon as retrieved and user-supplied text share a prompt.
Chapter 9 quiz
13 questions on this chapter, marked instantly, with an explanation for every answer.