Free live cohort on Google Meet — register your interest →

Chapter 9 · Domain 3 · 28% of the exam

RAG, Vector Stores, and the Customization Cost Ladder

27 min read · Chapter 9 of 17

On this page
  1. Certification Blueprint
  2. What This Chapter Covers
  3. What Is Retrieval Augmented Generation?
  4. RAG in Business Terms
  5. Amazon Bedrock Knowledge Bases
  6. Where Embeddings Live: the Four Named Services
  7. The Customization Cost Ladder
  8. Fresh Facts, or Wrong Behaviour?
  9. Decision Rules and Exam Signals
  10. Distractor Patterns
  11. Scenario Walkthrough
  12. Key Concepts
  13. Revision Flashcards
  14. The Five-Beat Answer
  15. Why This Helps You
  16. Chapter Checklist
  17. After the Chapter

Certification Blueprint

Field Coverage
Exam AWS Certified AI Practitioner (AIF-C01), exam guide v1.1
Domain Content Domain 3 — Applications of Foundation Models
Exam weight 28% of scored content — the largest domain on the exam
Task statement 3.1 Design considerations for applications that use foundation models
Objectives 3.1.3 RAG and its business applications · 3.1.4 AWS services that store embeddings in vector databases · 3.1.5 cost tradeoffs of FM customization

What This Chapter Covers

A foundation model was trained on an enormous amount of general text. It has never seen your refund policy, your product catalogue, or the price change that went live on Tuesday. Every question in this chapter follows from that one fact.

There are only two families of answer, and telling them apart is the skill:

Give it the material Change the model
What moves The prompt — new content arrives with each request The weights — the model itself is different afterwards
When it knows something new Immediately, as soon as the source changes After a training cycle completes
What it costs Tokens per request, plus retrieval infrastructure Training, hosting, and a permanent maintenance obligation
What it fixes Missing facts Wrong behaviour

Almost every wrong answer in Task 3.1 is the second column applied to a first-column problem. It is attractive because it sounds more thorough — and "more thorough" is exactly what the exam punishes here.

What Is Retrieval Augmented Generation?

RAG is fetching relevant material from your own sources at the moment a question is asked, and putting it into the prompt so the model answers from it rather than from memory.

Two ways to hold it:

  • An open-book exam. The candidate has not learned more; they have been handed the textbook and told which page. The same person, with the page open, now answers correctly.
  • A doctor with your file on the desk. Their training did not change when they picked up your notes. What changed is what is in front of them while they answer.

Both analogies carry the load-bearing point: the person is unchanged. Nothing was learned and nothing was forgotten. That is what makes retrieval immediate and cheap, and it is the property the exam tests hardest.

Left-to-right flow in which a user question is embedded into a vector, used to search a vector store for the nearest chunks, returning top matching chunks from the organisation's own sources, which are assembled into a prompt alongside the question and sent to a foundation model whose weights are unchanged, producing a grounded answer with citations

What to remember from this diagram: every step happens per request, and none of them is a training step. That is why a document published a minute ago is answerable a minute later. Trace the flow once in your own words — question in, vector out, nearest passages back, prompt assembled, answer generated with citations — because a question that garbles any one of those steps is testing whether you can.

What RAG does not do

This definition is examined by its negative space more often than by its positive form.

RAG does not Because
Change the model's weights Retrieved text enters the prompt, not the parameters
Teach the model anything permanently The next request starts from the same unchanged model
Improve the model's writing style or tone Those are behaviours; retrieval supplies content
Make the model better at reasoning It supplies material to reason over, not the capability
Remove the need to evaluate A wrong retrieval produces a confidently wrong grounded answer

A scenario that says "so the model learns our documentation" is describing training, not retrieval. The word learns is very often planted in the distractor.

The last row deserves its own sentence, because it survives into real systems: a citation proves where the text came from, never that it was the right text. If a withdrawn policy is still in the index, an answer citing it looks exactly like a correct one.

RAG in Business Terms

Objective 3.1.3 asks for business applications, not architecture. The pattern repeats: an authoritative internal corpus, changing faster than any training cycle, where a wrong answer is expensive.

Business application What retrieval provides
Customer support over policies and terms Current wording, and a citation the agent can show
Internal knowledge search across wikis and tickets An answer instead of ten links to read
Product and catalogue questions This week's specification, not last quarter's
Regulated advice, claims and underwriting Traceability — every statement points at its source
Employee onboarding and HR self-service One answer that matches the current handbook

Citation is the business feature. In several of these an ungrounded answer is not merely worse, it is unusable — nobody may act on a claim that cannot be traced back to an approved source.

Amazon Bedrock Knowledge Bases

Objective 3.1.3 names this service explicitly. It is managed RAG: the retrieval pipeline offered as a service rather than assembled as an application.

You provide The service handles
Documents in a data source Splitting them into passages
A choice of embedding model Embedding every passage
A vector store, or a managed default Indexing, and keeping the index current
A question at request time Retrieval, prompt assembly, and citations

The exam signal is the phrase "without building a retrieval pipeline." A scenario that wants grounded answers over its own documents and does not want to operate the plumbing is naming this service.

Left-to-right ingestion flow in which source documents in Amazon S3 are split into passages, each passage is embedded with an embedding model, the vectors are indexed in a vector database, and the result is ready to retrieve at request time

What to remember from this diagram: this half runs when documents change, not per request. Confusing it with the retrieval flow above is the most common structural error in this objective — one is ingestion, the other is retrieval, and only the question is embedded per request.

Where Embeddings Live: the Four Named Services

Objective 3.1.4 names four AWS services. They are not four competing products to rank. Each answers a different question about where your data already sits and what shape the query takes.

Service What it is The scenario that selects it
Amazon OpenSearch Service Search and analytics engine with vector search Search is the primary workload; large corpora; hybrid keyword-plus-vector retrieval
Amazon Aurora Managed relational database, PostgreSQL-compatible, with vector support The data is already relational and vectors should sit beside it, at Aurora's scale
Amazon RDS for PostgreSQL Managed PostgreSQL with vector support An existing PostgreSQL estate; add vectors without adopting a new data store
Amazon Neptune Graph database, with vector search alongside graph traversal Relationships between entities are part of the answer, not just similarity

Decision flow asking first whether relationships between entities decide the answer, routing yes to Amazon Neptune for graph traversal plus vectors, and no to a second question about whether the data is already in a PostgreSQL database, routing to Amazon Aurora for vectors beside relational data, Amazon RDS for PostgreSQL for vectors in the existing database, and Amazon OpenSearch Service for search-first workloads at large scale

What to remember from this diagram: the first question is about the data, not the product. A scenario that describes an existing estate has usually already chosen for you. No exam question asks which vector store is best, because none of the four is a bad product — the wrong answers in this objective are true statements about the wrong service for this scenario, which is a much harder distractor to spot than a false one.

Amazon Aurora, and why it is here

Aurora entered the exam's in-scope service list at v1.1. Study material written against the earlier version of the guide will not contain it, and it appears in this objective.

What it is A managed relational database service, PostgreSQL-compatible, with vector storage and similarity search
Why it appears in an AI objective Vectors can live in the same database as the rows they describe
What that removes A second data store to synchronise, secure, back up and pay for
The scenario that names it Relational data already in Aurora, and a requirement to search it semantically
The pairing to keep straight Aurora and RDS for PostgreSQL are both PostgreSQL-compatible; the existing footprint decides

Vectors beside the rows they describe is the whole idea. A product's embedding sits in the same database as its price and stock level, and one query reaches both. If you have met Aurora only inside this objective it is easy to assume it must be a purpose-built vector product — it is not, and that assumption breaks the selection signal, which is "the data is already in Aurora."

The Customization Cost Ladder

Customization is any approach that makes a general foundation model behave usefully for your specific situation — from adding a sentence to a prompt, all the way to training a model.

Two ways to hold it:

  • Renting a hall for an event. Bring your own decorations; hire a stylist; refit the room; build a venue. Each step costs more, commits you further, and is occasionally the only thing that works.
  • Getting a suit. Wear it as sold; have the sleeves taken up; have it re-cut; have one made from scratch. Nobody commissions bespoke tailoring because the sleeves are long.

Objective 3.1.5 does not ask you to perform any of these. It asks you to explain the cost tradeoffs, which is a question about choosing.

Rung Approach What actually changes
1 In-context learning Nothing but the prompt — instructions and examples supplied per request
2 RAG Still nothing in the model; retrieved source material joins the prompt
3 Fine-tuning The model's weights, adapted on your examples
4 Distillation A new, smaller model trained to reproduce a larger one's behaviour
5 Pre-training A model built from scratch on a large corpus

Rungs 1 and 2 leave the model alone. Rungs 3, 4 and 5 produce a model you now own — and owning a model is a standing obligation to retrain, re-evaluate and host, not a one-off purchase.

Vertical ladder of five customization approaches, from in-context learning using prompt and examples only, climbing to RAG which retrieves the source at request time when the model lacks facts rather than behaviour, to fine-tuning where the model's weights change when behaviour or style is wrong rather than the facts, to distillation which trains a smaller student model when per-request cost or latency must fall at scale, to pre-training which builds a model from scratch when no existing model fits the domain at all

What to remember from this diagram: read the labels on the arrows, not the boxes. Each label is a condition that must be true before the climb is justified. If the condition is not established in the scenario, the answer is the rung you are already standing on. Climbing is justified by a condition, not by the failure of the rung below.

What each rung costs, and what it buys

Approach What it costs What it buys The condition that justifies it
In-context learning Tokens per request; nothing up front Immediate change, zero commitment Always try first — it is the baseline
RAG Retrieval infrastructure, embedding, larger prompts Current facts and citations, with no training The model lacks facts, and they change
Fine-tuning A training job, custom hosting, a maintenance obligation Behaviour, format and domain style prompting cannot reach The behaviour is wrong and prompting has genuinely failed
Distillation Training a student model against a teacher Lower per-request cost and latency at scale Volume makes the large model's unit cost the problem
Pre-training Enormous data, compute and expertise A model that exists on your terms No suitable model exists — very rarely the exam's answer

Distillation changes the shape of the cost

The other four rungs simply get more expensive as you climb. Distillation is the one that spends up front to make each request cheaper afterwards, which is why it is the only rung justified by volume rather than by capability.

Before distillation After distillation
Model used at request time Large, capable, expensive per call Smaller, faster, cheaper per call
Up-front cost None A training cycle against the larger model
Where it pays back — Only at volume — the saving is per request
What can go wrong — The smaller model loses capability the larger one had

Ask what the scenario is complaining about. "The model cannot do X" is not a distillation problem. "The model does X well and we cannot afford it at this volume" is exactly one.

Fresh Facts, or Wrong Behaviour?

Decision flow asking whether the missing thing is a fact the model never saw, routing yes to retrieving it with RAG while leaving the model unchanged, and no to a second question asking whether format, tone or style is wrong despite good prompting, routing yes to fine-tuning where the weights change and no to improving the prompt first with in-context learning

What to remember from this diagram: this single question separates most Task 3.1 answers. Facts are retrieved; behaviour is tuned — and if neither is genuinely established by the scenario, the answer is a better prompt. The two failures look very similar in a question stem, which is why the exam can build a pair of near-identical scenarios with opposite answers.

When RAG is not the answer

Retrieval is this task statement's usual answer, which makes the cases where it is wrong worth naming explicitly. A candidate who has learned "always RAG" has simply learned a different wrong rule.

The scenario says Why retrieval does not fix it What does
"Responses do not follow our house style" Style is behaviour; there is no document to retrieve Prompting first, then fine-tuning
"It must answer in a specialist notation it handles badly" The capability is missing, not the material Fine-tuning, or a different model
"Per-request cost is too high at our volume" Retrieval increases prompt size and cost Distillation, or a smaller model
"The facts are stable and already well known" There is nothing to keep current In-context learning
"Every claim must cite an approved source" — this is the retrieval case RAG, and grounding is mandatory

Decision Rules and Exam Signals

Rule 1 — retrieval fills the prompt; training changes the weights. If an option says the model "learns" your documents, it is describing the other family.

Rule 2 — documents are embedded on change, questions on request. Two pipelines, two triggers.

Rule 3 — facts are retrieved, behaviour is tuned. Diagnose before choosing.

Rule 4 — the existing estate picks the vector store. Relationships mean Neptune; already on PostgreSQL means Aurora or RDS for PostgreSQL; search-first at scale means OpenSearch Service.

Rule 5 — the cheapest rung that meets the requirement wins. In-context learning, then RAG, then fine-tuning, then distillation, then pre-training.

Rule 6 — climb on a condition, not on frustration. Each arrow has a stated condition; if the scenario has not established it, do not climb.

Rule 7 — distillation is about volume, not capability. It preserves behaviour at lower unit cost and adds nothing new.

Rule 8 — grounded is not the same as correct. Citation proves provenance; evaluation proves correctness.

Distractor Patterns

Pattern What it looks like How to defuse it
"So the model learns our documents" RAG described as if it trained the model Retrieval fills the prompt; weights never change
Fine-tuning for fresh facts Retraining offered for content that changes weekly Facts are retrieved; a training cycle is slower than the change
Ranking the vector stores "Which is the best vector database?" None is; the existing estate and query shape decide
Climbing past the working rung Fine-tuning offered before prompting was tried The cheapest approach that meets the requirement wins
Distillation for capability Offered to fix something the model cannot do Distillation reduces unit cost; it does not add capability
Pre-training as the thorough answer Building from scratch for a domain problem Almost never correct; name what it would actually require
Ingestion confused with retrieval Embedding described as happening per question Documents are embedded on change; questions per request
Grounded therefore correct Citation treated as proof A stale or wrong passage yields a confidently wrong grounded answer
True statement, wrong service An accurate description of OpenSearch attached to a graph requirement All four vector stores are good; match the stated data shape

The fourth and sixth are the two most reliable ways to lose marks in Objective 3.1.5, because the offending option is the one that sounds most rigorous.

Scenario Walkthrough

A hospital group wants an assistant that answers clinician questions about its own treatment protocols. Protocols are revised continuously by a clinical governance committee. Every answer shown to a clinician must be traceable to the approved protocol document it came from. The protocol text is already stored relationally in Aurora alongside approval status and revision metadata. Answers currently read as too informal for clinical use.

Requirement Reading Decision
Answers about the group's own protocols Facts the model has never seen Retrieval, not training
Protocols revised continuously Faster than any training cycle Confirms RAG; rules out fine-tuning for currency
Every answer traceable to its source Grounding and citation are mandatory RAG with citations — the business feature, not a nicety
Text already relational in Aurora The estate has chosen the store Amazon Aurora as the vector store
Answers read as too informal Behaviour, not facts Prompting first; fine-tuning only if that genuinely fails

Five requirements, and the last one is the test. Four of them point at retrieval; the fifth is a different kind of problem sitting inside the same scenario. A candidate who has decided "this is a RAG question" will answer the tone problem with retrieval too — and retrieving the style guide delivers the very instruction that prompting already failed with.

Key Concepts

Term Definition
Retrieval Augmented Generation (RAG) Fetching relevant material from your own sources at request time and placing it in the prompt, so the model answers from it rather than from memory; the model's weights are never modified
Grounding Constraining an answer to retrieved source material, so that every claim traces back to a document rather than to the model's memory
Citation The reference returned alongside a grounded answer identifying which source passage it came from; proves provenance, not correctness
Ingestion The pipeline that runs when documents change — splitting into passages, embedding them, and indexing the vectors
Retrieval The per-request step that embeds the question, finds the nearest stored passages, and returns them for the prompt
Amazon Bedrock Knowledge Bases Managed RAG — the retrieval pipeline offered as a service, handling splitting, embedding, indexing, retrieval and citation
Amazon OpenSearch Service A search and analytics engine with vector search; selected when search is the primary workload at large scale
Amazon Aurora A PostgreSQL-compatible managed relational database that can store vectors beside the relational rows they describe; new to this exam's scope at v1.1
Amazon RDS for PostgreSQL Managed PostgreSQL with vector support; selected when an existing PostgreSQL estate should gain vectors without a new data store
Amazon Neptune A graph database offering vector search alongside graph traversal; selected when relationships between entities are part of the answer
In-context learning Supplying instructions and examples in the prompt itself; rung 1, the cheapest form of customization, with nothing paid up front
Fine-tuning Adapting a model's weights on your own examples; rung 3, justified when behaviour rather than facts is wrong and prompting has failed
Model distillation Training a smaller student model to reproduce a larger teacher model's behaviour; rung 4, justified by per-request cost at volume rather than by capability
Pre-training Building a model from scratch on a large corpus; rung 5, justified only when no existing model fits the domain at all

Revision Flashcards

Say the answer aloud before revealing it.

1. Define RAG in one sentence, and name the thing it leaves unchanged. → Fetching relevant material from your own sources at the moment a question is asked and placing it in the prompt, so the model answers from it rather than from memory. What it leaves unchanged is the model itself — the weights are never modified, which is why a document published a minute ago is answerable a minute later.

2. A colleague says the RAG system means the model "has learned" the company handbook. What is wrong? → Nothing was learned. Retrieved text enters the prompt for that one request and is gone afterwards; the model is byte-identical before and after. There is no threshold of requests at which retrieval turns into training — a system that has served ten million requests has a model in exactly the state it started in.

3. When are source documents embedded, and when are questions embedded? → Documents are embedded during ingestion, which runs when the documents change. Questions are embedded per request, during retrieval. Two pipelines with two different triggers, and confusing them is the most common structural error in this objective.

4. What does a citation prove, and what does it not prove? → It proves provenance — which source passage the answer came from. It does not prove correctness. If a withdrawn policy is still in the index, an answer citing it looks exactly like a correct one, which is why evaluating a RAG application is examined separately.

5. Name the four AWS services in Objective 3.1.4 and the signal that selects each. → Amazon OpenSearch Service when search is the primary workload at large scale; Amazon Aurora when the data is already relational in Aurora and vectors should sit beside it; Amazon RDS for PostgreSQL when there is an existing PostgreSQL estate; Amazon Neptune when relationships between entities are part of the answer rather than just similarity.

6. Why is Amazon Aurora in an AI objective, and what is it actually? → It is a managed PostgreSQL-compatible relational database, not a purpose-built vector product. It appears because vectors can live in the same database as the rows they describe, removing a second data store to synchronise and secure. It entered this exam's scope at v1.1, so older study material will not mention it.

7. List the five customization approaches from cheapest to most expensive. → In-context learning, RAG, fine-tuning, distillation, pre-training. The first two leave the model untouched; the last three produce a model you then own and must retrain, re-evaluate and host.

8. What is the difference between a facts problem and a behaviour problem? → A facts problem is the model not having information — it never saw your policy, your catalogue, your protocol. That is retrieved. A behaviour problem is the model getting format, tone or domain style wrong. That is tuned. They look similar in a question stem and have opposite answers.

9. Why is distillation justified by volume rather than by capability? → Because it spends up front to make each request cheaper afterwards, and that saving only pays back at volume. It compresses behaviour a larger model already has — it adds nothing new. "The model cannot do X" is not a distillation problem; "the model does X well and we cannot afford it at this volume" is.

10. A catalogue changes weekly. Why is weekly fine-tuning the wrong answer? → Because it answers a facts problem with training. Even if every job finished on time, the model would answer from a snapshot averaging half a week old while stock changes hourly, and each cycle buys a permanent maintenance obligation. A scheduling fix cannot repair a category error.

11. Answers are factually right but do not sound like the organisation. Prompting has failed. What now, and why not RAG? → Fine-tune on examples written in the house style, because style is behaviour rather than fact. Retrieving the style guide would deliver a description of the style — the same instruction prompting has already proved insufficient with.

12. When is pre-training the reasonable choice? → When no existing model fits the domain at all. It is justified by the absence of an alternative, not by the difficulty of the problem or by a fine-tuned model missing its quality bar. A question offering it is usually testing whether thoroughness will be mistaken for correctness.

The Five-Beat Answer

The core question this chapter prepares you for: "How would you make a foundation model answer questions about our business, and how far would you go?"

Five beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.

  1. Diagnose — say first whether the gap is facts the model never saw or behaviour it gets wrong. Naming this before proposing anything is what separates a designed answer from a remembered one.
  2. Retrieve — for a facts problem, describe RAG: fetch from your sources at request time, place it in the prompt, answer from it, cite it. Say explicitly that the model is unchanged, because that is what makes it immediate.
  3. Store — name where the embeddings live and why: OpenSearch Service for search-first at scale, Aurora or RDS for PostgreSQL when the data is already relational there, Neptune when relationships decide the answer. Say that the existing estate usually chooses.
  4. Ladder — order the five approaches by cost and give the condition on each climb. Name distillation's as volume rather than capability; it is the one most often stated wrongly.
  5. Stop — state the cheapest rung that meets the requirement and say what the next one would have cost. Then add that grounding is not evaluation: a citation proves provenance, not correctness.

A strong answer diagnoses before it designs, and names a rung it deliberately did not climb. A weak answer proposes fine-tuning.

Why This Helps You

On the job: the single most expensive avoidable mistake in this area is commissioning a custom model for a problem retrieval would have solved. It looks thorough at the point of approval, and it converts a content problem into a permanent obligation to retrain, re-evaluate and host — while staying, by construction, permanently out of date. The diagnosis in beat one is what prevents it.

In interviews: "when would you fine-tune instead of using RAG?" is a standard screening question, and most candidates answer it as a preference. The strong answer is structural — facts versus behaviour, plus the cost the climb commits you to — and mentioning that a citation proves provenance rather than correctness marks out someone who has actually run one of these systems.

On the exam: Domain 3 is 28% of scored content, the largest on the exam, and these three objectives are the design-decision core of it. The highest-value habit is refusing to climb: when two options both work, the cheaper one is the answer, and the more thorough-sounding option is the trap.

Chapter Checklist

  • I can define RAG without describing training, and say what it leaves unchanged
  • I can walk the request flow from question to grounded, cited answer
  • I can separate ingestion (on document change) from retrieval (per request)
  • I can name business applications of RAG and say why citation is the business feature
  • I can say what Amazon Bedrock Knowledge Bases removes from a build
  • I can name the four vector stores and the scenario signal that selects each
  • I can explain why Amazon Aurora appears in this objective, and how it differs from RDS for PostgreSQL
  • I can order the five customization approaches by cost and state each climb condition
  • I can explain why distillation is justified by volume rather than capability
  • I can diagnose whether a scenario describes missing facts or wrong behaviour
  • I can say why a grounded answer is not automatically a correct one

After the Chapter

  1. Complete student/project.md — parts 34-36 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one.
  2. Take student/quiz.md closed-book, then review the reasoning for every question you guessed, including the ones you got right. Pay particular attention to questions 7 and 9 — they are deliberate mirror images, and missing both means the facts-versus-behaviour diagnosis has not landed.
  3. Open the official v1.1 exam guide's Domain 3 page and confirm you can attach a concept from this chapter to the third, fourth and fifth bullets under Task Statement 3.1. Then find Amazon Aurora on the In-Scope AWS Services page and note which category it sits in.
  4. Next: Chapter 10 — Prompt Engineering (Domain 3, Task 3.2). This chapter named the bottom rung of the ladder but did not teach it. Chapter 10 is that rung in full: the constructs of a prompt, the techniques that make in-context learning work, and the risks — exposure, poisoning, hijacking and jailbreaking — that arrive as soon as retrieved and user-supplied text share a prompt.

Chapter 9 quiz

13 questions on this chapter, marked instantly, with an explanation for every answer.

1. Which statement best describes retrieval augmented generation?
2. A team reports that their new RAG system means "the model has now learned our documentation." What is wrong with that description?
3. In an Amazon Bedrock Knowledge Base, when are the source documents embedded?
4. A fraud team wants to answer questions in which the connections between accounts, devices and payments are part of the answer, not merely similarity of text. Which vector store fits?
5. Which statement about Amazon Aurora is accurate in the context of Objective 3.1.4?
6. Which ordering of the five customization approaches runs from least to most expensive?
7. A retailer proposes fine-tuning a model weekly on its catalogue so that the assistant always knows current stock levels and pricing. What is the strongest objection?
8. An assistant performs well, but its per-request cost is unsustainable at the firm's volume. Which approach targets that specific problem, and why?
9. An assistant's answers are factually correct but do not follow the organisation's house style. Prompting has been tried extensively and has not fixed it. What is the appropriate next step?
10. A team wants grounded, cited answers over its own documents but does not want to build or operate a retrieval pipeline. Which service is named in Objective 3.1.3 for exactly this?
11. Which two of the following leave the foundation model's weights unchanged? (Select two.)
12. A grounded assistant cites a real internal page for every answer, yet a review finds several of its answers are wrong. What is the most likely explanation?
13. Under what circumstance does pre-training — the most expensive rung of the ladder — become the reasonable choice?