Free live cohort on Google Meet — register your interest →

Chapter 11 · Domain 3 · 28% of the exam

Training and Fine-Tuning Foundation Models

29 min read · Chapter 11 of 17

On this page
  1. Certification Blueprint
  2. What This Chapter Covers
  3. The Four Key Elements of Training
  4. Fine-Tuning or Continuous Pre-Training?
  5. Methods for Fine-Tuning
  6. Preparing the Data
  7. Decision Rules and Exam Signals
  8. Distractor Patterns
  9. Scenario Walkthrough
  10. Key Concepts
  11. Revision Flashcards
  12. The Five-Beat Answer
  13. Why This Helps You
  14. Chapter Checklist
  15. After the Chapter

Certification Blueprint

Field Coverage
Exam AWS Certified AI Practitioner (AIF-C01), exam guide v1.1
Domain Content Domain 3 — Applications of Foundation Models
Exam weight 28% of scored content — the largest domain on the exam
Task statement 3.3 Describe the training and fine-tuning process for FMs
Objectives 3.3.1 key elements of training an FM · 3.3.2 methods for fine-tuning an FM · 3.3.3 how to prepare data to fine-tune an FM

What This Chapter Covers

Chapter 09 gave you a ladder. In-context learning, RAG, fine-tuning, distillation, pre-training — ordered by cost, each with a condition that must be true before you climb. That chapter deliberately refused to explain the top rungs. It only ranked them.

This chapter opens them.

There is one question that separates every method in this task statement, and it is not about expense or sophistication:

What data does it consume, and what does it change?

Answer those two things and every method here becomes distinguishable. Miss them and the four names blur into "training, but more so" — which is exactly the confusion the exam's distractors are built from.

Method What data it consumes What it changes
Pre-training An enormous general corpus, unlabelled Creates the weights from nothing
Continuous pre-training Your raw domain text, unlabelled Extends existing weights with new subject matter
Fine-tuning Your example pairs, labelled Adjusts existing weights toward a task or behaviour
Distillation The larger model's own outputs Produces a different, smaller model

The labelled/unlabelled split in the middle column is the highest-value distinction in this task statement. It is what separates continuous pre-training from fine-tuning, and those two are the pair the exam most often asks you to tell apart.

The Four Key Elements of Training

Objective 3.3.1 names four operations. They are not four intensities of the same thing — they are four different jobs, and only two of them are usually available to an organisation.

Left-to-right flow showing pre-training building a model from scratch on an enormous unlabelled general corpus, then continuous pre-training extending an existing model with raw unlabelled domain text, then fine-tuning adapting the model on labelled example pairs, then distillation compressing it into a smaller model trained on the teacher's own outputs, with dashed skip arrows showing that most organisations start at fine-tuning because a model already exists, skip continuous pre-training when vocabulary is adequate, and skip distillation when per-request cost is acceptable, all converging on deployment

What to remember from this diagram: follow the dashed arrows, not the solid ones. The solid path is the full sequence; the dashed ones are the skips, and almost every real project takes them. A scenario that describes an organisation walking the whole solid path from the left is describing something very few organisations ever do.

Pre-training

Pre-training is building a foundation model from scratch — starting with no weights at all and learning language from an enormous general corpus.

What it consumes Vast quantities of general text, unlabelled — the model learns by predicting what comes next
What it produces A model that did not previously exist
What it costs Enormous data, compute and specialist expertise
Who does it Model providers, overwhelmingly — not their customers
The scenario that names it No existing model fits the domain at all

On the exam, pre-training is almost never the correct answer, and it is offered often. It is the option that sounds most rigorous, which is precisely what makes it a reliable trap. When you see it, ask what it would actually require — and whether the scenario has established that no existing model could work.

Continuous pre-training

Continuous pre-training is taking an existing foundation model and continuing to train it on your own raw text, so it absorbs the vocabulary and conventions of your domain.

What it consumes Your raw, unlabelled domain text — documents, transcripts, filings, manuals
What it produces The same model, extended with domain familiarity
What it does not need Labels. There are no correct answers to supply
The scenario that names it The model does not speak the language of the field

This is the method most often missed, because it sits between the two everyone knows. It uses pre-training's mechanism (raw text, no labels) but starts from an existing model like fine-tuning does.

⚠️ A genuine quirk of the v1.1 exam guide worth knowing: continuous pre-training is listed twice — as a key element of training in Objective 3.3.1, and as a method for fine-tuning an FM in Objective 3.3.2. That is not an error in your reading. It is a real boundary case, and the guide places it on both sides. If a question treats it as a fine-tuning method, that is consistent with the guide; if another treats it as a distinct training element, so is that.

Fine-tuning

Fine-tuning is adapting an existing model's weights using your own labelled examples, so it performs a specific task or behaves in a specific way.

What it consumes Labelled example pairs — an input and the response you wanted
What it produces The same model, adjusted toward your task
What it demands Someone to produce the labels, which is the real cost
The scenario that names it Behaviour, format or task performance is wrong, and prompting has genuinely failed

Distillation

Distillation is training a smaller student model to reproduce a larger teacher model's behaviour, so that requests can be served more cheaply.

What it consumes The teacher model's own outputs — generated, not hand-labelled
What it produces A different, smaller model — this is the only method here that does
What it buys Lower cost and latency per request
What it risks The student loses capability the teacher had
The scenario that names it The model works well and cannot be afforded at the required volume

Distillation is the only one of the four justified by volume rather than by capability. Chapter 09 made this point about cost; it is worth repeating as a mechanism. Distillation compresses behaviour that already exists. It cannot add any.

Fine-Tuning or Continuous Pre-Training?

This is the pair the exam most reliably asks you to separate, and one question does it.

Decision flow beginning with what the model is getting wrong, asking first whether it does not know the vocabulary and conventions of the domain, routing yes to continuous pre-training which feeds raw domain text with no labels required and teaches the language, and no to a second question asking whether the model understands the domain but responds in the wrong form format or behaviour, routing yes to fine-tuning which feeds prompt and response pairs requiring labels and teaches the task, and no to neither, returning to prompting or retrieval

What to remember from this diagram: the split is language versus task. Continuous pre-training teaches the model what the words mean in your field; fine-tuning teaches it what to do when asked. And the practical tell follows from that: if the scenario describes a pile of documents, that is continuous pre-training's input. If it describes examples of good answers, that is fine-tuning's.

Continuous pre-training Fine-tuning
Data required Raw domain text Input-and-response pairs
Labels needed No Yes
What it teaches The domain's language The required task or behaviour
Typical source Existing corpus you already hold Must usually be created
The complaint it answers "It does not understand our terminology" "It understands, but answers wrongly"

The labelling requirement is the practical difference, not the algorithm. An organisation with ten years of unlabelled filings can begin continuous pre-training immediately. The same organisation may need months to produce fine-tuning pairs — because someone has to write the correct answers.

Methods for Fine-Tuning

Objective 3.3.2 names four methods. Three are genuinely distinct approaches; the fourth is continuous pre-training, appearing here for the second time.

Decision flow starting from fine-tuning being justified and asking which method, first asking whether the model should follow instructions and answer in a required style routing yes to instruction tuning using pairs of instruction and desired response, then whether it needs terminology and conventions of a specialist field routing yes to domain adaptation using material from that field, then whether there is a related task whose learned representations can be reused routing yes to transfer learning, and otherwise advising to re-check the diagnosis because the problem may not be a tuning problem

What to remember from this diagram: the last box is the one that matters most. "Re-check the diagnosis" is a real exam answer far more often than any of the three methods above it — because most scenarios that look like tuning problems are prompting problems that were never properly tried.

Method What it is The scenario that selects it
Instruction tuning Training on pairs of instruction and desired response, so the model reliably follows directions and answers in the required form Answers are correct but ignore the instruction's form, tone or structure
Domain adaptation Adapting a general model to a specialist field's terminology, conventions and expectations The field has its own vocabulary the general model handles badly
Transfer learning Reusing what a model learned on one task as the starting point for a related task A related, already-learned capability exists to build on
Continuous pre-training Continuing to train on raw domain text without labels A large body of domain text exists and the gap is language rather than task

Transfer learning is the principle, not a separate product

This is worth stating plainly because the exam's phrasing can make it sound like a fourth tool sitting beside the others.

Transfer learning is the general idea that knowledge learned for one task can be reused for another. Fine-tuning a foundation model is an application of transfer learning — the model transferred what it learned in pre-training to your task. When the guide lists it as a method, it is naming the principle. A question describing "reusing a model trained on one task as the starting point for a related one" is describing transfer learning by definition.

Instruction tuning versus domain adaptation

The scenario says Reading Method
"It gives good information but ignores the format we asked for" Instruction-following failure Instruction tuning
"It does not know what our industry's terms mean" Vocabulary gap Domain adaptation or continuous pre-training
"It answers legal questions like a general assistant, not a lawyer" Domain conventions Domain adaptation
"We have a model for a near-identical task already" Reuse an existing capability Transfer learning
"The tone is wrong and we have never written a prompt for it" Nothing has been tried Neither — prompt first

Preparing the Data

Objective 3.3.3 is the objective candidates under-prepare, and it carries as many marks as the other two. It names six things: data curation, governance, size, labeling, representativeness, and RLHF.

The training is not the project. The data preparation is the project. A fine-tuning job is hours of compute against months of assembling, cleaning, labelling and checking what goes into it. A scenario that treats the training run as the hard part has usually mislocated the difficulty.

Left-to-right pipeline showing curation selecting and cleaning what goes in, then governance covering consent licensing privacy and retention, then size ensuring enough examples for the task, then labeling marking the correct response for each example, then representativeness asking whether the data reflects real users and cases, arriving at ready to fine-tune, with a dashed feedback arrow returning from representativeness to curation when gaps are found

What to remember from this diagram: the dashed arrow going backwards is the honest part. Representativeness is checked near the end and routinely sends you back to curation — which is why data preparation is iterative and why estimating it as a single up-front step is how these projects overrun.

Step What it means The failure it prevents
Curation Selecting and cleaning what goes in; removing duplicates, errors and irrelevant material Training on noise, and teaching the model mistakes that were in the source
Governance Consent, licensing, privacy, and permission to use the data for this purpose Training on data you had no right to use — which cannot be undone once it is in the weights
Size Enough examples for the task; more is not automatically better if quality falls A model that has not seen enough of the task to generalise
Labeling Marking the correct response for each example Fine-tuning without a target; this is the step that makes fine-tuning expensive
Representativeness Does the data reflect the real range of users, cases and conditions? A model that works for the majority case and fails everyone outside it
RLHF Using ranked human judgement to align the model with what people actually prefer A model that is technically correct and unhelpful

Governance cannot be fixed afterwards

Of the six, governance is the one with no remedy. Retrieved content can be removed from an index in minutes. Training data cannot be removed from weights. Once a model has been trained on material you had no licence, consent or lawful basis to use, the model itself is the problem, and the remedy is retraining — which means the cost of the mistake is the cost of the whole job again.

This is why governance appears in a technical objective about data preparation rather than only in Domain 5. It is a precondition, not a review step.

Representativeness is the one that causes harm

Size and representativeness are easily confused, and the exam separates them.

  • Size asks how much.
  • Representativeness asks of what.

A large dataset drawn entirely from one region, one customer segment or one time period is big and unrepresentative at the same time. The resulting model performs well on average and fails specific groups — which is where fine-tuning turns into a fairness problem. Chapter 14 examines that consequence under responsible AI; this objective examines the preparation step that prevents it.

Reinforcement learning from human feedback

RLHF is using human judgement about which model outputs are better to train the model toward what people actually prefer.

Left-to-right loop in which the model produces several candidate responses, humans rank them best to worst, a reward model learns what humans preferred, the model is updated to score higher against the reward model, and the cycle returns to the model producing candidates, with a branch noting that the human ranking step is the cost and the bottleneck of RLHF

What to remember from this diagram: humans rank outputs; they do not write them. That is the detail most often stated wrongly. RLHF does not require someone to author the perfect answer — it requires someone to say which of several candidates is better, which is a far cheaper judgement to make and a far harder one to scale.

What humans supply Preference — a ranking of candidate outputs
What humans do not supply The ideal answer text itself
What the reward model does Learns to predict human preference, so it can score outputs at scale
What it is for Alignment with what people find helpful — not factual accuracy
Why it appears in Objective 3.3.3 It is a data preparation approach: the preferences are the data

RLHF fixes helpfulness, not correctness. A model aligned to human preference produces answers people like. Whether those answers are true is a different question, examined in Chapter 12.

Decision Rules and Exam Signals

Rule 1 — ask what data it eats and what it changes. Those two properties identify every method in this task statement.

Rule 2 — labels are the dividing line. Raw text means continuous pre-training. Input-and-response pairs mean fine-tuning.

Rule 3 — distillation makes a new, smaller model; nothing else here does. The other three modify or create weights for the same model.

Rule 4 — pre-training is almost never the answer. It is justified by the absence of any suitable model, never by the difficulty of the problem.

Rule 5 — "re-check the diagnosis" beats every tuning method when the scenario has not established that prompting failed.

Rule 6 — the data preparation is the project. If a scenario asks what will take the longest, it is the labelling and curation, not the training run.

Rule 7 — governance has no undo. Data in weights cannot be withdrawn; only retraining removes it.

Rule 8 — size and representativeness are different questions. How much, versus of what.

Rule 9 — in RLHF humans rank, they do not write. Preference data, not model answers.

Rule 10 — RLHF aligns helpfulness, not truth. Accuracy is Chapter 12's subject.

Distractor Patterns

Pattern What it looks like How to defuse it
Fine-tuning offered with unlabelled data "Fine-tune the model on the company's document archive" Fine-tuning needs labelled pairs; a raw archive is continuous pre-training's input
Continuous pre-training described as needing labels "Continue pre-training on labelled question-and-answer pairs" The labels are the giveaway — that describes fine-tuning
Pre-training as the thorough answer Building from scratch for a domain problem Almost never correct; ask what it would actually require
Distillation to add capability Offered to make the model do something new Distillation compresses existing behaviour; it adds nothing
Tuning before prompting A training method offered with no evidence prompting was tried The cheapest approach that meets the requirement wins
RLHF as answer-writing "Experts write ideal responses for the model to learn from" That is supervised fine-tuning; RLHF collects rankings
RLHF for factual accuracy Offered to stop the model being wrong It aligns preference, not truth
Size offered as the fix for representativeness "Collect more data" for a model failing one user group More of the same skew is still skewed
Governance treated as a later review Compliance sign-off scheduled after training Training on unlicensed data cannot be undone without retraining
Transfer learning as a distinct product Presented as a fourth tool beside the others It is the underlying principle; fine-tuning is an instance of it

The first two are a matched pair and the most reliable marks in this task statement. Both are answered by looking at one thing: is the data labelled?

Scenario Walkthrough

A national insurer wants a claims assistant. Its models handle general English well but consistently misread the insurer's policy terminology, in which ordinary words carry specific contractual meanings. The insurer holds twenty years of claims correspondence and adjuster reports, none of it annotated. It also wants answers structured to a fixed internal template; prompting has been tried at length and the structure is still inconsistent. Legal has not yet confirmed whether historical correspondence may be used for model training. Most historical claims come from two regions where the insurer has operated longest.

Requirement Reading Decision
Misreads domain terminology A language gap, not a task gap Continuous pre-training
Twenty years of unannotated text Raw, unlabelled — exactly its input Confirms continuous pre-training; rules out fine-tuning for this part
Fixed template, prompting genuinely failed A behaviour gap, and the precondition is met Instruction tuning, needing labelled pairs
Legal has not confirmed permission Governance, and it is a blocker Resolve before training — it cannot be undone afterwards
Claims concentrated in two regions Representativeness, not size Check coverage before training, or the model fails other regions

Five requirements, two different training methods, and one that stops the project. The governance row is the test. A candidate optimising the technical answer will sequence the two training methods correctly and start immediately — and training on correspondence the insurer may not lawfully use puts the unusable data somewhere it cannot be removed from.

Note also what the last row is not: the fix for a two-region skew is not "collect more claims." It is collecting claims from the other regions.

Key Concepts

Term Definition
Pre-training Building a foundation model from scratch on an enormous unlabelled general corpus; produces a model that did not previously exist
Continuous pre-training Continuing to train an existing model on raw, unlabelled domain text so it absorbs a field's vocabulary and conventions; listed in the guide under both Objective 3.3.1 and Objective 3.3.2
Fine-tuning Adapting an existing model's weights using labelled input-and-response pairs so it performs a specific task or behaves in a specific way
Distillation Training a smaller student model to reproduce a larger teacher model's behaviour; the only method here that produces a different model, and justified by per-request cost at volume
Instruction tuning Fine-tuning on pairs of instruction and desired response so the model reliably follows directions and answers in the required form
Domain adaptation Adapting a general model to a specialist field's terminology, conventions and expectations
Transfer learning The principle that knowledge learned for one task can be reused as the starting point for a related task; fine-tuning a foundation model is an instance of it
Data curation Selecting and cleaning what enters a training set — removing duplicates, errors and irrelevant material
Data governance Establishing consent, licensing, privacy and lawful basis for using data in training; the one preparation step with no remedy after the fact
Representativeness Whether training data reflects the real range of users, cases and conditions; distinct from size, and the property whose absence produces a model that fails specific groups
Labeling Marking the correct response for each training example; the step that makes fine-tuning expensive, because it requires human effort per example
Reinforcement learning from human feedback (RLHF) Training a model toward human preference by having people rank candidate outputs, learning a reward model from those rankings, and updating the model to score well against it
Reward model A model trained on human preference rankings that can then score outputs at scale, standing in for the human judgement it learned from
Catastrophic forgetting The loss of previously held general capability when a model is trained heavily on narrow new data

Revision Flashcards

Say the answer aloud before revealing it.

1. Name the two questions that identify every training method in this task statement. → What data does it consume, and what does it change? Pre-training eats an enormous unlabelled general corpus and creates weights from nothing; continuous pre-training eats your raw unlabelled domain text and extends existing weights; fine-tuning eats labelled example pairs and adjusts existing weights; distillation eats the teacher model's own outputs and produces a different, smaller model.

2. What single property separates continuous pre-training from fine-tuning? → Labels. Continuous pre-training takes raw, unlabelled domain text and teaches the model the field's language. Fine-tuning takes labelled input-and-response pairs and teaches the model a task or behaviour. If a scenario describes an archive of documents, that is continuous pre-training's input; if it describes examples of good answers, that is fine-tuning's.

3. Why is pre-training almost never the right exam answer, and when is it right? → Because it is justified only by the absence of any suitable existing model, never by the difficulty of the problem. It requires enormous data, compute and specialist expertise, and is done overwhelmingly by model providers rather than their customers. It is offered often precisely because it sounds the most thorough.

4. What does distillation do that none of the other three methods do? → It produces a different, smaller model. The other three either create or modify the weights of the model in question; distillation trains a new student model to reproduce a teacher's behaviour, and it is justified by per-request cost at volume rather than by any capability it adds. It cannot add capability at all.

5. Continuous pre-training appears in two different objectives. Is that a mistake? → No. The v1.1 guide lists it as a key element of training in Objective 3.3.1 and again as a method for fine-tuning an FM in Objective 3.3.2. It is a genuine boundary case: it uses pre-training's mechanism of raw unlabelled text but starts from an existing model the way fine-tuning does. Both framings are consistent with the guide.

6. Distinguish instruction tuning from domain adaptation. → Instruction tuning trains on instruction-and-response pairs so the model follows directions and answers in the required form — it fixes a model that gives good information in the wrong shape. Domain adaptation adapts the model to a specialist field's terminology and conventions — it fixes a model that does not know what the field's terms mean.

7. What is transfer learning, and why does it feel like an odd item on the list? → It is the general principle that knowledge learned for one task can be reused as the starting point for a related one. It feels odd beside the others because it is not a separate product — fine-tuning a foundation model is itself an application of transfer learning, since the model transfers what it learned during pre-training to your task.

8. Name the six things Objective 3.3.3 lists for preparing fine-tuning data. → Data curation, governance, size, labeling, representativeness, and reinforcement learning from human feedback. The objective carries as many marks as the other two in this task statement and is the one candidates under-prepare because it sounds administrative.

9. Why is governance the preparation step with no remedy? → Because training data cannot be removed from weights. Content in a retrieval index can be deleted in minutes; material a model was trained on is embedded in the model itself, and the only way to remove it is to retrain — meaning the cost of the mistake is the cost of the entire job again. That is why it is a precondition rather than a review step.

10. What is the difference between size and representativeness? → Size asks how much data; representativeness asks what it is data of. A dataset drawn entirely from one region or customer segment can be very large and completely unrepresentative. The resulting model performs well on average and fails everyone outside the majority case, and the fix is not more of the same data but data from the missing groups.

11. In RLHF, what exactly do the humans provide? → Rankings. They compare candidate outputs the model produced and say which is better. They do not write ideal answers — that would be supervised fine-tuning. Those rankings train a reward model, which can then score outputs at scale, and the model is updated to score well against it.

12. What does RLHF align, and what does it not? → It aligns the model with what humans find helpful and prefer. It does not make the model factually correct. A model tuned to preference produces answers people like; whether they are true is a separate question, and evaluating that is Chapter 12's subject.

The Five-Beat Answer

The core question this chapter prepares you for: "We want to train a model on our own data. How would you approach it?"

Five beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.

  1. Diagnose — say what is actually wrong before naming a method. Is the gap the domain's language, the required task or behaviour, or per-request cost? Each points at a different method, and stating this first is what separates a designed answer from a remembered one.
  2. Check the precondition — say whether prompting and retrieval have genuinely been tried. If they have not, that is the answer, and any training method is premature.
  3. Match the method to the data you have — raw unlabelled text supports continuous pre-training today; fine-tuning needs labelled pairs that usually do not exist yet. Say which you hold, because it decides what is actually available rather than merely appropriate.
  4. Name the data work — curation, governance, size, labelling and representativeness, and say that this is the majority of the effort. Put governance first and say why: it cannot be fixed after training.
  5. Say how you would know it worked — and hand off to evaluation. Naming that you would evaluate, rather than assuming a completed training run is a success, is the beat that most answers omit.

A strong answer diagnoses before it selects, and treats the data as the project. A weak answer names a training method in the first sentence.

Why This Helps You

On the job: the most common way these projects fail is not a bad training run. It is discovering at the labelling stage that the data required does not exist, or discovering after training that the organisation never had permission to use it. Both are visible in advance, and both are found by the diagnosis in beats one and four rather than by any technical step.

In interviews: "when would you fine-tune versus continue pre-training?" separates candidates who have run one of these from those who have read about them. The strong answer is about the data — labelled pairs versus raw text — and about which one the organisation actually has. Mentioning that humans rank rather than write in RLHF is another reliable marker.

On the exam: Domain 3 is 28% of scored content, and this task statement's three objectives split their marks fairly evenly — which means the data preparation objective is worth as much as the two about methods. Candidates who study the method names and skip the preparation step lose a third of the available marks to an objective they considered administrative.

Chapter Checklist

  • I can name the four key elements of training and say what data each consumes
  • I can say what each of the four changes, and which one produces a different model
  • I can separate continuous pre-training from fine-tuning by the labelling requirement
  • I can explain why continuous pre-training appears in two objectives
  • I can say when pre-training is genuinely the right choice, and why it usually is not
  • I can explain why distillation cannot add capability
  • I can distinguish instruction tuning, domain adaptation and transfer learning
  • I can name all six data preparation elements from Objective 3.3.3
  • I can explain why governance cannot be remedied after training
  • I can distinguish size from representativeness and say what each failure looks like
  • I can describe the RLHF loop and say what humans actually supply
  • I can say what RLHF aligns and what it does not

After the Chapter

  1. Complete student/project.md — parts 40-42 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one.
  2. Take student/quiz.md closed-book, then review the reasoning for every question you guessed, including the ones you got right. Pay particular attention to questions 3 and 4 — they are deliberate mirror images, and missing both means the labelled-versus-unlabelled split has not landed.
  3. Open the official v1.1 exam guide's Domain 3 page and confirm you can attach a concept from this chapter to each of the three bullets under Task Statement 3.3. Note where continuous pre-training appears, and satisfy yourself that it is listed under both 3.3.1 and 3.3.2.
  4. Next: Chapter 12 — Evaluating Foundation Models (Domain 3, Task 3.4). This chapter produced a model; it gave you no way to know whether it is any good. Chapter 12 is that: human-in-the-loop evaluation, benchmark datasets, Amazon Bedrock Model Evaluation, the metric families including ROUGE, BLEU and BERTScore, and the business alignment metrics that decide whether the model earned its place at all.

Chapter 11 quiz

13 questions on this chapter, marked instantly, with an explanation for every answer.

1. Which pair of properties most reliably distinguishes the key elements of training a foundation model from one another?
2. A manufacturer holds thirty years of maintenance logs, engineering reports and service bulletins. None of it is annotated. The model handles general English well but misreads the firm's technical vocabulary. Which approach fits what they have and what they need?
3. A support team has assembled eight thousand pairs of customer question and approved agent reply. They want the model to answer new questions the way those replies do. Which approach does this data support?
4. A team proposes "fine-tuning the model on our document archive" — a large collection of unannotated internal reports. What is the strongest technical objection?
5. An assistant performs its task well, but per-request cost at the organisation's volume is unsustainable. Which approach targets that, and what is its known risk?
6. Under what circumstance is pre-training a foundation model from scratch the reasonable choice?
7. A model returns accurate information about a specialist field but consistently ignores the required response structure, despite the structure being stated clearly in the prompt. Prompting has been tried extensively. Which fine-tuning method fits?
8. Which statement best describes transfer learning as the exam guide uses the term?
9. In reinforcement learning from human feedback, what do the human participants actually provide?
10. An organisation adopts RLHF hoping it will stop the model stating incorrect facts. What is wrong with that expectation?
11. Which two of the following require **labelled** training data? (Select two.)
12. A model has been trained on historical customer correspondence. After deployment, legal review establishes that the organisation never had permission to use that correspondence for training. Why is this materially worse than the equivalent problem in a retrieval system?
13. A model performs well overall but consistently underserves customers in rural areas. The team proposes collecting substantially more training data. What is the flaw in that plan as stated?