Imagine an interviewer asks how to adapt Amazon Nova 2 Lite for a specialist domain: should you feed it a vast library of raw documents, or show it carefully prepared examples of the answers you want? Those choices correspond to continued pre-training and supervised fine-tuning, and they solve different problems. Continued pre-training develops deeper familiarity with a domain from large volumes of unlabeled text, while supervised fine-tuning teaches tasks and behaviors from labeled prompt-response examples; choosing correctly depends on whether the gap is primarily missing domain fluency or unclear behavior.
The strongest answer separates knowledge acquisition from behavior shaping
Interview question: What is the difference between fine-tuning and continued pre-training a foundation model on AWS, and what data does each need?
Model answer: Continued pre-training, or CPT, extends a foundation model’s pre-training by exposing it to additional raw, unlabeled text from a particular domain or collection. Its purpose is to help the model acquire deeper domain knowledge, terminology, writing patterns, and familiarity with particular content types.
Supervised fine-tuning, or SFT, instead trains the model on labeled input-output pairs. Each pair demonstrates a prompt and the response that the model should learn to produce, allowing the model to align with a specific task, instruction, workflow, format, tone, or style.
The simplest distinction is therefore knowledge acquisition versus behavior shaping. CPT answers, “What body of language and subject matter should the model become more fluent in?” SFT answers, “Given an input, what should the model do?”
The data requirements follow directly from that distinction. CPT needs large volumes of domain-specific raw documents, potentially tens of billions of tokens. SFT needs high-quality labeled examples in which each input is paired with the desired output. CPT generally needs subsequent instruction-tuning stages so that the model can use its newly acquired knowledge to complete useful tasks.
Think of it as preparing a specialist for work. CPT resembles giving the specialist an extensive professional library to absorb, while SFT resembles showing the specialist worked examples of how to respond to particular assignments. The analogy distinguishes broad exposure from explicit demonstrations; it does not imply that either method automatically guarantees correct answers.
Continued pre-training builds domain fluency, while fine-tuning demonstrates desired behavior
The mistake weak candidates make: In practical interview terms, a weak answer often treats both methods as interchangeable ways to “add company data.” That misses the decisive difference between learning from unlabeled documents and learning from labeled demonstrations.
The right choice begins with the gap you need to close
Follow-up question 1: When should you choose supervised fine-tuning?
Model answer: Choose SFT when you can express the desired behavior through high-quality input-output examples. It is suited to teaching a specific response format, tone, or style, as well as domain-specific instructions and workflows. It can also adapt the model for multimodal tasks involving text, images, video, or tool calling.
For example, suppose the model already understands the vocabulary in technical service records but does not consistently transform a record into the required structured response. A training set could pair representative records with correctly formatted outputs. That example illustrates SFT because every input is accompanied by a target response that demonstrates the expected task.
SFT is also the clearer choice when success can be shown directly: “For an input like this, produce an output like that.” This selection rule is a practical judgment, but it follows the structure of the required training data.
Supervised fine-tuning starts with examples that show the required response
The mistake weak candidates make: A weak answer may say that SFT merely gives the model more subject matter. In fact, its defining input is a set of labeled prompt-response demonstrations, and its stated purpose is alignment with tasks, instructions, or desired behaviors.
Raw domain corpora point toward continued pre-training
Follow-up question 2: When should you choose continued pre-training?
Model answer: Choose CPT when you have a very large volume of domain-specific text and want the model to develop deeper knowledge and more native fluency in that domain. Suitable material can include legal documents, medical literature, technical documentation, or proprietary business content. The documents do not need prompt-response labels because CPT trains on raw text.
CPT is particularly valuable at a scale of tens of billions of tokens. It can help the model learn specialized terminology, characteristic writing patterns, new subject matter, and particular content types. That makes CPT a different commitment from assembling a comparatively targeted collection of demonstrations: its stated use case assumes a large corpus rather than a list of correct answers.
For example, a collection of specialist manuals may contain the language and concepts the model should absorb without specifying a task for each passage. That is CPT-shaped data. By contrast, pairs such as “manual excerpt plus question” and “approved answer” are SFT-shaped data because they demonstrate how the model should respond.
Think of CPT as immersion in a domain’s literature. Immersion can build familiarity with the domain, but it does not by itself specify the useful task the model should perform with that familiarity.
The presence or absence of target responses determines the training-data shape
The mistake weak candidates make: A weak answer may recommend CPT simply because some proprietary documents exist. The stronger answer checks whether there is enough domain text for the documented large-volume use case and whether the real gap is domain fluency rather than response behavior.
Continued pre-training may be only the first adaptation stage
Follow-up question 3: Does continued pre-training make the model ready for a specific task?
Model answer: Not necessarily. After CPT, the model generally needs additional instruction-tuning stages so that it can use the newly acquired knowledge and complete useful tasks. CPT can deepen knowledge of a domain, but its raw documents do not directly provide the labeled demonstrations that define a desired response.
A sensible design is to treat CPT and SFT as potentially complementary rather than mutually exclusive. CPT can first expose the model to a large specialist corpus, after which instruction tuning can demonstrate how that knowledge should be applied. The first stage develops domain familiarity; the later stage shapes task performance.
For example, raw technical documentation could support CPT so the model encounters specialized terms and writing patterns. Labeled examples could then show how to answer a diagnostic question, follow a domain-specific workflow, or return a required response format. The example demonstrates why “knowing the material” and “performing the task” are separate adaptation goals.
Domain adaptation can continue from raw-text training into instruction-focused training
The mistake weak candidates make: In interview terms, weak candidates often imply that pouring raw documents into CPT automatically teaches question answering, formatting, or workflow compliance. The source makes the limitation explicit: additional instruction tuning is generally needed after CPT.
Data quality and data shape are different requirements
Follow-up question 4: What exactly should the training dataset look like for each method?
Model answer: A CPT dataset should be a large corpus of relevant, unlabeled domain text. Its useful signal lies in the documents themselves: their subject matter, terminology, writing patterns, and content types. The documented examples include legal, medical, technical, and proprietary business text.
An SFT dataset should contain labeled input-output pairs that demonstrate the correct behavior. Inputs are prompts or task examples, while outputs are the responses the model is meant to learn from. Those pairs can demonstrate a format, tone, style, instruction, workflow, or multimodal task.
“Unlabeled” does not mean “unrelated,” and “labeled” does not mean “automatically useful.” In practical terms, CPT material should represent the domain the model needs to absorb, while SFT examples should clearly represent the behavior it needs to reproduce. That is a judgment about dataset design rather than a separate sourced requirement.
For example, a folder of domain documents does not become an SFT dataset merely because the documents have filenames or categories. SFT requires the stronger relationship of an input paired with its intended output. Conversely, converting every raw passage into a demonstration is unnecessary for CPT because CPT is specifically designed to train on raw documents.
Choose the method by asking whether the dataset contains desired answers
The mistake weak candidates make: A weak answer discusses only the amount of data. Volume matters to CPT’s documented use case, while SFT is defined by high-quality labeled pairs; the structure and purpose of the data are therefore as important as its quantity.
Nova 2 Lite supports both paths, but model selection remains a tradeoff
Follow-up question 5: How does this distinction apply to Amazon Nova 2?
Model answer: Amazon Nova 2 is the current model generation in the question, and Nova 2 Lite supports both CPT and SFT. The documentation also lists Nova 1.0 Micro, Lite, and Pro for both techniques, but that is the older model generation.
Nova 2.0 is the documented choice when the use case needs enhanced reasoning with an explicit reasoning mode, broader multilingual performance, improved performance on complex work such as coding and tool use, or more accurate and stable handling of extended context. Nova 1.0 remains an option when standard language understanding is sufficient, lower training and inference costs are a priority, the focus is domain-specific knowledge and behavior rather than complex reasoning, or its performance has already been validated for the use case.
Model generation and adaptation method answer separate questions. Selecting Nova 2 Lite identifies the base model, while selecting CPT or SFT identifies how that model will learn from the available data. A sensible interview answer keeps those choices distinct and evaluates accuracy, speed, cost, and business requirements rather than assuming that a larger model is always preferable.
For example, choosing Nova 2 Lite because a task needs stronger reasoning does not establish whether CPT or SFT is appropriate. The data and adaptation goal still decide between raw-text domain learning and labeled behavior learning.
The mistake weak candidates make: A weak answer treats “use Nova 2” as a complete adaptation strategy. The model choice does not remove the need to determine whether the problem calls for domain knowledge from raw text, task behavior from demonstrations, or a sequence involving both.
Key takeaways
- CPT extends pre-training with raw, unlabeled domain documents so the model can acquire deeper knowledge, terminology, writing patterns, and domain fluency.
- SFT trains on labeled prompt-response pairs so the model can learn specific tasks, instructions, workflows, formats, tones, styles, or multimodal behaviors.
- CPT is particularly relevant when the available corpus contains tens of billions of domain-specific tokens, while SFT depends on high-quality examples of the desired behavior.
- CPT generally needs later instruction tuning before the adapted model can apply its new knowledge to useful tasks.
- Nova 2 Lite supports both CPT and SFT, while the documentation also lists the older Nova 1.0 Micro, Lite, and Pro models.
- The practical decision rule is to choose CPT for a domain-fluency gap, SFT for a demonstrated-behavior gap, and consider both in sequence when both gaps exist.
This distinction settles what each technique is meant to teach and what form of data it requires. It does not settle which method will meet a particular quality target, how much labeled SFT data is sufficient, or whether the cost-performance tradeoff favors adaptation over another design for a specific workload.