Chapter 10 · Domain 3 · 28% of the exam
Prompt Engineering: Techniques, Practice, and Risk
28 min read · Chapter 10 of 17
On this page
- Certification Blueprint
- What This Chapter Covers
- The Five Constructs
- Techniques: How Much Demonstration to Supply
- Benefits and Best Practices
- Risks and Limitations
- Prompt Versioning and Management
- Decision Rules and Exam Signals
- Distractor Patterns
- Scenario Walkthrough
- Key Concepts
- Revision Flashcards
- The Five-Beat Answer
- Why This Helps You
- Chapter Checklist
- After the Chapter
Certification Blueprint
| Field | Coverage |
|---|---|
| Exam | AWS Certified AI Practitioner (AIF-C01), exam guide v1.1 |
| Domain | Content Domain 3 — Applications of Foundation Models |
| Exam weight | 28% of scored content — the largest domain on the exam |
| Task statement | 3.2 Choose effective prompt engineering techniques |
| Objectives | 3.2.1 concepts and constructs · 3.2.2 techniques · 3.2.3 benefits and best practices · 3.2.4 risks and limitations · 3.2.5 prompt versioning and management with Amazon Bedrock Prompt Management |
| Service families | Amazon Bedrock Prompt Management |
What This Chapter Covers
Chapter 09 ended by placing prompting at the bottom of the customization cost ladder — the rung you try first, and the one that resolves most problems. It did not say how. This chapter is that rung.
Everything in it follows from a single structural fact, and it is worth stating before any terminology:
The model receives one string. Whatever you wrote as instructions, whatever was retrieved from your documents, and whatever the user typed all arrive as a single undifferentiated block of text. There is no channel, no privilege level, no separator the model is obliged to respect.
That one fact explains both halves of this chapter. It is why prompting works at all — you can change what a model does without touching its weights, purely by changing text. And it is why all four of the named risks exist, because anything that reaches the string is read as though the author had written it.
Hold that sentence. The rest of the chapter is its detail.
What to remember from this diagram: the five constructs are things you assemble, but the arrow into the model carries only one payload. The boundaries you can see in the diagram do not exist by the time the model reads it. That is the whole chapter in one picture.
The Five Constructs
Objective 3.2.1 names context, instruction and negative prompts as examples. A complete prompt is usually built from five parts, and being able to name the missing one is more useful on the exam than being able to say a prompt is "too vague".
| Construct | What it contributes | What its absence looks like |
|---|---|---|
| Context | Background the model would not otherwise have — the role it should adopt, the audience, the situation | Answers that are correct in general but wrong for this setting |
| Instruction | The task itself, stated as a directive. The one construct a prompt cannot omit | The model summarises when you wanted a comparison |
| Negative prompt | What to exclude, avoid, or never do | Disclaimers, apologies, or forbidden content that keeps reappearing |
| Input data | The specific item to act on this time — the ticket, the paragraph, the record | A generic answer that ignores the particular case |
| Output indicator | The required shape of the response: format, length, schema, structure | Prose where the system needed JSON; a parser that breaks intermittently |
The exam signal is diagnostic. A scenario that describes output which is accurate but arrives in the wrong shape is naming a missing output indicator, not a model that needs replacing. A scenario where the model keeps adding a safety disclaimer nobody asked for is naming a missing negative prompt.
Negative prompts deserve their own paragraph
A negative prompt states what must not appear. It is the construct most often left out, and the one most often described incorrectly in distractors.
Two things to hold:
- It is an instruction, not a filter. The model is asked to avoid something; nothing intercepts the output afterwards to check. A distractor that describes a negative prompt as removing content from a response after generation is describing output filtering, which is a Domain 5 control and belongs to Chapter 16.
- It is not a guarantee. Instructing a model to never do something reduces the frequency; it does not make it impossible. This is a limitation of prompt engineering in the sense Objective 3.2.4 asks about, and it is why the risks section later in this chapter matters.
Techniques: How Much Demonstration to Supply
Objective 3.2.2 names five techniques. They are not five alternatives on one list — they sit on two different axes, and the most common exam error is treating them as one.
Axis one — how many examples you supply.
| Technique | What you supply | When it is the right choice |
|---|---|---|
| Zero-shot | The instruction alone, no examples | The model already does this task reliably. Start here; it is the cheapest per request |
| Single-shot | One worked example | The task is understood but the shape of a good answer needs pinning down |
| Few-shot | Several examples, consistently formatted | A pattern needs demonstrating — a classification scheme, a house format, an edge case convention |
Axis two — whether you ask for reasoning.
Chain-of-thought asks the model to work through intermediate steps rather than leap to an answer. It is not "more examples"; it is a different request entirely, and it can be combined with any of the three above.
What to remember from this diagram: the branch point is what kind of gap the output has, not how hard the task feels. Every path ends at a template, because a prompt that works is an asset worth reusing rather than retyping.
Why the shot ladder cannot fix a reasoning failure
This is the single highest-value idea in Objective 3.2.2.
Examples demonstrate what a good answer looks like. They do not demonstrate how to get there. If a task fails because it requires several linked deductions — arithmetic, ordering, comparing against multiple constraints — then showing the model more finished answers gives it more targets to imitate and no additional capacity to reason.
The result is a prompt that is longer, more expensive on every single request, and just as wrong.
What to remember from this diagram: the right-hand path produces something the left-hand path never does — a trail. That is a second, quieter benefit of chain-of-thought: when the answer is wrong, you can see which step went wrong. A direct prompt fails silently.
Prompt templates
A template is a prompt with the varying parts parameterised, so the wording that was tested is the wording that ships every time.
| A template gives you | Why it matters |
|---|---|
| One place to change the wording | Otherwise the prompt is copy-pasted across services and drifts |
| Consistent structure across requests | Output stays parseable because the output indicator is always present |
| A unit you can version | Which is exactly what Objective 3.2.5 is about |
Templates are the bridge to the last objective. Once a prompt is a named, reusable artefact, the question "which version of it is in production?" becomes answerable — and Amazon Bedrock Prompt Management is the service that answers it.
Benefits and Best Practices
Objective 3.2.3 names six things: response quality improvement, experimentation, guardrails, discovery, specificity and concision, and using multiple comments. All six are examinable, and two of them are routinely skipped in study material.
| Named item | What it actually means | The practice |
|---|---|---|
| Response quality improvement | The headline benefit — better output with no change to the model, no training cost, and effect on the very next request | Treat the prompt as the first lever, not the last |
| Experimentation | Prompting is cheap enough to try variants and compare them, which is not true of any other rung on the ladder | Change one thing at a time and keep what you tried |
| Guardrails | In this objective's narrow sense: using the prompt to constrain what the model will do — scope, tone, refusal behaviour | State boundaries in the prompt rather than assuming defaults |
| Discovery | Using prompting to find out what the model can already do before deciding you need something more expensive | Probe capability first; most "we need fine-tuning" conclusions are untested |
| Specificity and concision | Say exactly what is wanted, and no more. These pull the same way, not opposite ways | Remove words that do not change the output; add words that do |
| Using multiple comments | The exam guide's own phrasing — breaking a prompt into clearly delineated commented sections rather than one undifferentiated paragraph | Label the parts, so context, instruction and data are visually separable |
⚠️ "Using multiple comments" is the exam guide's literal wording, verified against the live Domain 3 page. It is easy to misread as "multiple components" and it means what it says: annotate and delimit the sections of a prompt. The practical effect is the same thing the constructs table teaches — make the parts distinguishable to a human maintainer, since they are not distinguishable to the model.
Specificity and concision are not in tension, and the exam may test that. A vague prompt and a padded prompt fail for the same reason: the words that would have determined the output are not there, or are buried among words that do nothing. Adding length is not adding specificity.
Risks and Limitations
Objective 3.2.4 names four: exposure, poisoning, hijacking, jailbreaking. Learn them by where the attacker's text enters, because that is the only thing that reliably separates them.
What to remember from this diagram: three arrows go in and one comes out. Three of the risks are about text getting into the prompt; one is about text getting out of it.
| Risk | Where it enters | What the attacker achieves |
|---|---|---|
| Prompt injection / hijacking | The user input | The model is redirected to the attacker's task instead of yours — "ignore your instructions and do this instead" |
| Jailbreaking | The user input | The model's safety constraints are escaped, producing content it was built to refuse |
| Poisoning | The source material the prompt will later retrieve or the examples it is given | The attack is planted in advance and fires when a legitimate request pulls it in |
| Exposure | Nothing enters — it leaves | The system prompt, private context, or another user's data appears in the response |
Telling the confusable pairs apart
Hijacking versus jailbreaking. Both arrive through user input, which is why they are the pair most often confused.
- Hijacking changes the task. The model does something for the attacker rather than for you — it is still behaving, just working for someone else.
- Jailbreaking changes the limits. The model does the thing it was constrained not to do.
A support assistant tricked into writing marketing copy has been hijacked. A support assistant tricked into producing content the operator forbade has been jailbroken.
Poisoning versus hijacking. Both plant instructions, and the difference is timing and route. Hijacking is delivered by the attacker at request time through the input. Poisoning is placed in a document, a knowledge base or an example set beforehand, and is triggered later by an innocent user's ordinary question. That delay is the signal: a scenario where the attacker is not present when the attack fires is describing poisoning.
Exposure is the odd one out and it is the one candidates most often mislabel. Nothing hostile has to enter for exposure to happen — a prompt that includes a system instruction, or context assembled from another customer's record, can simply be repeated back. A scenario in which no attacker appears at all, and confidential text appears in an answer, is exposure.
The limitation, stated plainly
All four exist because of the fact this chapter opened with. The model cannot distinguish your instructions from anything else in the string, so instruction-based defences are advisory. Telling a model to ignore attempts to override it is itself just more text in the same undifferentiated block.
This is why Objective 3.2.4 is a risks and limitations objective rather than a controls objective. The controls — input validation, output filtering, guardrails as a service, least-privilege on what the application can do with the answer — are Domain 5 and are examined in Chapter 16.
Prompt Versioning and Management
Objective 3.2.5 names Amazon Bedrock Prompt Management specifically. The idea underneath it: a prompt is code. It determines system behaviour, it changes, and a change can break production. Everything that follows is what you already do for code.
What to remember from this diagram: the application points at a version, never at the draft. That single arrow is what makes every other property possible.
| Capability | What it gives you |
|---|---|
| Store prompts as named artefacts | The prompt has one home instead of being duplicated in application code |
| Draft versus version | A draft is editable; a version is an immutable snapshot that cannot change under a running application |
| Test against saved inputs | A change can be compared before it reaches production |
| Deploy by version reference | Production behaviour is pinned and reproducible |
| Roll back | A regression is undone by pointing at the previous version, not by remembering what the wording used to be |
The exam signal is "which version was live?" A scenario describing an application whose answers changed and a team that cannot say what the prompt used to say is describing the absence of prompt management. The fix named in this objective is versioning, not more testing and not a better model.
Why deploying a draft is the wrong answer even when it works. A draft is mutable. An application referencing one has behaviour that can change without any deployment, any review, or any record — someone edits the draft and production shifts. The immutability of a version is not a convenience feature; it is the entire point.
Decision Rules and Exam Signals
Rule 1 — the model reads one string. Instructions, retrieved text and user input are indistinguishable by the time they arrive. Every risk in this task statement follows from this.
Rule 2 — name the missing construct. "The prompt is bad" is not a diagnosis. Wrong shape means a missing output indicator; unwanted content means a missing negative prompt; wrong setting means missing context.
Rule 3 — examples fix shape, not reasoning. If the failure is multi-step, few-shot makes the prompt longer and no better. Chain-of-thought is the technique that addresses reasoning.
Rule 4 — start at zero-shot and climb only on evidence. Same discipline as Chapter 09's cost ladder, one level down: every example is paid for on every request.
Rule 5 — chain-of-thought is a different axis from the shot count. It combines with zero-, single- or few-shot; it is not a fourth position on the same ladder.
Rule 6 — specificity and concision agree. Padding is not precision. Remove words that do not change the output.
Rule 7 — classify a risk by where the text entered. User input at request time is hijacking or jailbreaking; planted in advance in a source is poisoning; leaving in the response is exposure.
Rule 8 — task changed is hijacking, limits escaped is jailbreaking.
Rule 9 — prompt defences are advisory, not enforcing. An instruction telling the model to resist override attempts is more text in the same string. Real controls are Domain 5.
Rule 10 — applications reference versions, never drafts. Immutability is what makes rollback and reproducibility possible.
Distractor Patterns
| Pattern | What it looks like | How to defuse it |
|---|---|---|
| More examples for a reasoning failure | Few-shot offered for a task failing on arithmetic or multi-step logic | Examples show finished answers; they add no reasoning capacity. Chain-of-thought is the technique |
| Fine-tune it | Training offered before prompting has been genuinely tried | Discovery and experimentation come first; this is Chapter 09's cost ladder restated |
| Temperature as a prompt technique | Changing an inference parameter offered as prompt engineering | Temperature is Objective 3.1.2 and changes no words in the prompt |
| Negative prompt as a filter | Described as removing content after generation | It is an instruction before generation. Post-generation removal is output filtering, Domain 5 |
| Hijacking labelled jailbreaking | Any injected-text scenario called jailbreaking | Ask what changed: the task (hijacking) or the limits (jailbreaking) |
| Poisoning labelled injection | A planted document called prompt injection | Poisoning is placed in advance and fires on an innocent request |
| Exposure needs an attacker | Assuming a leak implies someone attacked | A prompt can repeat its own system instructions with nobody hostile involved |
| "Instruct the model to refuse overrides" | Offered as the fix for injection | Advisory only — it is more text in the same undifferentiated string |
| Longer prompt = more specific | Padding presented as precision | Specificity and concision pull the same way |
| Deploy the draft | Application pointed at an editable draft | A version is immutable; that is what makes rollback and reproducibility work |
| Better model instead of versioning | Model swap offered for "answers changed and we don't know why" | That is a change-management gap, and versioning is the named fix |
The fifth and sixth rows are the two most reliable ways to lose marks in Objective 3.2.4, because all three risks involve text that was not written by the prompt's author.
Scenario Walkthrough
An insurer runs an assistant that answers policy questions. It retrieves clauses from an internal document library that brokers can upload to. Answers are factually right but arrive as prose, while the downstream system needs structured fields. A separate complaint says the assistant sometimes repeats an internal instruction beginning "You are a policy assistant; never discuss pricing." The team also reports that after someone edited the wording last week, answer quality dropped, and nobody can say what the prompt said before.
| Requirement | Reading | Decision |
|---|---|---|
| Right facts, wrong shape | The content is fine; the form is unspecified | Missing output indicator — add the required schema |
| Brokers can upload to the retrieved library | Untrusted text enters the prompt in advance | Poisoning exposure, not hijacking — the uploader need not be present when it fires |
| Internal instruction appears in answers | The system prompt is leaving in the response | Exposure — and no attacker is implied |
| Wording edited, quality dropped, no history | Change management, not model quality | Amazon Bedrock Prompt Management — version it, deploy the version, roll back |
Four requirements, and the second one is the test. A candidate who has decided "this is a prompt injection question" will label the broker upload as hijacking. It is not: nobody is injecting anything at request time. The attack is stored, and it fires when an ordinary user asks an ordinary question that happens to retrieve it. Route of entry, not intent, is what names these.
Key Concepts
| Term | Definition |
|---|---|
| Prompt engineering | Shaping a model's output by changing the text it is given, without modifying its weights |
| Context | The construct supplying background the model would not otherwise have — role, audience, situation |
| Instruction | The construct stating the task as a directive; the one part a prompt cannot omit |
| Negative prompt | The construct stating what must not appear; an instruction before generation, not a filter after it |
| Input data | The construct carrying the specific item to act on for this request |
| Output indicator | The construct specifying the required shape of the response — format, length, schema |
| Zero-shot | Prompting with an instruction and no examples |
| Single-shot | Prompting with exactly one worked example, typically to pin down the shape of a good answer |
| Few-shot | Prompting with several consistently formatted examples to demonstrate a pattern |
| Chain-of-thought | Asking the model to produce intermediate reasoning steps rather than leap to an answer; addresses reasoning failures, and leaves a trail that can be checked |
| Prompt template | A prompt with its varying parts parameterised so tested wording is reused rather than retyped |
| Exposure | A risk in which the system prompt or private context appears in the response; requires no attacker |
| Poisoning | A risk in which malicious content is planted in advance in a source or example set, and fires when a legitimate request retrieves it |
| Prompt injection / hijacking | A risk in which text supplied at request time redirects the model to the attacker's task |
| Jailbreaking | A risk in which crafted input escapes the model's safety constraints, producing content it was built to refuse |
| Amazon Bedrock Prompt Management | The AWS service for storing, versioning, testing and deploying prompts as managed artefacts |
| Prompt version | An immutable snapshot of a prompt that an application can reference, enabling reproducible behaviour and rollback |
Revision Flashcards
Say the answer aloud before revealing it.
1. State the one structural fact that this whole chapter follows from. → The model receives a single undifferentiated string. Your instructions, any retrieved documents and the user's input are all the same text by the time it reads them — there is no channel, privilege level, or separator it is obliged to respect. That is why prompting can change behaviour without touching weights, and why every risk in this task statement exists.
2. Name the five constructs of a prompt and the one that cannot be omitted. → Context, instruction, negative prompt, input data, output indicator. The instruction is the one a prompt cannot do without — it states the task. The others sharpen, constrain, supply or shape, but without an instruction there is no task.
3. Output is accurate but arrives as prose when the system needed JSON. What is missing? → The output indicator — the construct that specifies the required shape. This is a prompt-construction gap, not a model-quality problem, and the distractor to avoid is the one offering a different or larger model.
4. What is a negative prompt, and what is it not? → It is an instruction stating what must not appear, applied before generation. It is not a filter: nothing inspects the output afterwards to enforce it. An option describing content being removed after generation is describing output filtering, which is a Domain 5 control.
5. Distinguish zero-shot, single-shot and few-shot. → Zero-shot supplies the instruction alone; single-shot adds exactly one worked example, usually to pin down the shape of a good answer; few-shot supplies several consistently formatted examples to demonstrate a pattern. They are one axis — how much demonstration — and every example is paid for on every request.
6. Why can few-shot prompting not fix a multi-step reasoning failure? → Because examples demonstrate what a good answer looks like, not how to reach one. Showing more finished answers gives the model more to imitate and no additional capacity to reason. You get a longer prompt, a bigger per-request bill, and the same wrong answer. Chain-of-thought is the technique that addresses reasoning.
7. What does chain-of-thought produce besides a better answer? → A visible trail. Because intermediate steps appear in the response, a reviewer can see which step went wrong when the answer is wrong. A direct prompt fails silently, which is why chain-of-thought is valuable even when the final answer would have been right.
8. Name the six items in Objective 3.2.3, including the two usually skipped. → Response quality improvement, experimentation, guardrails, discovery, specificity and concision, and using multiple comments. The two usually skipped are discovery — probing what the model can already do before paying for something more expensive — and experimentation, which is only affordable at this rung.
9. Are specificity and concision in tension? → No, they pull the same way. A vague prompt and a padded prompt fail for the same reason: the words that would determine the output are absent or buried. Adding length is not adding specificity. Remove words that do not change the output; add words that do.
10. Distinguish hijacking from jailbreaking. → Both arrive through user input. Hijacking changes the task — the model works for the attacker instead of you. Jailbreaking changes the limits — the model produces something it was constrained to refuse. A support assistant tricked into writing marketing copy was hijacked; one tricked into forbidden content was jailbroken.
11. A broker uploads a document containing hidden instructions; weeks later an ordinary customer question retrieves it and the assistant misbehaves. Name the risk. → Poisoning. The attacker was not present when the attack fired — the malicious text was planted in advance in a source the prompt would later retrieve, and an innocent request triggered it. The delay between planting and firing is the signal that separates it from hijacking.
12. Why must an application reference a prompt version rather than a draft? → Because a version is immutable and a draft is not. An application pointed at a draft has behaviour that can change with no deployment, no review and no record — someone edits the draft and production shifts underneath it. Immutability is what makes reproducible behaviour and rollback possible, and it is the whole point of versioning.
The Five-Beat Answer
The core question this chapter prepares you for: "How would you get better results out of a foundation model without changing the model — and what could go wrong?"
Five beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.
- Construct — say what a prompt is made of before saying how to improve one: context, instruction, negative prompt, input data, output indicator. Diagnosing a weak prompt means naming the missing part, not calling it vague.
- Technique — pick from the kind of gap. Zero-shot first; examples for shape and pattern; chain-of-thought for reasoning. Say explicitly that examples cannot repair reasoning, because that is the distinction most candidates miss.
- Practise — name what good prompting work looks like: experiment cheaply and one change at a time, use discovery to test capability before paying for more, be specific and concise, and delimit the sections of the prompt.
- Risk — give the four named risks classified by where the text entered: hijacking and jailbreaking through user input, poisoning planted in advance, exposure on the way out. Then state the limitation plainly — the model cannot tell your instructions from anything else, so prompt-level defences are advisory.
- Manage — finish on versioning. A prompt is code: store it, version it, test it, deploy a version rather than a draft, and roll back when a change regresses. Name Amazon Bedrock Prompt Management.
A strong answer diagnoses the gap before choosing a technique, and names a risk that prompting alone cannot close. A weak answer proposes few-shot for everything.
Why This Helps You
On the job: the most common waste in this area is escalating past prompting without ever testing it properly — commissioning fine-tuning for a problem an output indicator would have solved. The second most common is treating a prompt as a string literal in application code, which is how a team ends up unable to say what changed when quality drops. Both are avoided by the same habit: treat the prompt as an engineered, versioned artefact.
In interviews: "how would you stop prompt injection?" is a standard screening question and most candidates answer it with a better instruction. The strong answer names the structural reason that cannot work — the model reads one undifferentiated string — and moves to controls outside the prompt. Being able to separate hijacking, jailbreaking, poisoning and exposure by route of entry marks out someone who has actually operated one of these systems.
On the exam: Domain 3 is 28% of scored content and this chapter owns an entire task statement within it. The highest-value habits are diagnosing the gap before choosing a technique, and classifying a risk by where the text entered rather than by how the scenario feels.
Chapter Checklist
- I can state why the model receiving one undifferentiated string explains both prompting's power and its risks
- I can name the five constructs and say which one a described failure is missing
- I can explain what a negative prompt is, and why it is not a filter
- I can distinguish zero-shot, single-shot and few-shot, and say what each costs per request
- I can explain why examples cannot fix a multi-step reasoning failure
- I can say what chain-of-thought provides beyond a better answer
- I can explain what a prompt template gives me and how it leads to versioning
- I can name all six items in the benefits and best practices objective
- I can explain why specificity and concision are not in tension
- I can classify exposure, poisoning, hijacking and jailbreaking by where the text entered
- I can tell hijacking from jailbreaking by whether the task or the limits changed
- I can explain why prompt-level defences against injection are advisory
- I can say why an application must reference a version rather than a draft
After the Chapter
- Complete
student/project.md— parts 37-39 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one. - Take
student/quiz.mdclosed-book, then review the reasoning for every question you guessed, including the ones you got right. Pay particular attention to questions 8 and 10 — they are deliberate mirror images, and missing both means the route-of-entry classification has not landed. - Open the official v1.1 exam guide's Domain 3 page and confirm you can attach a concept from this chapter to each of the five bullets under Task Statement 3.2. Note the exact wording of the third bullet — it says "using multiple comments", and reading it as "components" changes what it asks for.
- Next: Chapter 11 — Training and Fine-Tuning Foundation Models (Domain 3, Task 3.3). This chapter's boundary is the point where prompting has genuinely failed. Chapter 11 starts there: what pre-training, fine-tuning, continuous pre-training and distillation actually involve, and the data preparation each one demands.
Chapter 10 quiz
13 questions on this chapter, marked instantly, with an explanation for every answer.