Chapter 5 · Domain 2 · 24% of the exam
Token-Based Pricing and Context Engineering
23 min read · Chapter 5 of 17
On this page
Certification Blueprint
| Field | Coverage |
|---|---|
| Exam | AWS Certified AI Practitioner (AIF-C01), exam guide v1.1 |
| Domain | Content Domain 2 — Fundamentals of GenAI |
| Exam weight | 24% of scored content |
| Task statement | 2.1 Explain the basic concepts of generative AI (GenAI) |
| Objectives | 2.1.4 the token-based pricing model and its effect on cost and performance for inference · 2.1.5 the role of context engineering in FM applications |
What This Chapter Covers
Chapter 04 gave you the unit. This chapter attaches the consequences.
| Chapter 04 established | Chapter 05 asks |
|---|---|
| What a token is | What a token costs, and what it costs you in time |
| That both directions are counted | Which direction moves the bill, and which moves the latency |
| That models have a bounded input | How to spend that bound deliberately |
| That prompt engineering is a term | What the larger discipline around it actually is |
Two objectives is the smallest brief in Domain 2, and it is not a light chapter. These are the two objectives most people are confident about before they start — almost everyone has used an AI assistant, and that experience feels like understanding. It is not. Experience teaches you that replies cost something and that long conversations feel slower. It does not teach you which leg is metered at which rate, or why a long conversation costs more per request rather than merely more in total.
Every wrong answer in this pair of objectives sounds economical. That is the thing to watch for.
Token-Based Pricing
Token-based pricing charges for the quantity of tokens processed, counted separately for what goes in and what comes out. Not per request. Not per user. Not per unit of time.
Two ways to hold it:
- A metered utility. You are not billed for having the tap; you are billed for the water that goes through it. A request that moves more text costs more, and an idle account costs nothing.
- Postage charged by weight, at a different rate in each direction. The envelope you send and the reply you receive are both weighed, and the return leg is charged at the higher rate.
The unit of account is not the question you asked. It is the volume of text on both legs of the exchange.
The anatomy of a billed request
What to remember from this diagram: the user's question is usually the smallest part of what you pay for. Three of the four input components — the system instruction, the conversation history and the retrieved context — were added by the application, not by the person typing.
Input and output are not priced the same
| Leg | What it contains | Typical rate |
|---|---|---|
| Input | System instruction, conversation history, retrieved context, user message | Lower per token |
| Output | Everything the model generates | Higher per token |
The reason matters more than the fact, because it explains two things at once:
- Input is read in one pass. The model takes in the whole prompt before it begins.
- Output is produced one token after another, each one depending on the ones before it.
That single mechanical difference is why the output leg is typically dearer and why output length dominates how long the user waits. Learn it once, apply it twice.
The exam never asks you for a price. Published rates change and vary by model and region. What is examinable is the shape of the model: metered, two-sided, asymmetric.
What actually drives the bill
| Driver | Why it grows | Who added it |
|---|---|---|
| A long system instruction | Resent in full on every request | The application |
| Conversation history | Prior turns resent every turn | The application |
| Retrieved context | Every injected chunk is input tokens | The application |
| Few-shot examples | Repeated on every call | The application |
| The user's message | Usually short | The user |
| The model's answer | Billed at the higher rate | The design |
Five of the six are decisions rather than traffic. Cost is a design output, not a usage fact.
The system instruction row is the one that surprises people. A long standing instruction looks like configuration and behaves like a recurring charge — it is resent, in full, on every single request.
The conversation that costs more every turn
What to remember from this diagram: the model is stateless. What looks like memory is the whole conversation being sent again. So a long chat costs more per request, not merely more in total — and that per-request growth is the version the exam asks about.
Inference Performance
The objective says "cost and performance for inference". Performance here means latency, not accuracy. A question about whether the answers are good belongs to a different objective.
| Number | What it measures | Moved mainly by |
|---|---|---|
| Time to first token | The wait before anything appears | Input length |
| Total response time | The wait until the answer is complete | Output length |
| Throughput | How much work the system sustains overall | Concurrency and request size |
A scenario that mentions users waiting is asking about this table. Which of the first two rows it means depends on one word — whether the answer is slow to start or slow to finish.
What to remember from this diagram: reading the prompt happens in one pass; writing the answer happens one token at a time. Prefill governs the time to first token and is driven by input length. Decode governs the rest and is driven by output length. That is why output length is the stronger lever on total wait.
The levers, and what each one costs you
| Lever | Moves | Trade-off |
|---|---|---|
| Shorten the system instruction | Cost | Less standing guidance |
| Cap the response length | Cost and latency | Truncated or terser answers |
| Inject fewer retrieved chunks | Cost and time to first token | Risk of missing the relevant passage |
| Summarise conversation history | Cost | Older detail is lost |
| Choose a smaller model | Cost and latency | Capability |
| Batch non-urgent work | Cost | Not available when a user is waiting |
No lever is free. Every one trades tokens for something the system used to have. Naming a lever without naming its trade-off is half an answer.
Diagnose before you pull a lever
| Symptom | What is actually happening | The reading |
|---|---|---|
| "Cost tripled but traffic was flat" | Whole conversations kept; each turn resends everything before it | Growth is in history, not usage |
| "The first word takes forever" | Many retrieved chunks injected per question | Long input delays prefill |
| "We shortened every prompt and the bill barely moved" | Long structured reports requested | The spend was on the output leg all along |
The third is the expensive one — real effort spent on the wrong side of the meter.
Context Engineering
Context engineering is the practice of deciding what information occupies the model's context window on each request, and how it is arranged, so that the model can do the job.
Two ways to hold it:
- Packing a case for a specialist you have hired for one hour. They are extremely capable and know nothing about your situation. What you put in the case decides what they can do — and the case is only so big.
- A one-page brief for a stand-in presenter. Give them the wrong page and they will present the wrong thing confidently. Give them forty pages and they will not find the point.
The model brings general capability. Context supplies the particulars — and every particular is paid for, on every request.
What occupies the context window
What to remember from this diagram: the window is refilled from empty on every request. Nothing carries over by itself — which is exactly why history has to be resent, which is exactly why a long conversation costs more each turn. The stateless fact and the growing bill are the same fact seen twice.
Note the last occupant: room has to be reserved for the answer, which shares the same window. A request that fills the window with input leaves nothing to answer into.
Context engineering is not prompt engineering
| Prompt engineering | Context engineering | |
|---|---|---|
| Governs | The wording of the instruction | The whole payload sent with it |
| Asks | How do I phrase this? | What should be in here at all? |
| Decides | Tone, structure, examples, constraints | What to retrieve, what to keep, what to drop |
| Fails as | A vague or ambiguous instruction | A window full of the wrong things |
| Examined in | Task 3.2 — Chapter 10 | This objective, 2.1.5 |
Prompt engineering writes the instruction. Context engineering decides what surrounds it. One is a component of the other, not a synonym for it — and they are examined under different objectives in different domains.
Assembling the context
What to remember from this diagram: two questions do most of the work — does this request need it, and does it change between requests. The first keeps the window clean. The second decides between a system instruction and per-request retrieval.
And note the trim order. When the window fills, something goes. If you have not decided what, the system decides for you. Compress history first because it is the most redundant; drop examples last because they shape the output format.
When more context makes the answer worse
| Failure | What it looks like | Why it happens |
|---|---|---|
| Dilution | The answer drifts to a nearby but wrong topic | The relevant passage is present but outnumbered |
| Truncation | Part of the input silently disappears | The window filled; something had to go |
| Contradiction | A confident answer built on stale material | Two injected sources disagree |
| Leakage | Sensitive data appears in an answer | It was placed in the context |
Dilution is the counter-intuitive one: the correct passage is right there, correct, and the answer is still wrong because it is outnumbered. Being in the window is not the same as being used.
"Add everything, just in case" raises cost, raises latency, and can lower accuracy at the same time. It is wrong on all three axes, which is precisely what makes it such a good distractor.
Why this is the cheapest way to specialise a model
This is the "role" the objective actually asks about. A general model does your particular job because of what you put in front of it — not because it was rebuilt for you.
| Approach | What changes | What it costs |
|---|---|---|
| Context engineering | The request | Tokens per request |
| Fine-tuning | The model's weights | A training cycle, then serving |
| Pre-training | A model from scratch | Rarely justifiable |
Context engineering is the first thing to try, not the fallback when training fails. The full comparison — in-context learning, RAG, fine-tuning, distillation, pre-training — is the customization cost ladder in Chapter 09.
Decision Rules and Exam Signals
Rule 1 — both legs are metered, and they are not priced alike. Output is typically dearer per token. Any reasoning that costs only the prompt is incomplete.
Rule 2 — cost and latency have different levers. Establish which number the question says moved before choosing anything.
Rule 3 — "slow to start" is input; "slow to finish" is output. Prefill versus decode. One word in the scenario decides it.
Rule 4 — capacity is not a discount. A larger context window lets you send more; it does not make what you send cheaper.
Rule 5 — the model is stateless. Anything that looks like memory is being resent and rebilled.
Rule 6 — more context is not more accuracy. Dilution, truncation and contradiction are all real and all get worse as you add material.
Rule 7 — context engineering decides the payload; prompt engineering decides the wording. If the question is about phrasing, it is Chapter 10's objective, not this one.
| Signal in the question | Points at |
|---|---|
| "Cost rose, traffic flat" | Conversation history or injected context growing |
| "Slow before anything appears" | Input length → prefill → time to first token |
| "Slow to finish" | Output length → decode |
| "Answers drift off-topic since we added documents" | Dilution, not model quality |
| "Needs current or private facts" | Context, not retraining |
| "Same instruction on every call" | System instruction, not per-request injection |
Distractor Patterns
| Pattern | What it looks like | How to defuse it |
|---|---|---|
| Output is free | Only the prompt is costed | Both legs are metered; output usually costs more per token |
| A bigger window is cheaper | "Move to a larger context window to cut cost" | Capacity is not a discount; you pay for what you put in it |
| Fine-tune to save tokens | Retraining offered as a cost fix | That is a training cost, and it does not help with current or private facts |
| More context is more accurate | "Inject all the documents" | Dilution and truncation both degrade the answer |
| Latency is about the prompt | A shorter prompt offered for a slow finish | Prompt length moves first-token time; output length moves total time |
| Memory is stored server-side | "The model remembers the conversation" | It is stateless; history is resent and rebilled |
| Context engineering = prompt engineering | Treated as synonyms | One writes the instruction; the other decides the payload |
The second is the highest-value one to recognise. It sounds like an upgrade, and it is the opposite of a saving.
Scenario Walkthrough
A retailer runs an assistant that answers policy questions. It injects the twenty most similar policy passages into every request, carries the full conversation, and asks for a detailed written explanation each time. Costs have risen sharply while traffic stayed flat, users complain the reply takes a long time to start, and since the policy library grew the answers sometimes cite the wrong region's rules.
| Observation | Reading | Action |
|---|---|---|
| Cost up, traffic flat | Growth is inside the request, not in usage | Summarise history; stop resending whole conversations |
| Slow to start | Large input delays prefill | Inject fewer, better-targeted chunks |
| Wrong region's rules | Dilution — twenty chunks outnumber the right one | Retrieve fewer, and filter by region |
| Detailed explanation every time | Output leg, at the higher rate | Cap the length; ask for detail only when needed |
Three different complaints and three different mechanisms. The wrong-region symptom is the one most often mislabelled as a model-quality problem — it is not. The model is doing exactly what it was given.
One change happens to help all three: injecting fewer, better chunks touches cost, prefill and dilution simultaneously. That is a property of this scenario, not a general rule.
Key Concepts
| Term | Definition |
|---|---|
| Token-based pricing | Charging by the quantity of tokens processed, counted separately for input and output, rather than per request, per user or per unit of time |
| Input tokens | Everything sent to the model — system instruction, conversation history, retrieved context and the user's message |
| Output tokens | Everything the model generates in response, metered separately and typically priced higher per token |
| System instruction | Standing guidance sent in full with every request; looks like configuration, behaves like a recurring charge |
| Context window | The bounded space a request occupies, refilled from empty each time, shared between the input and the room reserved for the answer |
| Prefill | The stage in which the model reads the whole prompt in one pass; governs time to first token |
| Decode | The stage in which output tokens are produced one after another; governs total response time |
| Time to first token | The wait before any output appears, driven mainly by input length |
| Total response time | The wait until the answer is complete, driven mainly by output length |
| Statelessness | The property that nothing carries between requests, so apparent memory is history being resent and rebilled |
| Context engineering | Deciding what information occupies the context window on each request and how it is arranged, so the model can do the job |
| Dilution | Degradation caused by relevant material being outnumbered by irrelevant material in the context |
| Truncation | Silent loss of part of the input when the context window fills |
| Trim order | The deliberate sequence for reducing a request that does not fit: compress history, then fewer chunks, then fewer examples |
Revision Flashcards
Say the answer aloud before revealing it.
1. What exactly is token-based pricing charging for? → The quantity of tokens processed, counted separately for input and output. Not per request, not per user, not per unit of time.
2. Which four things make up the input side of a billed request, and which did the user supply? → The system instruction, the conversation history, the retrieved context and the user's message. Only the last came from the user; the other three were added by the application.
3. Which leg is usually priced higher, and why? → Output. Input is read in a single pass, whereas output is produced one token after another, each depending on the ones before it.
4. Why does a long conversation cost more per request, not just more in total? → The model is stateless. Prior turns are resent with every request, so the input grows each turn even when the new message is short.
5. What is the difference between time to first token and total response time? → Time to first token is the wait before anything appears and is driven mainly by input length through prefill. Total response time is the wait until the answer is complete and is driven mainly by output length through decode.
6. A user says the reply is slow to start. Which side of the request do you look at? → The input side. Long input delays prefill. If they had said slow to finish, it would be the output side.
7. Does moving to a model with a larger context window reduce cost? → No. A larger window increases how much you may send; you still pay for every token you actually send. Capacity is not a discount.
8. Which single lever moves both cost and latency? → Capping the response length. It reduces output tokens, which are both the dearer leg and the driver of total response time. The trade-off is terser or truncated answers.
9. Define context engineering in one sentence. → Deciding what information occupies the model's context window on each request, and how it is arranged, so the model can do the job.
10. How is context engineering different from prompt engineering? → Prompt engineering governs the wording of the instruction. Context engineering governs the whole payload around it — what to retrieve, what history to keep, what to drop. Prompt engineering is a component of context engineering.
11. Name three ways that adding more context makes the answer worse. → Dilution, where the relevant passage is outnumbered; truncation, where the window fills and something is silently dropped; and contradiction, where two injected sources disagree and the model answers confidently from the wrong one.
12. What is the trim order when a request does not fit the window? → Compress the conversation history first because it is the most redundant, then inject fewer retrieved chunks, then reduce examples last because they shape the output format.
The Four-Beat Answer
The core question this chapter prepares you for: "How does token-based pricing work, and what do you do about it?"
Four beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.
- The unit. Charging is metered by token quantity, counted on both legs, with output typically dearer per token because it is generated sequentially rather than read in one pass.
- What is actually in the request. The system instruction, the conversation history, the retrieved context and the user's message — and most of that is the application's own decisions rather than user traffic. This is the beat that makes an answer sound like it came from someone who has built something.
- Which number moves. Input length drives prefill and therefore time to first token. Output length drives decode and therefore total response time. Both drive cost. Name the lever and what it trades away.
- The discipline. Context engineering: treat the window as a finite budget, allocate it on purpose, and know the trim order for when it does not fit. It is the cheapest way to make a general model do a particular job.
A weak answer describes billing. A strong answer names the levers and what each one costs you.
Why This Helps You
On the job: almost every foundation-model system that becomes unexpectedly expensive got there the same way — a growing system instruction, unmanaged conversation history, and retrieval tuned for recall rather than precision. All three are invisible until the bill arrives, and all three are design decisions rather than traffic.
In interviews: "how would you reduce the cost of this feature?" is a standard system-design follow-up, and "use a cheaper model" is the answer everyone gives. Naming which leg you would cut, which number it moves, and what capability it costs is the answer that sounds like experience.
On the exam: Domain 2 is 24% of scored content, and these two objectives are unusually distractor-rich because every wrong option sounds thrifty. The habit of asking which number does this actually move is worth more here than anywhere else on the paper.
Chapter Checklist
- I can describe token-based pricing as metered, two-sided and asymmetric
- I can name the four components of an assembled prompt and say which one the user supplied
- I can explain why a multi-turn conversation costs more per request over time
- I can separate prefill from decode, and time to first token from total response time
- I can say which lever moves cost, which moves latency, and which moves both
- I can name the trade-off that comes with every lever I propose
- I can define context engineering without describing prompt engineering
- I can treat the context window as a budget and state a trim order
- I can explain three ways in which more context makes an answer worse
- I can say why context is the cheapest way to specialise a general model
- I can reject "a bigger context window will reduce our costs" and say why
After the Chapter
- Complete
student/project.md— parts 19-22 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one. - Take
student/quiz.mdclosed-book, then review the reasoning for every question you guessed, including the ones you got right. In this chapter especially, a right guess is not a right understanding — the distractors are designed to sound thrifty. - Open the official v1.1 exam guide's Domain 2 page and confirm you can attach a concept from this chapter to the fourth and fifth bullets under Task Statement 2.1.
- Next chapter: Chapter 06 — Agentic AI: MCP, Multi-Agent Patterns, Memory, Tools, and Orchestration (Domain 2, Objective 2.1.6). It is the last objective in Task 2.1 and the newest material on the exam. ⚠️ Note the domain carefully: version 1.1 moved the Model Context Protocol out of Domain 3 and into Domain 2, so anyone revising from older material has it filed in the wrong place. Memory management appears on that objective's list — and memory is this chapter's context window put under deliberate control.
Chapter 5 quiz
13 questions on this chapter, marked instantly, with an explanation for every answer.