Free live cohort on Google Meet — register your interest →

Chapter 5 · Domain 2 · 24% of the exam

Token-Based Pricing and Context Engineering

23 min read · Chapter 5 of 17

On this page
  1. Certification Blueprint
  2. What This Chapter Covers
  3. Token-Based Pricing
  4. Inference Performance
  5. Context Engineering
  6. Decision Rules and Exam Signals
  7. Distractor Patterns
  8. Scenario Walkthrough
  9. Key Concepts
  10. Revision Flashcards
  11. The Four-Beat Answer
  12. Why This Helps You
  13. Chapter Checklist
  14. After the Chapter

Certification Blueprint

Field Coverage
Exam AWS Certified AI Practitioner (AIF-C01), exam guide v1.1
Domain Content Domain 2 — Fundamentals of GenAI
Exam weight 24% of scored content
Task statement 2.1 Explain the basic concepts of generative AI (GenAI)
Objectives 2.1.4 the token-based pricing model and its effect on cost and performance for inference · 2.1.5 the role of context engineering in FM applications

What This Chapter Covers

Chapter 04 gave you the unit. This chapter attaches the consequences.

Chapter 04 established Chapter 05 asks
What a token is What a token costs, and what it costs you in time
That both directions are counted Which direction moves the bill, and which moves the latency
That models have a bounded input How to spend that bound deliberately
That prompt engineering is a term What the larger discipline around it actually is

Two objectives is the smallest brief in Domain 2, and it is not a light chapter. These are the two objectives most people are confident about before they start — almost everyone has used an AI assistant, and that experience feels like understanding. It is not. Experience teaches you that replies cost something and that long conversations feel slower. It does not teach you which leg is metered at which rate, or why a long conversation costs more per request rather than merely more in total.

Every wrong answer in this pair of objectives sounds economical. That is the thing to watch for.

Token-Based Pricing

Token-based pricing charges for the quantity of tokens processed, counted separately for what goes in and what comes out. Not per request. Not per user. Not per unit of time.

Two ways to hold it:

  • A metered utility. You are not billed for having the tap; you are billed for the water that goes through it. A request that moves more text costs more, and an idle account costs nothing.
  • Postage charged by weight, at a different rate in each direction. The envelope you send and the reply you receive are both weighed, and the return leg is charged at the higher rate.

The unit of account is not the question you asked. It is the volume of text on both legs of the exchange.

The anatomy of a billed request

Four components — the system instruction, the conversation history, retrieved context chunks and the user message — combining into one assembled prompt that is metered as input tokens, passing through the foundation model, which emits a response metered separately as output tokens, with the billed request being the input count times the input rate plus the output count times the output rate

What to remember from this diagram: the user's question is usually the smallest part of what you pay for. Three of the four input components — the system instruction, the conversation history and the retrieved context — were added by the application, not by the person typing.

Input and output are not priced the same

Leg What it contains Typical rate
Input System instruction, conversation history, retrieved context, user message Lower per token
Output Everything the model generates Higher per token

The reason matters more than the fact, because it explains two things at once:

  • Input is read in one pass. The model takes in the whole prompt before it begins.
  • Output is produced one token after another, each one depending on the ones before it.

That single mechanical difference is why the output leg is typically dearer and why output length dominates how long the user waits. Learn it once, apply it twice.

The exam never asks you for a price. Published rates change and vary by model and region. What is examinable is the shape of the model: metered, two-sided, asymmetric.

What actually drives the bill

Driver Why it grows Who added it
A long system instruction Resent in full on every request The application
Conversation history Prior turns resent every turn The application
Retrieved context Every injected chunk is input tokens The application
Few-shot examples Repeated on every call The application
The user's message Usually short The user
The model's answer Billed at the higher rate The design

Five of the six are decisions rather than traffic. Cost is a design output, not a usage fact.

The system instruction row is the one that surprises people. A long standing instruction looks like configuration and behaves like a recurring charge — it is resent, in full, on every single request.

The conversation that costs more every turn

Turn one sending the system instruction plus the first message, turn two resending the system instruction plus turn one plus the second message, turn three resending everything above plus a third message, showing input tokens per request rising every turn even when each new message is short, and three ways to stop the growth — summarising older turns, keeping a sliding window of recent turns, and starting a fresh session when the task changes

What to remember from this diagram: the model is stateless. What looks like memory is the whole conversation being sent again. So a long chat costs more per request, not merely more in total — and that per-request growth is the version the exam asks about.

Inference Performance

The objective says "cost and performance for inference". Performance here means latency, not accuracy. A question about whether the answers are good belongs to a different objective.

Number What it measures Moved mainly by
Time to first token The wait before anything appears Input length
Total response time The wait until the answer is complete Output length
Throughput How much work the system sustains overall Concurrency and request size

A scenario that mentions users waiting is asking about this table. Which of the first two rows it means depends on one word — whether the answer is slow to start or slow to finish.

Input tokens determining the prefill stage in which the model reads the whole prompt, producing the time to first token; output tokens determining the decode stage in which tokens are produced one after another in sequence; both feeding total response time, with the conclusion that a shorter prompt improves time to first token while a shorter answer improves total time and output length is the stronger lever

What to remember from this diagram: reading the prompt happens in one pass; writing the answer happens one token at a time. Prefill governs the time to first token and is driven by input length. Decode governs the rest and is driven by output length. That is why output length is the stronger lever on total wait.

The levers, and what each one costs you

Lever Moves Trade-off
Shorten the system instruction Cost Less standing guidance
Cap the response length Cost and latency Truncated or terser answers
Inject fewer retrieved chunks Cost and time to first token Risk of missing the relevant passage
Summarise conversation history Cost Older detail is lost
Choose a smaller model Cost and latency Capability
Batch non-urgent work Cost Not available when a user is waiting

No lever is free. Every one trades tokens for something the system used to have. Naming a lever without naming its trade-off is half an answer.

Diagnose before you pull a lever

Symptom What is actually happening The reading
"Cost tripled but traffic was flat" Whole conversations kept; each turn resends everything before it Growth is in history, not usage
"The first word takes forever" Many retrieved chunks injected per question Long input delays prefill
"We shortened every prompt and the bill barely moved" Long structured reports requested The spend was on the output leg all along

The third is the expensive one — real effort spent on the wrong side of the meter.

Context Engineering

Context engineering is the practice of deciding what information occupies the model's context window on each request, and how it is arranged, so that the model can do the job.

Two ways to hold it:

  • Packing a case for a specialist you have hired for one hour. They are extremely capable and know nothing about your situation. What you put in the case decides what they can do — and the case is only so big.
  • A one-page brief for a stand-in presenter. Give them the wrong page and they will present the wrong thing confidently. Give them forty pages and they will not find the point.

The model brings general capability. Context supplies the particulars — and every particular is paid for, on every request.

What occupies the context window

The context window shown as a fixed budget refilled from empty on every request, divided among the system instruction that is resent every time, few-shot examples repeated every time, conversation history that grows unless managed, retrieved chunks in whatever quantity is injected, the user message, and room reserved for the answer, with a check on whether the total fits — proceeding if it does, and dropping something if it does not

What to remember from this diagram: the window is refilled from empty on every request. Nothing carries over by itself — which is exactly why history has to be resent, which is exactly why a long conversation costs more each turn. The stateless fact and the growing bill are the same fact seen twice.

Note the last occupant: room has to be reserved for the answer, which shares the same window. A request that fills the window with input leaves nothing to answer into.

Context engineering is not prompt engineering

Prompt engineering Context engineering
Governs The wording of the instruction The whole payload sent with it
Asks How do I phrase this? What should be in here at all?
Decides Tone, structure, examples, constraints What to retrieve, what to keep, what to drop
Fails as A vague or ambiguous instruction A window full of the wrong things
Examined in Task 3.2 — Chapter 10 This objective, 2.1.5

Prompt engineering writes the instruction. Context engineering decides what surrounds it. One is a component of the other, not a synonym for it — and they are examined under different objectives in different domains.

Assembling the context

A decision flow that asks whether a candidate piece of information is needed for this request and not already in the model's general knowledge, leaving it out if not, then asking whether it is identical on every request — placing it in the system instruction if so, or retrieving only the relevant part per request if not — then checking whether the total is within budget, sending if it is, and otherwise trimming in a fixed order of compressing history, then fewer chunks, then fewer examples

What to remember from this diagram: two questions do most of the work — does this request need it, and does it change between requests. The first keeps the window clean. The second decides between a system instruction and per-request retrieval.

And note the trim order. When the window fills, something goes. If you have not decided what, the system decides for you. Compress history first because it is the most redundant; drop examples last because they shape the output format.

When more context makes the answer worse

Failure What it looks like Why it happens
Dilution The answer drifts to a nearby but wrong topic The relevant passage is present but outnumbered
Truncation Part of the input silently disappears The window filled; something had to go
Contradiction A confident answer built on stale material Two injected sources disagree
Leakage Sensitive data appears in an answer It was placed in the context

Dilution is the counter-intuitive one: the correct passage is right there, correct, and the answer is still wrong because it is outnumbered. Being in the window is not the same as being used.

"Add everything, just in case" raises cost, raises latency, and can lower accuracy at the same time. It is wrong on all three axes, which is precisely what makes it such a good distractor.

Why this is the cheapest way to specialise a model

This is the "role" the objective actually asks about. A general model does your particular job because of what you put in front of it — not because it was rebuilt for you.

Approach What changes What it costs
Context engineering The request Tokens per request
Fine-tuning The model's weights A training cycle, then serving
Pre-training A model from scratch Rarely justifiable

Context engineering is the first thing to try, not the fallback when training fails. The full comparison — in-context learning, RAG, fine-tuning, distillation, pre-training — is the customization cost ladder in Chapter 09.

Decision Rules and Exam Signals

Rule 1 — both legs are metered, and they are not priced alike. Output is typically dearer per token. Any reasoning that costs only the prompt is incomplete.

Rule 2 — cost and latency have different levers. Establish which number the question says moved before choosing anything.

Rule 3 — "slow to start" is input; "slow to finish" is output. Prefill versus decode. One word in the scenario decides it.

Rule 4 — capacity is not a discount. A larger context window lets you send more; it does not make what you send cheaper.

Rule 5 — the model is stateless. Anything that looks like memory is being resent and rebilled.

Rule 6 — more context is not more accuracy. Dilution, truncation and contradiction are all real and all get worse as you add material.

Rule 7 — context engineering decides the payload; prompt engineering decides the wording. If the question is about phrasing, it is Chapter 10's objective, not this one.

Signal in the question Points at
"Cost rose, traffic flat" Conversation history or injected context growing
"Slow before anything appears" Input length → prefill → time to first token
"Slow to finish" Output length → decode
"Answers drift off-topic since we added documents" Dilution, not model quality
"Needs current or private facts" Context, not retraining
"Same instruction on every call" System instruction, not per-request injection

Distractor Patterns

Pattern What it looks like How to defuse it
Output is free Only the prompt is costed Both legs are metered; output usually costs more per token
A bigger window is cheaper "Move to a larger context window to cut cost" Capacity is not a discount; you pay for what you put in it
Fine-tune to save tokens Retraining offered as a cost fix That is a training cost, and it does not help with current or private facts
More context is more accurate "Inject all the documents" Dilution and truncation both degrade the answer
Latency is about the prompt A shorter prompt offered for a slow finish Prompt length moves first-token time; output length moves total time
Memory is stored server-side "The model remembers the conversation" It is stateless; history is resent and rebilled
Context engineering = prompt engineering Treated as synonyms One writes the instruction; the other decides the payload

The second is the highest-value one to recognise. It sounds like an upgrade, and it is the opposite of a saving.

Scenario Walkthrough

A retailer runs an assistant that answers policy questions. It injects the twenty most similar policy passages into every request, carries the full conversation, and asks for a detailed written explanation each time. Costs have risen sharply while traffic stayed flat, users complain the reply takes a long time to start, and since the policy library grew the answers sometimes cite the wrong region's rules.

Observation Reading Action
Cost up, traffic flat Growth is inside the request, not in usage Summarise history; stop resending whole conversations
Slow to start Large input delays prefill Inject fewer, better-targeted chunks
Wrong region's rules Dilution — twenty chunks outnumber the right one Retrieve fewer, and filter by region
Detailed explanation every time Output leg, at the higher rate Cap the length; ask for detail only when needed

Three different complaints and three different mechanisms. The wrong-region symptom is the one most often mislabelled as a model-quality problem — it is not. The model is doing exactly what it was given.

One change happens to help all three: injecting fewer, better chunks touches cost, prefill and dilution simultaneously. That is a property of this scenario, not a general rule.

Key Concepts

Term Definition
Token-based pricing Charging by the quantity of tokens processed, counted separately for input and output, rather than per request, per user or per unit of time
Input tokens Everything sent to the model — system instruction, conversation history, retrieved context and the user's message
Output tokens Everything the model generates in response, metered separately and typically priced higher per token
System instruction Standing guidance sent in full with every request; looks like configuration, behaves like a recurring charge
Context window The bounded space a request occupies, refilled from empty each time, shared between the input and the room reserved for the answer
Prefill The stage in which the model reads the whole prompt in one pass; governs time to first token
Decode The stage in which output tokens are produced one after another; governs total response time
Time to first token The wait before any output appears, driven mainly by input length
Total response time The wait until the answer is complete, driven mainly by output length
Statelessness The property that nothing carries between requests, so apparent memory is history being resent and rebilled
Context engineering Deciding what information occupies the context window on each request and how it is arranged, so the model can do the job
Dilution Degradation caused by relevant material being outnumbered by irrelevant material in the context
Truncation Silent loss of part of the input when the context window fills
Trim order The deliberate sequence for reducing a request that does not fit: compress history, then fewer chunks, then fewer examples

Revision Flashcards

Say the answer aloud before revealing it.

1. What exactly is token-based pricing charging for? → The quantity of tokens processed, counted separately for input and output. Not per request, not per user, not per unit of time.

2. Which four things make up the input side of a billed request, and which did the user supply? → The system instruction, the conversation history, the retrieved context and the user's message. Only the last came from the user; the other three were added by the application.

3. Which leg is usually priced higher, and why? → Output. Input is read in a single pass, whereas output is produced one token after another, each depending on the ones before it.

4. Why does a long conversation cost more per request, not just more in total? → The model is stateless. Prior turns are resent with every request, so the input grows each turn even when the new message is short.

5. What is the difference between time to first token and total response time? → Time to first token is the wait before anything appears and is driven mainly by input length through prefill. Total response time is the wait until the answer is complete and is driven mainly by output length through decode.

6. A user says the reply is slow to start. Which side of the request do you look at? → The input side. Long input delays prefill. If they had said slow to finish, it would be the output side.

7. Does moving to a model with a larger context window reduce cost? → No. A larger window increases how much you may send; you still pay for every token you actually send. Capacity is not a discount.

8. Which single lever moves both cost and latency? → Capping the response length. It reduces output tokens, which are both the dearer leg and the driver of total response time. The trade-off is terser or truncated answers.

9. Define context engineering in one sentence. → Deciding what information occupies the model's context window on each request, and how it is arranged, so the model can do the job.

10. How is context engineering different from prompt engineering? → Prompt engineering governs the wording of the instruction. Context engineering governs the whole payload around it — what to retrieve, what history to keep, what to drop. Prompt engineering is a component of context engineering.

11. Name three ways that adding more context makes the answer worse. → Dilution, where the relevant passage is outnumbered; truncation, where the window fills and something is silently dropped; and contradiction, where two injected sources disagree and the model answers confidently from the wrong one.

12. What is the trim order when a request does not fit the window? → Compress the conversation history first because it is the most redundant, then inject fewer retrieved chunks, then reduce examples last because they shape the output format.

The Four-Beat Answer

The core question this chapter prepares you for: "How does token-based pricing work, and what do you do about it?"

Four beats, checked in this order. Missing a beat is a failure state — you will be probed on whichever one you skipped.

  1. The unit. Charging is metered by token quantity, counted on both legs, with output typically dearer per token because it is generated sequentially rather than read in one pass.
  2. What is actually in the request. The system instruction, the conversation history, the retrieved context and the user's message — and most of that is the application's own decisions rather than user traffic. This is the beat that makes an answer sound like it came from someone who has built something.
  3. Which number moves. Input length drives prefill and therefore time to first token. Output length drives decode and therefore total response time. Both drive cost. Name the lever and what it trades away.
  4. The discipline. Context engineering: treat the window as a finite budget, allocate it on purpose, and know the trim order for when it does not fit. It is the cheapest way to make a general model do a particular job.

A weak answer describes billing. A strong answer names the levers and what each one costs you.

Why This Helps You

On the job: almost every foundation-model system that becomes unexpectedly expensive got there the same way — a growing system instruction, unmanaged conversation history, and retrieval tuned for recall rather than precision. All three are invisible until the bill arrives, and all three are design decisions rather than traffic.

In interviews: "how would you reduce the cost of this feature?" is a standard system-design follow-up, and "use a cheaper model" is the answer everyone gives. Naming which leg you would cut, which number it moves, and what capability it costs is the answer that sounds like experience.

On the exam: Domain 2 is 24% of scored content, and these two objectives are unusually distractor-rich because every wrong option sounds thrifty. The habit of asking which number does this actually move is worth more here than anywhere else on the paper.

Chapter Checklist

  • I can describe token-based pricing as metered, two-sided and asymmetric
  • I can name the four components of an assembled prompt and say which one the user supplied
  • I can explain why a multi-turn conversation costs more per request over time
  • I can separate prefill from decode, and time to first token from total response time
  • I can say which lever moves cost, which moves latency, and which moves both
  • I can name the trade-off that comes with every lever I propose
  • I can define context engineering without describing prompt engineering
  • I can treat the context window as a budget and state a trim order
  • I can explain three ways in which more context makes an answer worse
  • I can say why context is the cheapest way to specialise a general model
  • I can reject "a bigger context window will reduce our costs" and say why

After the Chapter

  1. Complete student/project.md — parts 19-22 of the AI/ML Decision Sheet you began in Chapter 01. Bring the same sheet; do not start a new one.
  2. Take student/quiz.md closed-book, then review the reasoning for every question you guessed, including the ones you got right. In this chapter especially, a right guess is not a right understanding — the distractors are designed to sound thrifty.
  3. Open the official v1.1 exam guide's Domain 2 page and confirm you can attach a concept from this chapter to the fourth and fifth bullets under Task Statement 2.1.
  4. Next chapter: Chapter 06 — Agentic AI: MCP, Multi-Agent Patterns, Memory, Tools, and Orchestration (Domain 2, Objective 2.1.6). It is the last objective in Task 2.1 and the newest material on the exam. ⚠️ Note the domain carefully: version 1.1 moved the Model Context Protocol out of Domain 3 and into Domain 2, so anyone revising from older material has it filed in the wrong place. Memory management appears on that objective's list — and memory is this chapter's context window put under deliberate control.

Chapter 5 quiz

13 questions on this chapter, marked instantly, with an explanation for every answer.

1. Which statement best describes what a token-based pricing model charges for?
2. A support assistant's costs have tripled over a quarter while the number of conversations has stayed flat. Each conversation now runs to many turns, and the assistant carries the full history. What is the most likely explanation?
3. Why is the output leg of a request typically priced higher per token than the input leg?
4. Users of a document assistant report that the answer takes a long time to *begin* appearing, though once it starts it completes quickly. Which change most directly addresses the complaint?
5. An architect proposes moving to a model with a much larger context window "so we stop paying so much per request." What is wrong with this reasoning?
6. A team writes a long standing instruction that defines tone, scope and refusal rules for its assistant. Which statement about that instruction is correct?
7. Which of the following best defines context engineering?
8. After a company grew its policy library from a few hundred to several thousand documents, its assistant began citing rules from the wrong region, even though the correct policy is retrieved and present in the request. What is happening?
9. A team fills the context window almost entirely with retrieved passages to maximise the chance that the answer is present. What problem does this create, beyond cost?
10. A colleague says: "Context engineering and prompt engineering are two names for the same thing." What is the most accurate correction?
11. Which two of the following reduce the **total response time**, rather than only the time to first token? (Select two.)
12. A request no longer fits within the context window. Which trim order best reflects the reasoning taught for this objective?
13. A team proposes fine-tuning a model on its internal policy documents specifically to avoid sending those documents as context on every request. Which assessment is most accurate for this objective?