1. Anatomy of a Messages API request
One call carries everything: client.messages.create(). There is no session to
open and nothing to close.
| Field | Required | What it does |
|---|---|---|
model |
yes | Which Claude model answers the request |
messages |
yes | The conversation history you are sending, oldest first |
max_tokens |
yes | Ceiling on how many tokens the response may contain |
system |
no | Standing instructions, outside the conversation |
tools |
no | The tool schemas the model may call — designed in D2.1 |
tool_choice |
no | Whether the model may answer in prose or must call a tool — developed in D4.3 |
temperature |
no | How much randomness enters token selection, from 0 to 1 |
Two of these are read wrong often enough to be worth stating plainly.
max_tokens is a safety limit, not a target. Set it to 1000 and Claude stops
after 1000 tokens even if it had more to say; it never tries to reach the
number. It bounds the output, not the request you send.
system is a sibling of messages, not one of them. Every other field on this
list is a sibling too: tools and tool_choice sit at the top level of the
body, never inside a message.
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"system": "You are a financial analysis assistant. Answer concisely.",
"messages": [
{ "role": "user", "content": "What was company X's revenue in Q3?" }
]
}
system sits beside messages, never inside it. Sliding the standing
instructions in as a first user message is the most common shape error on
this endpoint, and it changes how the model weighs them.
Exact model identifiers are not exam material. temperature is not either.
The guide never names it, and no task statement turns on it; the row is here so
you recognize the field in a request body you are handed, not to be memorized.
To test — the self-assessment questions of this page.
2. Roles, turns, and a stateless API
messages is a list of dictionaries, each with a role and a content. There
are two roles:
| Role | Who writes it |
|---|---|
user |
You — what you send to Claude |
assistant |
Claude — what it generated |
A conversation alternates between the two: user turn, assistant turn, user turn.
The API stores nothing. Every request is independent and carries no memory of the ones before it. Ask "What is quantum computing?", get an answer, then send "Write another sentence" on its own, and the model has no idea what it is being asked to extend, so it writes a sentence about something else entirely.
Multi-turn conversation is therefore work you do, not a service the API renders:
send the first user message, append Claude's reply to your list as an
assistant message, append the follow-up as a user message, and send the
whole list again.
"messages": [
{ "role": "user", "content": "Define quantum computing in one sentence" },
{ "role": "assistant", "content": "Quantum computing uses quantum states to process information." },
{ "role": "user", "content": "Write another sentence" }
]
Only because the first two turns are still in the array does "Write another sentence" mean anything on the third.
Everything that must still be visible on turn five has to be in the array you send on turn five. This is the rule behind most "the model forgot" bugs, and it applies to tool results exactly as it applies to text. The same rule is why a spawned subagent has to be handed its context explicitly, since it inherits nothing.
Sending only the latest message and expecting the earlier turns to still be there. Nothing expires server-side, because nothing is stored server-side; rebuilding the history is the caller's job, on every single request.
Test yourself on this section
Q3 A stateless API and a vanished tool_result
Scenario: A build pipeline drives Claude across several requests. From one request to the next, a tool_result produced earlier is no longer visible to the model.
Question: What is the most likely cause?
A) The API keeps no state, and the caller did not resend the earlier tool_result blocks.
B) A tool_result block expires server-side soon after it is submitted, so later requests miss it.
C) The max_tokens value is too low to hold the tool_result, so the oldest blocks drop first.
D) A session_id parameter is missing, so the server cannot attach this request to the earlier turn.
Answer: the API is stateless and the history was not resent
Why: nothing carries over between requests. Every call must ship the complete messages array, tool results included; rebuilding that history is the caller's job. The same rule is what forces you to pass context explicitly to a spawned subagent.
Why the others are wrong:
- Nothing expires server-side, because nothing is stored there.
- That budget bounds the output length, not the size of the history you send; and no block is dropped for you, because none is kept for you.
- The Messages API has no such parameter.
3. Content blocks
A reply's content is never a bare string but a list of blocks, each with
its own type. Reading a reply as text works only because you indexed into that
list: message.content[0].text is the first block's text, not the message. The
two directions are not symmetrical: a request accepts either form, a plain
string or a list of blocks.
| Block type | Direction | Carries |
|---|---|---|
text |
both ways | Prose, for the user to read |
tool_use |
from the assistant | The tool name and the arguments the model chose |
tool_result |
from you, in a user message |
The output of the tool you ran |
A single assistant reply routinely holds several blocks at once, for instance a
text block explaining what it is about to do plus a tool_use block requesting
it. A tool_result block carries three fields: tool_use_id, matching the id
of the tool_use block it answers; content, the function output serialized as
a string; and is_error, a boolean. It travels back inside a user message,
alongside the complete history.
{ "role": "assistant", "content": [
{ "type": "text", "text": "Let me look that order up." },
{ "type": "tool_use", "id": "toolu_01X", "name": "lookup_order",
"input": { "order_id": "ORD-67890" } }
]},
{ "role": "user", "content": [
{ "type": "tool_result", "tool_use_id": "toolu_01X",
"content": "{ \"status\": \"shipped\", \"total\": 42.90 }",
"is_error": false }
]}
Between those two messages, your code ran lookup_order. The model never
executes anything; it emits a structured request and waits.
Appending only the text you displayed and dropping the tool_use block. The
pairing with tool_result is what the next request depends on, so the history
becomes unanswerable. Append all of response.content, and keep the
original tool schemas on the follow-up request even when no further call is
expected.
A tool is declared as a JSON schema with three parts: name, description —
what the tool does, when to use it, what it returns — and input_schema for the
arguments. That is the anatomy, and it stops there. Which wording makes the
model pick the right tool among several similar ones is tool design, developed
in D2.1; using a schema to force a conforming payload is structured output,
developed in D4.3.
4. stop_reason
When the model stops, the response says why in stop_reason, and that field,
not the presence of a text block, is what tells you whether a reply is finished.
| Value | What ended generation |
|---|---|
end_turn |
The model reached a natural end and had nothing more to say |
max_tokens |
The output budget you set ran out mid-generation |
stop_sequence |
One of your stop_sequences strings appeared |
tool_use |
The model is calling one of your tools and is waiting for you to run it |
pause_turn |
A server-tool loop hit its iteration limit, 10 per request by default |
refusal |
The model declined for safety-policy reasons |
model_context_window_exceeded |
The response filled the model's context window |
stop_details is null for every value but refusal, where it carries the
policy category. model_context_window_exceeded is the recent one: Sonnet 4.5
and newer models return it without a beta header, and earlier models need the
model-context-window-exceeded-2025-08-26 header to enable it.
{
"id": "msg_01ABC",
"role": "assistant",
"content": [
{ "type": "text", "text": "Let me look that order up." },
{ "type": "tool_use", "id": "toolu_01X", "name": "lookup_order",
"input": { "order_id": "ORD-67890" } }
],
"stop_reason": "tool_use",
"stop_details": null
}
That response carries text and a tool request, which is exactly why the presence of text cannot serve as a completion signal.
end_turn is the only value that means complete. A response with
stop_reason: "max_tokens" holds a text block like any other, but it was cut
off mid-sentence. Display code that checks for text instead of checking
stop_reason will present a truncated draft as a finished one. That is the
classic distractor on this field.
Two values are worth telling apart even though only one pair is tested. A reply
waiting on your tool always says tool_use, and you continue it by sending
tool_result blocks. pause_turn concerns a server tool that ran out of
iterations, and you continue it by sending the assistant content back unchanged.
The guide only ever asks you to tell tool_use from end_turn. That pair
drives the loop, and it is all TS 1.1 tests. Turning the check into a loop —
call, inspect stop_reason, execute, append, call again — is the agentic loop,
developed in D1.1. D0 stops at naming the signal.
Test yourself on this section
Q1 A stop_reason your display code ignores
Scenario: A drafting agent produces long documents. Your display code shows the API response to the user whenever content holds a text block. On one long document the response cuts off mid-sentence, yet the code still presents it as the finished draft. Inspecting the response, stop_reason is "max_tokens".
Question: What does that value mean, and what should the display code check before treating a response as final?
A) The response is complete; "max_tokens" caps the size of the request being sent, not the length of the output.
B) The token budget ran out mid-generation; only "end_turn" means the response is complete.
C) The model refused to continue for safety reasons, so nothing should be shown and the attempt should be escalated for human review instead.
D) A server-side tool call was interrupted mid-turn, so resending the same request unchanged lets the paused turn pick up where it left off.
Answer: "max_tokens" means truncated output, not completion
Why: stop_reason distinguishes several reasons generation stopped, and only "end_turn" marks a natural, complete stopping point. "max_tokens" means the output budget ran out mid-generation — the presence of a text block in content says nothing about whether that text is the whole answer. Scope note: among stop reasons, the exam guide names only the pair that drives the agentic loop, and that loop is what the task this sheet prepares tests. "max_tokens", "refusal" and "pause_turn" come from the API reference — D0 prerequisite material, not extra exam surface.
Why the others are wrong:
- Reverses cause and effect: the output token budget is exactly why this value was returned.
- Describes
"refusal", a distinct value with a distinct cause. Nothing was refused here, so withholding the draft and escalating would stall a run that only needs continuing. - Describes
"pause_turn", which concerns an interrupted server tool, not an exhausted token budget. Nothing is paused waiting to resume, so an identical resend would simply regenerate the same truncated draft.
5. System prompts
The system parameter is a plain string passed alongside messages, and it
sits outside the conversation: it is not a turn, it has no role, and it is
not something the model answers. It shapes how Claude responds, not what it
responds.
Take a maths tutor. Without a system prompt, "How do I solve 5x + 2 = 3 for x?" gets a complete step-by-step solution immediately, which is correct and useless for teaching. With:
You are a patient math tutor.
Do not directly answer a student's questions.
Guide them to a solution step by step.
the same question comes back as "What do you think would be a good first step to isolate x?". Same model, same user turn; the standing instruction changed the behaviour.
That is what a system prompt is for: assign a role, and the model answers the way someone in that role would; state durable constraints, and they hold across every turn without being restated. They hold because you resend them every request, not because the API remembers them, and you pay for them every time.
Putting the system prompt in as the first user message. It then reads as
conversational context rather than a standing instruction, and it competes with
the turns around it instead of sitting above them.
A system instruction carries further than any per-turn wording, far enough that a single absolute keyword can override tool descriptions that are otherwise well written. "Always verify the customer's identity" will drag the nearest-matching tool onto every turn. How that overriding works, and how to word around it, is developed in D2.1.
One practical note: the API does not accept system=None, so a reusable chat
helper has to add the key only when a prompt was actually given.
6. The context window
The context window is the finite token budget every request must fit into, and the two previous facts set it against you: the API keeps no state, so every request ships the entire history. The history is therefore not free. It is a recurring cost that grows monotonically with the conversation, unless you do something about it.
What accumulates, per request:
| What you resend | Why it is there |
|---|---|
The system prompt |
It holds across turns because you resend it, not because it is stored |
Every past user and assistant turn |
Nothing is stored server-side |
Every tool_use and tool_result block |
Dropping them breaks the pairing |
| The full tool schemas | Required on every request, even when no call follows |
Tool results are the part people underestimate. A tool that returns a large record puts that record into the history for the rest of the run, and you pay for it on every subsequent turn, not once.
lookup_order → 42 fields returned, 5 useful for return eligibility
order_id, status, total, items, return_eligible ← useful
warehouse_route, carrier_scac, pick_wave, audit_flags,
promo_ledger, tax_jurisdiction, … (37 more) ← dead weight here
3 orders looked up = 126 fields resent on every later turn, 15 of them earning it.
The guide names three failure modes of a full window, and D0 goes exactly as far as naming them: lost-in-the-middle, where material buried in the middle of a long context is missed while the start and the end are read reliably; accumulation, the growth shown above; and progressive summarization, where compressing the history to make room trades away the numbers, dates and specifics first.
Recognize the three failure modes here; the fixes are not D0's. Trimming verbose tool outputs, extracting structured facts, and ordering input by position are developed in D5.1. Retrieval and chunking are not the answer the guide asks for. It puts them out of scope, and so does this corpus.
7. Self-assessment questions
Q2 Which part of a tool definition decides the routing
Scenario: An HR assistant exposes search_policy_doc ("Retrieves policy content") and search_payroll_record ("Retrieves payroll data"). Both accept employee identifiers of a similar shape, and the model routinely picks the wrong one.
Question: Which part of a tool definition drives that choice?
A) The position of each tool in the request's tools array.
B) The tool name, which outranks the rest of the definition.
C) The description, which is the primary selection mechanism.
D) The size of the input_schema: the model prefers the simplest one.
Answer: the description
Why: the model selects a tool by reading the definitions, and the description is where the meaning lives. Two one-line descriptions give it nothing to tell the tools apart, so misrouting becomes structural. The fix is to state input formats, example requests, edge cases and the boundary with neighboring tools.
Why the others are wrong:
- Declaration order is not a selection criterion.
- The name helps, but it cannot carry input formats, edge cases or usage boundaries.
- The schema describes the arguments once a tool has been chosen; it does not decide which tool to choose.
Q4 A standing rule that drags a tool onto every turn
Scenario: A greenhouse operations assistant opens its system prompt with "Every answer must reflect the current sensor readings." Logs show read_sensors firing on each turn, including when a grower asks how a cultivar name should be spelled in the planting log. Each tool definition spells out its inputs and where its job ends.
Question: What stops the needless calls?
A) Add few-shot examples of clerical turns answered without any sensor call, so the assistant can copy the pattern on the next spelling question.
B) Have the client drop read_sensors from the tools array whenever it judges the incoming question clerical.
C) Name the conditions that call for a reading, and retire the absolute sentence.
D) Cache the last reading for ten minutes so repeat turns reuse it instead of calling out again.
Answer: state when a reading is required, and drop the blanket rule
Why: "every answer" reads as a permanent trigger, and the model binds it to whichever tool the wording resembles most. A rule phrased with no exceptions outranks the tool definitions, however careful they are — which is why the scenario tells you those definitions are sound. Spelling out the situations that genuinely need a reading removes the trigger while keeping the check where it earns its cost.
Why the others are wrong:
- Examples cannot outvote a rule that still admits no exception; the model treats them as odd cases rather than as the boundary.
- Moves the decision to the caller, which must now classify a question before the model has read it, and leaves the misleading rule in place.
- Cheaper calls, identical behavior: a spelling question still travels through the sensor path, and a stale value is the wrong answer where readings matter.
Q5 A triage loop that never writes its answer
Scenario: A ticket-triage client runs its own loop: send the request, run whatever tool comes back, append the result, send again. To stop the assistant from chatting while work remains, the client sets tool_choice to the value that guarantees a call — and it sets it on every request of the loop. Runs now end only when the client's iteration cap trips, and no written answer ever arrives.
Question: Why does the run never reach one?
A) That value binds the first request alone; the later ones revert to the automatic setting, so the assistant keeps reaching for lookup_account long after the answer is available.
B) Each tool result is appended as a user turn, and a request whose last turn came from the user can only be answered by another call.
C) Repeating that value on every request removes the assistant's only way of finishing, which is a turn that ends in text; the client should guarantee the opening call and leave the remaining requests on the automatic setting.
D) The iteration cap sits below the number of declared tools, so the client cuts the run short before classify_ticket and its neighbours have all been tried.
Answer: the setting also forbids the text-only turn a run has to end on
Why: tool_choice: "any" promises exactly one thing, on every request that carries it: a tool will be called. Sent once, that pins the opening step. Sent on every request, it also outlaws the closing one, since a turn answering in prose is no longer permitted. Handing the later requests back to tool_choice: "auto" restores the choice between calling again and answering.
Why the others are wrong:
- The field is read afresh on each request; it neither expires nor decays, and the loop here restates it every time.
- A user turn carrying a tool result can be answered in prose just as readily as with another call — that is precisely what the automatic setting permits.
- The cap is what halts a run that would otherwise not halt at all; raising it buys further calls, never a written answer.
Q6 A mandatory field the source cannot fill
Scenario: An invoice extraction tool lists discount_pct among its required properties, yet plenty of invoices grant no discount. Reviewers keep finding discount figures that appear nowhere on the invoice.
Question: What is happening, and what removes the cause?
A) The API validates the payload against the schema and rejects the call as invalid, so nothing reaches the pipeline.
B) The model invents a figure to satisfy the constraint; make the property nullable and no longer mandatory.
C) The model leaves the property empty, and the pipeline should read an empty value as absent.
D) The model skips the tool call whenever a mandatory property has no source value.
Answer: the model invents a figure, so make the property nullable and optional
Why: a mandatory property leaves the model choosing between breaking the schema and making something up, and it makes something up. Declaring it "type": ["number", "null"] and taking it out of required gives the model an honest way to report absence.
Why the others are wrong:
- Validation covers schema conformance, never truthfulness: an invented number conforms perfectly.
- Nothing empties the property while the schema still demands a value.
- The tool does get called; it is the payload that is wrong.