- 1.1 — Design and implement agentic loops for autonomous task execution
- 1.2 — Orchestrate multi-agent systems with coordinator-subagent patterns
- 1.3 — Configure subagent invocation, context passing, and spawning
- 1.4 — Implement multi-step workflows with enforcement and handoff patterns
- 1.5 — Apply Agent SDK hooks for tool call interception and data normalization
- 1.6 — Design task decomposition strategies for complex workflows
- 1.7 — Manage session state, resumption, and forking
What this course owns, what it defers elsewhere
D0 is assumed. This course starts where D0 stopped: D0 named stop_reason and
said what each value means; D1 turns that check into a loop, then into a system
of agents around it.
The guide places neighbouring notions elsewhere, and this course honours that.
tool_choice and JSON Schemas are tested in TS 2.3 and TS 4.3; structured error
responses in TS 2.2 and TS 5.3; the .claude/agents/ files and the /agents
command are Claude Code configuration, tested in domain 3. Each of those gets a
sentence and a pointer here, never a second development.
1. 1.1 — Design and implement agentic loops for autonomous task execution
A single Messages API call produces one answer. A real task — find the customer, look up the order, decide on a refund — needs several actions, each chosen in the light of what the previous one returned. The agentic loop is the mechanism: request, response, tool execution, new request, until the model says it is done.
The cycle, five steps:
- You send a request with the conversation history and the
toolsparameter. - Claude answers. The reply may hold text, one or more
tool_useblocks, or both. - Your code executes each requested tool and collects the results.
- You append the assistant reply and the
tool_resultblocks to the history. - You send the whole history again. Repeat until
stop_reasonis terminal.
stop_reason is the field on the response that says why generation stopped;
its full set of values is catalogued in D0. Only two of them drive this loop:
stop_reason |
What your code does |
|---|---|
"tool_use" |
Execute the requested tools, append their results, iterate |
"end_turn" |
The model has nothing more to ask for — leave the loop |
while True:
response = client.messages.create(
model="claude-sonnet-5",
messages=messages,
tools=tools,
)
# The whole reply goes back, text blocks and tool_use blocks together.
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
break # terminal — D0 catalogues the other values
tool_results = execute_tools(response.content)
messages.append({"role": "user", "content": tool_results})
Two lines carry the whole task statement. The break is on stop_reason, on
nothing else. And the second append is what makes the next iteration a
reasoning step rather than a repetition: tool results enter the conversation
history, so the model's next choice of tool is made in the light of what the
last one returned. A loop that runs the tools but never appends their results
asks the same question forever.
Several calls in one reply
One reply can also carry several tool_use blocks at once, for instance two
lookups the model wants done together. Run them all, then send every
tool_result back in a single user message, one block per call. That is the
same mechanism a coordinator uses to start several subagents at once, developed
in 1.3.
The loop is model-driven: Claude reads the accumulated context and decides
which tool comes next. That is what separates an agent from a workflow. A
pre-configured decision tree — your code fixing the sequence
read_schema → plan_changes → apply_migration in advance — is legitimate
engineering, but it can only ever handle the cases whoever wrote it foresaw.
The guide tests your ability to tell the two apart and to pick by how
predictable the task is.
The guide names three wrong ways to drive the loop, and they are the distractors on this task statement:
- Parsing the assistant's text for "task completed" or "I'm done". Natural
language is not a signal; Claude can write "done" while a
tool_useblock is still pending in the same reply. - Using an arbitrary iteration cap (
max_iterations = 5) as the primary stopping mechanism. A cap is a guardrail against a runaway agent, not a statement about whether the work is finished, and it cuts legitimate runs short. - Checking for text content as a completion indicator. Text and
tool_usetravel in the same reply; the presence of one says nothing about the other.
Caps are guardrails, not stop conditions
Caps do have a place, as long as they stay caps. The Agent SDK exposes
max_turns / maxTurns (agentic round-trips) and max_budget_usd /
maxBudgetUsd (spend); both default to no limit, and a budget is a sound
production default on an open-ended prompt. Neither is a termination condition,
and the SDK marks a run that hit one as an error subtype rather than a success.
Whenever a question asks how to decide that the loop should continue or stop,
the answer is stop_reason, and specifically "tool_use" versus "end_turn".
Every plausible-looking alternative in the option list — text parsing, an
iteration ceiling, a check for a text block — is one of the three anti-patterns
above, dressed up.
Test yourself on this section
Q4 Who decides the next tool call
Scenario: A migration agent needs to run read_schema, then plan_changes, then apply_migration for each table. To keep the process predictable, the orchestrator pre-computes this exact call sequence before the first request and feeds each tool result into the next hardcoded step, never letting the model's own response decide what happens next. When a table turns out to need an extra validation step the sequence doesn't account for, the agent has no way to trigger it.
Question: What should change about how the loop is driven?
A) Keep the fixed sequence, but add a fourth hardcoded step for validation, since the gap that appeared today is now a known case.
B) Let the model choose the next action each turn from the latest tool_result, and stop when it stops requesting tools.
C) Run all three tools on every turn regardless of need, so the missing step is always covered and no table can slip through, accepting the wasted work as the price of completeness.
D) Keep the fixed sequence, but let the orchestrator retry a failed step before moving to the next one.
Answer: hand control back to the model, turn by turn
Why: a sequence fixed in advance by the orchestrator is a hardcoded decision tree — it can only be as flexible as whoever wrote it anticipated. An agentic loop is model-driven: the model reads each tool result and decides the next action itself, the loop continues for as long as it keeps requesting tools, and it ends when the model signals it is done. That is what lets it insert a step nobody hardcoded.
Why the others are wrong:
- Fixes this one gap but leaves the orchestrator guessing every future case in advance — the same limitation returns with the next table that needs something different.
- Running every tool regardless of need does not add judgment, it just hides the missing decision behind unconditional execution — and the completeness it buys is an illusion, since it still only covers steps that were written down.
- A retry addresses a step failing, not a step nobody planned for.
2. 1.2 — Orchestrate multi-agent systems with coordinator-subagent patterns
Delegating to a subagent is a decision that comes before any architecture. A subagent runs in its own context window, does its work, and returns only a summary; the intermediate reads, searches and tool calls are discarded with it. That trade has one question: does the intermediate work matter to the caller? If you only need the answer, delegate. If you need to see and react to each step, keep it in the main thread.
The hub-and-spoke shape
When a task is genuinely too broad for one agent — a research brief needing web search, document analysis and synthesis — you get a hub-and-spoke architecture: one coordinator (the hub) drives several specialized subagents (the spokes), and all traffic goes through the hub.
Coordinator (hub)
/ | \
search subagent analysis synthesis
(web) (documents) (report)
every result, every error, every retry
travels hub ↔ spoke — never spoke ↔ spoke
Nothing in the SDK forces that shape. It is a design choice, and the guide names its three motives: observability (one place sees the whole run), consistent error handling (one policy, not one per agent), and controlled information flow (the hub decides what each spoke is told).
What the coordinator owns
The coordinator's job, as the guide states it:
| Responsibility | What it means in practice |
|---|---|
| Task decomposition | Cut the request into subtasks that jointly cover it |
| Dynamic delegation | Select which subagents a given query actually needs, rather than routing every query through the full pipeline |
| Result aggregation | Collect the spokes' outputs and assemble the answer |
| Error handling | Decide to retry, adapt the request, or continue with partial results |
Two skills the questions turn on
Two skills are worth stating on their own because they are what the scenario questions turn on.
Partition the scope before delegating. Give each search subagent a distinct subtopic or a distinct source type. Overlapping briefs produce duplicated work and, worse, correlated blind spots.
Refine iteratively. The coordinator reads the synthesis output, looks for gaps, re-delegates to the search and analysis subagents with targeted queries, and re-invokes synthesis until coverage is sufficient. This is a quality loop sitting above the execution loop of 1.1, and it is the guide's named remedy for incomplete coverage.
Subagents start from an isolated context: they do not inherit the coordinator's conversation history, and they share no memory between invocations. Everything a spoke needs travels in the prompt the hub writes for it, and the mechanics of that passing are TS 1.3, below.
Decomposition that is too narrow
The risk the guide names by name is overly narrow decomposition by the coordinator. Asked for a brief on "AI impact on creative industries", a coordinator splits it into "AI in digital art", "AI in graphic design" and "AI in photography". Each subagent executes its brief correctly. The report covers visual art only; music, literature and film are simply absent. Nothing failed downstream. The assignment was already incomplete.
When every subagent did what it was asked and the final report still misses whole areas of the topic, the defect is the coordinator's decomposition, not any spoke. The distractors will accuse a real component — search queries too narrow, no gap detection in synthesis, relevance criteria too strict in analysis — and they are seductive because they name something that exists. The scenario has already told you each agent executed its assignment correctly. What was done badly is the assigning.
Letting one subagent hand its output straight to another to save a hop. It looks efficient and it breaks the three things the hub was for: no single place sees the run, error handling stops being uniform, and nobody controls what the second agent was actually told.
Routing every query through the full pipeline is the mirror mistake: the coordinator is supposed to analyze the query and invoke only the subagents it needs.
Errors travel through the hub too
Errors travel the same way. The expected pattern is local recovery first: the subagent handles transient failures itself and propagates to the coordinator only what it cannot resolve, together with its partial results and what it already tried. The coordinator continues on partial results and annotates the final output with what stayed uncovered. The shape of those error payloads — categories, retryability, business versus permission failures — is structured error design, developed in D2.2.
Test yourself on this section
Q2 Forty pages about one of five areas
Scenario: A coordinator was asked for a pre-launch risk brief on a payments feature. The request names five areas to cover: regulatory exposure, fraud, support load, vendor dependencies, data residency. It mentions chargebacks once, as an example under fraud. Six workers ran over two rounds and each returned a well-formed section; the brief that came back is forty pages, all of them about chargebacks. The coordinator's plan is still in the log, and it opens by restating the job as chargeback exposure.
Question: Which change keeps this from recurring?
A) Gate the delivery step: hold the brief back until every area named in the request has a section of its own, and send the job round again whenever one of them is missing.
B) Split the work into five independent runs, one per area, and have the requester stitch the five briefs together afterwards.
C) Have the coordinator derive one assignment per area named in the request, before it delegates.
D) Raise the worker count so each round covers more ground.
Answer: derive the assignments from the areas the request names, ahead of any delegation
Why: the plan in the log names the defect — a five-part request became a one-part job before any worker started. Coverage that was never assigned cannot be recovered by anything downstream, so the fix has to sit ahead of the delegation.
Why the others are wrong:
- The gate proves the brief incomplete only once six workers have been paid for, and it hands the job back to the same planning step that read five areas as one — so the next round can return just as narrow.
- Guarantees the coverage by removing the coordinator from the job. Whatever sits across two areas is lost, and the requester now does the merging on every request.
- Capacity was never the constraint. Six workers already produced forty pages; more of them widen nothing while every assignment points at the same area.
3. 1.3 — Configure subagent invocation, context passing, and spawning
The spawning mechanism is a tool: a coordinator spawns a subagent by calling the
Task tool, and its allowedTools must include "Task" for those
invocations to be auto-approved. Leave it out and the call falls back to the
permission flow: under the default mode, with no canUseTool callback to answer
it, it is refused. The coordinator does not delegate, whatever its prompt says.
Recent SDK versions renamed this tool from Task to Agent, and current code
emits Agent in tool_use blocks. The exam guide v1.0 still says Task, and
names including "Task" in allowedTools as the requirement. On the exam,
Task is the expected answer. In real code, write Agent and test for both
names when you detect an invocation.
What crosses, and what does not
What a subagent gets, and what it does not, is the single most tested fact of this domain:
| The subagent receives | The subagent does not receive |
|---|---|
| Its own system prompt, from its agent definition | The parent's conversation history and tool results |
| The prompt string of the spawning call | The parent's system prompt |
| Its tool definitions — inherited, or the subset it was given | Any memory of a previous invocation of itself |
The consequence is decisive: the only content that crosses from parent to subagent is the prompt string of the spawning call. If file paths, error messages, prior search results or decisions already taken are not in that string, the subagent does not have them.
Bad Task: "Analyse the document."
→ which document? no prior findings, no expected output shape.
Good Task: "Analyse the document below and return the clause table.
DOCUMENT (source: vendor-contract-2026.pdf, pages 4-9):
<full text>
PRIOR FINDINGS (from the search subagent):
- claim: 'net-45 payment terms'
source_url: https://… retrieved: 2026-08-14
OUTPUT FORMAT: one row per clause — clause_type, text, page."
Note what the good version does beyond being long. It uses a structured format that separates content from metadata. Source URL, document name, page number and retrieval date live in their own fields, not woven into prose. That is what preserves attribution all the way to the final report, and the guide names it as a skill of this task statement.
Nothing is inherited: the prompt string is the whole channel. A subagent that returns a thin or off-target answer was under-briefed, not under-resourced. That sorts the options in one move: enlarging its context window, handing it the parent's transcript, or letting it read another subagent's work all treat isolation as a defect to be worked around, when it is the mechanism itself. Whatever the subagent needs goes into the spawning prompt.
Parallelism: one reply, several calls
Parallelism comes from one response, several calls. A coordinator that
emits three Task calls in a single reply gets three subagents running at the
same time; the same three calls spread over three consecutive turns run one
after another. This is the same mechanism as parallel tool use in 1.1 — several
tool_use blocks in one assistant message — applied to the Task tool.
agents = {
"code-reviewer": AgentDefinition(
description="Expert code review specialist. Use for quality, security "
"and maintainability reviews. Tell the agent precisely "
"which files to review.",
prompt="You are a code review specialist…\n"
"Return: 1. Summary 2. Critical issues 3. Major issues "
"4. Minor issues 5. Approval status 6. Obstacles encountered",
tools=["Read", "Grep", "Glob"],
model="sonnet",
),
}
The four fields of an agent definition
Four fields, four decisions:
| Field | Required | What it decides |
|---|---|---|
description |
yes | When the coordinator launches this subagent — and what it writes in the delegation prompt |
prompt |
yes | The subagent's system prompt: its role, its method, its output format |
tools |
no | The tools it may use; omitted, it inherits everything available to subagents |
model |
no | opus, sonnet, haiku, inherit, or a full model id |
The description has that double role and it is easy to miss. The name and
description of every available subagent are placed in the main agent's system
prompt, so they decide when delegation happens; the main agent then uses the
description as guidance while writing the input prompt, so they also shape
what the subagent is told to do. A description that says "you must tell the
agent precisely which files to review" produces delegation prompts that list
files. A vague description produces vague delegations.
Two more design levers. The first comes from the Subagents course; the second is the fourth skill the guide lists under this task statement:
- Define an output format in the system prompt. It gives the subagent a checklist, so it knows when it is finished; without one, subagents struggle to decide when enough research is enough and run far longer than needed. An "Obstacles encountered" section is worth adding on purpose: workarounds and environment quirks the subagent discovered otherwise die with its context and have to be rediscovered by the main thread.
- Write coordinator prompts as goals and quality criteria, not procedures. "The brief must cover at least five sectors with dated sources" leaves the subagent free to adapt; a step-by-step script does not.
Restricting a subagent's tools
Tool restrictions belong here too, since the guide lists them among this task
statement's knowledge. A read-only research subagent gets Read, Grep,
Glob; a reviewer adds Bash to run git diff; only an agent whose job is to
change code gets Edit and Write. A tool you leave out is not in the
subagent's session at all: no permission prompt, no error, the agent simply
works without it. Why a large tool set degrades selection reliability, and
how to scope tools across agents, is tool distribution, developed in D2.3.
Three moves that look like fixes and are not:
- Enlarging the subagent's context window when it produced a thin answer. Nothing was passed to it; the size of an empty window changes nothing.
- Spawning across successive turns and calling it parallel. Consecutive turns are sequential execution with extra round-trips.
- Counting on inheritance, on the idea that "the subagent will see what the coordinator found". Isolation is the defining property of a subagent, not a limitation to work around.
Forking a session to give two subagents a shared analysis baseline is session
management, developed in 1.7 below. Defining subagents as markdown files
under .claude/agents/, and the /agents command that creates them, is Claude
Code configuration, developed in domain 3.
Test yourself on this section
Q7 Delegating a contract review (Select the 2 correct answers.)
Scenario: A coordinator has a clause-extraction subagent read two vendor contracts while a pricing subagent prepares a quote, and both outputs must land in one report.
Question: Which two moves are correct?
A) Put the contract text, the clause types wanted, and the output schema into the delegation prompt itself.
B) Emit both delegations in one coordinator response so the subagents run at the same time.
C) Count on each subagent inheriting the coordinator's conversation history.
D) Let the clause subagent hand its output straight to the pricing subagent, saving a hop.
E) Give both subagents the coordinator's whole tool set so neither is ever blocked.
Answers: carry the needed context in the delegation prompt, and issue both delegations in one response
Why: a subagent starts from an isolated context, so whatever it needs — source text, target clause types, output shape — travels in the prompt. Parallelism comes from several delegations inside one coordinator response; spread over consecutive turns they just run one after another.
Why the others are wrong:
- Nothing is inherited: an isolated context is the defining property of a subagent.
- Every exchange goes through the hub, which keeps the run observable and error handling uniform.
- A wider tool set degrades selection reliability; each spoke gets only the tools its role needs.
4. 1.4 — Implement multi-step workflows with enforcement and handoff patterns
Some orderings in a workflow must hold every single time it runs: verify the customer's identity before processing a refund; check drug interactions before dispensing. You can write that in the system prompt, and a prompt instruction is a request the model usually honours. The guide's formulation is blunt: when deterministic compliance is required, prompt instructions alone have a non-zero failure rate. Over thousands of interactions, that remainder is refunds paid to the wrong customer.
The prerequisite gate
Programmatic enforcement closes it by making the violation impossible in code rather than improbable in prose. The concrete shape is a prerequisite gate: a downstream tool call is blocked until an upstream step has returned.
async def require_verified_customer(input_data, tool_use_id, context):
if input_data["tool_name"] != "process_refund":
return {}
if not session_state.get("verified_customer_id"):
return {"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "deny",
"permissionDecisionReason":
"Call get_customer first: process_refund requires a verified "
"customer ID for this session.",
}}
return {}
process_refund cannot run until get_customer has returned a verified id.
Not "is unlikely to" — cannot. The reason string matters as much as the block:
the model reads it, so it retries the right thing instead of retrying blindly.
The hook API itself — events, matchers, decision fields — is TS 1.5, below.
Deterministic compliance is a property of code, not of wording. The wrong answers here are usually genuine improvements: a firmer instruction, few-shot examples showing the right order, an audit that catches violations afterwards. The first two lower the failure rate and leave a remainder; the third reports the breach after the action it was meant to stop. When the scenario names a financial, legal or safety consequence, only a mechanism that makes the violation impossible clears the requirement.
The multi-concern request
A second pattern of this task statement is the multi-concern request. A customer writes: "I want a refund on order #1234 and the delivery address changed on order #5678." Handling that sequentially and answering in two disconnected blocks is the failure. The expected move is to decompose it into distinct items, investigate each in parallel against the shared customer context, then synthesize one unified resolution.
The structured handoff
The third is the structured handoff. When the agent escalates mid-process, the human who takes over has no access to the conversation transcript. The summary therefore has to stand alone:
{
"customer_id": "CUST-12345",
"order_id": "ORD-67890",
"issue_summary": "Refund request for damaged item",
"root_cause": "Item arrived damaged; photos attached",
"actions_taken": [
"Identity verified via get_customer",
"Order confirmed via lookup_order",
"Replacement offered — customer insists on refund"
],
"refund_amount": "$89.99",
"recommended_action": "Approve full refund",
"escalation_reason": "Customer asked to speak with a manager"
}
Customer identity, root cause, amount, recommended action: the four the guide names. A handoff that assumes the human will read the conversation is a failed handoff.
The distractors on enforcement questions are the measures that genuinely lower the failure rate without ever reaching zero:
- Strengthening the system prompt to say "verification is mandatory".
- Adding few-shot examples showing the verification tool called first.
- Auditing transactions afterwards to catch violations.
- Deploying a routing classifier that only enables the relevant tool subset.
The first three lower the rate; the question asked for a guarantee, and an audit finds the violation after the money is gone. The fourth addresses which tools are available, not the order in which they are called, which is a different problem.
The rule to carry in: financial, legal or safety consequences → programmatic enforcement. Preferences, tone and formatting → prompt instructions. If the scenario contains the word "guarantee", or describes an irreversible action, no amount of prompt engineering is the answer.
When to escalate at all — the criteria, honouring a customer's stated preference, spotting a policy gap — is escalation decision-making, developed in D5.2.
Test yourself on this section
Q1 A dose released before the interaction check
Scenario: A pharmacy agent fills prescription requests through MCP tools (check_interactions, dispense, page_pharmacist). Reviewers pulled last month's transcripts and found doses released to patients whose interaction check never ran: in those requests the agent took the allergy list pasted into the ticket at face value and went straight to dispense.
Question: Which change makes the skipped check impossible?
A) Have the agent write out the sequence it intends to follow before its first call, so a skipped step shows up in the transcript.
B) Cache interaction results per patient so the check costs almost nothing and the agent has no reason to route around it.
C) Refuse dispense in code until check_interactions has returned for that patient.
D) Add a hook that rejects malformed prescription identifiers before the outgoing call leaves the agent.
Answer: refuse the downstream call in code until the upstream check has returned
Why: an ordering that guards a patient-safety boundary needs a guarantee, not a tendency. A prerequisite evaluated in code cannot be argued out of by a convincing-looking allergy list sitting in the ticket.
Why the others are wrong:
- A stated plan is still the model's own output, nothing holds it to what it wrote, and the transcript gets read after the dose is out.
- Cost was never the obstacle — the agent skipped a step it believed it already had the answer to.
- Enforces a property of the arguments, not the fact that anything ran first. A well-formed identifier says nothing about the check.
5. 1.5 — Apply Agent SDK hooks for tool call interception and data normalization
A hook is a callback of yours that the SDK runs when an agent event fires. It executes in your application process, not in the agent's context window — so it costs no tokens, and it is ordinary code with ordinary guarantees.
Two events carry this task statement:
| Event | Fires | Can it still change the outcome? |
|---|---|---|
PreToolUse |
A tool call has been requested, before it runs | Yes — allow, deny, or rewrite the input |
PostToolUse |
The tool has returned, before the model reads the result | The call already happened; you can replace what the model reads |
PreToolUse — making something impossible
PreToolUse is the enforcement primitive. Its callback returns a
hookSpecificOutput carrying permissionDecision ("allow", "deny",
"ask", "defer"), a permissionDecisionReason the model reads, and
optionally updatedInput to rewrite the call rather than refuse it.
async def cap_refunds(input_data, tool_use_id, context):
if input_data["tool_name"] != "process_refund":
return {}
if input_data["tool_input"].get("amount", 0) > 500:
return {"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "deny",
"permissionDecisionReason":
"Refunds above $500 require a human approver. "
"Call escalate_to_human with the case summary instead.",
}}
return {}
The block is half the work; the reason redirects the agent to the alternative
workflow instead of leaving it to guess. When several hooks and permission
rules disagree, the SDK resolves them by precedence — deny beats defer
beats ask beats allow — so a single deny is enough.
PostToolUse — changing what the model reads
PostToolUse is the normalization primitive. Its updatedToolOutput field
replaces the tool's output before Claude sees it; additionalContext adds to
it instead.
async def normalize_dates(input_data, tool_use_id, context):
if input_data["tool_name"] not in ORDER_TOOLS:
return {}
raw = json.loads(input_data["tool_response"]["content"])
return {"hookSpecificOutput": {
"hookEventName": "PostToolUse",
"updatedToolOutput": json.dumps({
**raw,
"shipped_at": to_iso8601(raw["shipped_at"]), # 1743800400 → "2025-04-04"
"status": STATUS_NAMES[raw["status"]], # 2 → "shipped"
}),
}}
That is the guide's case, literally: several MCP tools return dates as Unix timestamps, others as ISO 8601, statuses as numeric codes. Mapping them is a fixed table, not a judgement call — so it belongs in code that runs on every result, not in a prompt asking the model to cope with three formats.
The choice between a hook and a prompt instruction is the same choice as in 1.4, one level down:
| Hook | Prompt instruction | |
|---|---|---|
| Guarantee | Deterministic — it is code that runs | Probabilistic — no guarantee |
| Use for | Business rules, financial and regulatory limits | Preferences, tone, formatting |
| Example | Block refunds above $500 | "Try to resolve the issue before escalating" |
Reaching for PostToolUse to stop something. It fires after the tool has run,
by which point the refund is paid and the rows are deleted. Blocking is
PreToolUse.
The mirror error is a PreToolUse block used to fix a reading problem. If a
tool returns an overloaded field the model misinterprets, blocking the
downstream call suppresses the correct calls too; the fix is to rewrite the
result on the way back so the ambiguity never reaches the model.
Two questions separate the two hooks cleanly. Must something not happen? →
PreToolUse. Must the model read something different from what the tool
returned? → PostToolUse. Distractors will place the right mechanism at the
wrong point in the loop.
Claude Code's own configuration-level hooks, declared in settings rather than in SDK code, are developed in domain 3.
Test yourself on this section
Q6 Five hundred rows, and then a signature
Scenario: A retention agent trims expired records through purge_records. Policy lets it delete up to 500 rows in one call; a larger purge needs a data steward's sign-off. The rule is written in the system prompt, and last week a single call removed 40,000 rows from an audit table.
Question: What makes the ceiling hold?
A) Give the agent a request_signoff tool, with a description telling it when a purge has grown too large to run alone.
B) A hook that blocks the outgoing call when the row count passes the ceiling.
C) State the ceiling in the tool's own description, so the limit travels with the tool wherever it is used.
D) Have the agent announce the row count it is about to delete before it issues the call.
Answer: intercept the outgoing call and stop it above the ceiling
Why: a ceiling that must hold every time belongs in code that sees the actual argument. The hook runs before the tool does, so an oversized purge never reaches the database, and a deletion is not something an apology can undo.
Why the others are wrong:
- Offers a way out that the agent still has to choose, and the incident was precisely a case where it did not.
- Moves the rule closer to the call site without changing its nature: it stays text the model reads, not a condition anything evaluates.
- An announcement is a claim about an intent, made by the same run that is about to ignore it.
Q8 One integer, three meanings
Scenario: A storefront agent checks availability through check_stock before it confirms an order through place_order. The available field carries three meanings in one integer: a positive number is units on hand, -1 marks an item the supplier ships on request, and 0 marks a SKU the warehouse does not track. The agent treats anything that is not positive as sold out, and it has spent a month declining orders that could have been filled.
Question: Where does the fix belong?
A) Ask the inventory team to retire the overloaded integer and publish an explicit status field, since one number carrying three meanings will mislead every consumer of that API and not only this agent.
B) A PreToolUse hook on place_order that blocks the confirmation whenever the availability figure behind it was not positive.
C) Log every non-positive figure and review the pattern each week.
D) A hook that rewrites the figure into an explicit state before the model reads the result.
Answer: map the integer onto an explicit state in a hook, before the model reads the result
Why: the three readings are a fixed table, so applying them is a transformation and not a judgment call. A PostToolUse hook runs it on every result on the way back, which is the last moment at which the ambiguous integer can be kept out of what the model reads.
Why the others are wrong:
- The right thing to fix and the wrong thing to wait for: the field belongs to another team, every existing consumer has to move with it, and orders keep being declined until that lands.
- Hardens the misreading rather than removing it. The confirmations it would block are precisely the ones that should go through, and the agent is still reading the same integer.
- Counts the symptom every week and changes nothing about what the model reads. The mapping is already known, which is what makes measurement the wrong instrument.
6. 1.6 — Design task decomposition strategies for complex workflows
Not every task splits the same way, and picking the wrong regime costs you either rigidity or chaos. There are two.
Fixed sequential pipeline — prompt chaining. Every step is decided in advance and runs in order. Use it when the structure of the task is predictable, the steps are all known, and you want stable, comparable results run after run.
Dynamic adaptive decomposition. Subtasks emerge from what the previous step found. Use it for open-ended investigation, where the scope is unknown at the start.
The criterion is predictability, not size. A twelve-file pull request has a known, enumerable set of targets and takes a fixed pipeline. A one-file bug of unknown origin does not.
The guide's two examples
The guide's own example of prompt chaining is the multi-aspect code review. A single pass over a large diff produces uneven output: detailed comments on some files, superficial on others, and contradictions between them. Split it:
Pass 1 — per file, one focused analysis each
→ local bugs, quality issues, vulnerabilities
Pass 2 — one cross-file integration pass over the results of pass 1
→ inconsistent types, circular dependencies, broken data flow
Splitting this way is what avoids attention dilution on a large review.
The guide's example of adaptive decomposition is "add comprehensive tests to a legacy codebase". You cannot list the subtasks up front, because they are what you are about to discover:
1. Map the structure (Glob, Grep)
2. Discover: 3 modules untested, 2 partially covered
3. Prioritize by impact → payments module first (high risk)
4. Surprise: an unmocked external API dependency
5. Adapt the plan → build the mock, then write the tests
The guide's wording is: map structure, identify high-impact areas, produce a prioritized plan that adapts as dependencies surface. Step 4 is the whole point: the plan changed because of what step 1 found.
Three mismatches, and they are exactly the distractors:
- A fixed pipeline on unknown scope demands the list of steps before you start, which is precisely what you do not have.
- A multi-pass review offered for an open-ended investigation is built for a bounded, enumerable set of targets like the files of a PR.
- A single exhaustive prompt meant to cover everything at once assumes you already know the scope.
And the fourth, quieter one: adaptive decomposition on frozen targets. It buys nothing once the targets never vary, and it makes successive runs harder to compare.
Read the scenario for one thing: is the list of targets knowable before the first call? Yes → prompt chaining. No → dynamic decomposition. Words like "the same six adapters, every quarter" point one way; "nobody knows how many stages it has" points the other.
Test yourself on this section
Q3 Two backlog items, two decomposition strategies
Scenario: Each quarter you audit the same six connector adapters against the same three concerns: retry policy, secret handling, schema versioning. Separately, you must make an undocumented ingestion pipeline observable; nobody knows how many stages it has or which already emit metrics.
Question: Which strategy fits each item?
A) Prompt chaining for the audit, dynamic adaptive decomposition for the observability work.
B) Dynamic adaptive decomposition for the audit, prompt chaining for the observability work.
C) Prompt chaining for both, because a fixed pipeline is always more reproducible.
D) Dynamic adaptive decomposition for both, because letting the model replan is always safer.
Answer: prompt chaining for the audit, adaptive decomposition for the pipeline
Why: the criterion is whether the scope is known before you start. The audit repeats a fixed set of targets, so a chained sequence of focused passes gives even depth and comparable results. The observability work is open-ended: each stage you uncover changes what to look at next.
Why the others are wrong:
- Inverts the criterion: the fixed sequence lands on the item whose steps nobody knows, the adaptive one on the item that never varies.
- A fixed pipeline presupposes knowing the stages and their order, precisely the unknown here.
- Adaptivity buys nothing once targets are frozen, and it makes successive audits harder to compare.
7. 1.7 — Manage session state, resumption, and forking
A session is the conversation history the SDK accumulates while the agent works — your prompt, every tool call, every result, every reply — written to disk automatically. Long investigations span several working sessions, and between them files change and stored tool results go stale.
Three operations, three different jobs:
| Operation | What it does |
|---|---|
--resume <session-name> |
Continues one specific prior conversation, by name |
fork_session |
Creates an independent branch from an existing session's history; the original is untouched and the fork gets its own id |
| Fresh session + injected summary | Starts empty and receives, in the prompt, a structured summary of what still holds |
# Fork the analysis baseline to explore a second approach.
async for message in query(
prompt="Instead of Redux, outline how the Context API would work here",
options=ClaudeAgentOptions(
resume=session_id, # fork always branches an existing session
fork_session=True,
max_turns=5,
),
):
if isinstance(message, ResultMessage):
forked_id = message.session_id # distinct from session_id
Two facts about forking
Two facts about forking that questions turn on. It always combines with resumption, since you fork an existing session, never out of nothing. And it branches the conversation history, not the filesystem: two forked agents editing files in the same directory are editing the same files.
Choosing between the three
Choosing between the three is the skill the guide tests:
| Situation | Choice |
|---|---|
| Prior context still valid, files unchanged | --resume |
| A few specific files changed, the rest still holds | --resume, naming the changed files so the agent re-analyses those rather than re-exploring everything |
| Tool results are stale — files changed extensively or unverifiably | New session, seeded with a structured summary |
| Comparing two approaches from a shared analysis baseline | fork_session |
The third row is the guide's own judgement call, and it is counter-intuitive enough to be tested: starting a new session with a structured summary is more reliable than resuming with stale tool results. A resumed session carries recorded responses that describe a system that no longer exists, and the agent reasons on them as if they were current. A written summary carries the conclusions that survived and leaves the dead payloads behind.
A session is a record of what was read, not the state of what is on disk. Every choice among the three turns on how much of that record is still true. Resumption fits while the record is mostly valid, and naming the files that changed is what makes it targeted rather than a gamble. A fork branches the conversation history and not the filesystem, which is why it compares two approaches and never two versions of a file. A fresh session with a written summary keeps your conclusions and leaves the expired readings behind.
Three moves the distractors put in place of the three answers above:
- Resuming without saying what changed. The agent keeps a false picture of the files and reasons on dead tool results, the very failure resumption was supposed to avoid.
- Re-exploring everything from scratch when only three files moved. It throws away work that is still valid.
- Forking to compare an old and a new version of a file. A fork explores divergent approaches from a shared baseline; it does not reconcile a disk state.
Three situations, three answers, and the distractors trade them around. Context
mostly valid → --resume plus a list of the changed files. Tool results
largely stale → fresh session with an injected summary. Two approaches to
compare → fork_session. Note also that the guide says --resume <session-name> while the SDK documents resume taking a session id; the
concept — resuming a named investigation across working sessions — is what is
tested, not the syntax.
Test yourself on this section
Q5 The vendor retired the version you probed
Scenario: A session named fx-integration spent a long afternoon probing a currency vendor's sandbox: sample payloads, error shapes, throttling behavior. The vendor has since shipped v3 and retired v2, so every recorded response describes an endpoint that now answers 410. What that session concluded about the vendor's pagination model and its throttling window still holds.
Question: How do you take the integration work forward?
A) Start a new session and seed it with a summary of what still holds.
B) Resume the session and let the agent re-probe each endpoint as it needs one, so the recorded v2 responses get corrected along the way.
C) Start a new session and replay the entire v2 transcript into it, so nothing observed that afternoon is lost.
D) Start a new session with only the vendor's v3 changelog.
Answer: a new session seeded with what survived the version change
Why: resumption pays off while the stored tool results still describe the system you are talking to. Here they describe endpoints that answer 410, so replaying them makes the agent reason about a vendor that no longer exists. A summary carries the pagination model and the throttling window forward and leaves the dead payloads behind.
Why the others are wrong:
- Re-probing corrects the record only where the agent happens to look; everything it does not revisit sits in context as though it were still true.
- Carries the obsolete payloads into the new session, which is the one thing a fresh start was meant to avoid.
- A changelog says what moved on the vendor's side, not what you learned. Pagination and throttling would have to be rediscovered.