- 4.1 — Design prompts with explicit criteria to improve precision and reduce false positives
- 4.2 — Apply few-shot prompting to improve output consistency and quality
- 4.3 — Enforce structured output using tool use and JSON schemas
- 4.4 — Implement validation, retry, and feedback loops for extraction quality
- 4.5 — Design efficient batch processing strategies
- 4.6 — Design multi-instance and multi-pass review architectures
What this course owns, what it defers elsewhere
D0 is assumed. It named the anatomy of a tool — name, description,
input_schema — the tool_use block, and stop_reason. It deliberately stopped
there and pointed here, twice: tool_choice and the use of a JSON schema to
force a conforming payload are taught in 4.3, in full, and D1 defers the same
pair to the same place. That debt is paid below.
The guide places neighbouring notions elsewhere, and this course honours that.
Which wording makes a model pick the right tool among several similar ones is
tool design, tested in TS 2.1 and 2.5; the agentic loop that turns one
tool_use block into a conversation is D1's TS 1.1; running a reviewer inside a
pipeline is D3's TS 3.6, which applies what 4.6 explains. Each of those gets a
sentence and a pointer here, never a second development.
1. 4.1 — Design prompts with explicit criteria to improve precision and reduce false positives
An automated reviewer flags a pattern on Monday and approves the same pattern on Thursday. Nothing changed in the code, and nothing changed in the prompt. What changed is that "be careful about risky migrations" has no fixed meaning, so every run reinvents the threshold.
The remedy the guide names is not a better adjective. It is explicit categorical criteria: named conditions, each one something you could check by hand. The guide's own example is the pair to memorise: "flag comments only when claimed behavior contradicts actual code behavior", against "check that comments are accurate".
# Vague — the version that produces a different verdict every run
Check code comments for accuracy. Be conservative — report only
high-confidence findings.
# Explicit criteria — the version that produces the same verdict twice
Flag a comment ONLY IF one of these holds:
1. It describes behaviour that contradicts what the code does.
2. It refers to a function or variable that does not exist.
3. A TODO or FIXME refers to a bug already fixed in the code.
DO NOT flag:
- dated phrasing or style that does not contradict the code;
- imprecise wording that is still consistent with the code;
- missing comments (a separate category, reported separately).
For each finding, output: file:line, which condition matched, the quoted
comment, and the code fragment that contradicts it.
The two halves are equally load-bearing. Criteria that only say what to report leave everything else to inference; the skip list is what stops a reviewer from drifting into minor style and local convention.
"Be conservative" and "only report high-confidence findings" do not improve precision. The guide states it flatly, and the reason is worth holding: those instructions ask the model to move a threshold it has no shared scale for. "Confidence" is not a property you and the model agree on the way you agree on "the comment names a function that does not exist".
A confidence dial filters; a categorical criterion decides. Only the second one produces the same answer twice.
Severity that classifies consistently
The guide's third skill is severity, and it asks for two things: explicit criteria for each level, and a concrete code example per level. The definition alone leaves "medium" to taste; the worked example fixes it.
CRITICAL — fails at runtime for users.
Example: payment_total = order.get("total") + order["tip"] # KeyError when no tip
HIGH — security vulnerability.
Example: db.execute("SELECT * FROM users WHERE id = " + user_id)
MEDIUM — logic bug with no immediate outage.
Example: for i in range(len(items) - 1): # last item never processed
LOW — code quality only.
Example: the same six-line block duplicated in three handlers.
Why one bad category poisons the good ones
The guide's knowledge point on false positives is a claim about people, not about models: a high false positive rate in one category undermines developer confidence in the accurate categories too. Once the bot has cried wolf about naming conventions three times, its genuine SQL-injection finding is skimmed with the same eye.
That is what makes the guide's remedy counter-intuitive and worth remembering: temporarily disable the high false-positive category, keep the precise ones running, improve the prompt for the disabled one, and re-enable it once its criteria are specific. Reviewing less, on purpose, is how the remaining output keeps being read.
Two signatures, two answers.
Verdicts that vary run to run on identical input → replace the adjective with named conditions. Developers ignoring the tool → look for the one noisy category and turn it off while its prompt is rewritten.
The distractors are the plausible next moves, and each one leaves the criterion undefined. Adding "be conservative", "only high-confidence findings" or a numeric confidence threshold tunes a dial nobody has calibrated. Running the check twice and keeping what both runs agree on intersects two unstable verdicts and drops genuine findings. Narrowing the scope to fewer files reduces the volume without making a single remaining decision consistent. And switching to a larger model answers a specification problem with capacity.
Turning the whole reviewer off because one category is noisy. The guide's move is surgical: the accurate categories keep their value and keep being read. Silence costs you the findings that were correct.
Writing criteria you cannot check. "Flag anything that could confuse a future maintainer" is a categorical sentence in shape only. If you cannot decide the verdict yourself from the rule, neither can the model.
Test yourself on this section
Q1 A rule the reviewer reads differently every run
Scenario: A reviewer inspects database migration scripts under the instruction "flag risky migrations, and be careful". The same DROP COLUMN statement is flagged on Monday and approved on Thursday.
Question: What change makes the verdicts stable?
A) Add few-shot examples of migrations the team called risky, letting the reviewer infer the threshold.
B) Spell out the criteria: a dropped column, an in-place table rewrite, or an unbatched backfill.
C) Run the check twice and keep the findings both runs agree on, taking agreement for stability.
D) Restrict the check to migrations touching production schemas, where an inconsistent verdict actually costs something.
Answer: replace the adjective with named, testable conditions
Why: "risky" carries no operational meaning, so every run reinvents the threshold. Named conditions — flag a migration only when it drops a column, rewrites a table in place, or backfills rows without batching — turn a matter of taste into something the model applies the same way twice.
Why the others are wrong:
- Illustrates a rule that still has not been defined.
- Intersecting two unstable verdicts drops genuine findings.
- Shrinks the volume without making any remaining decision consistent.
2. 4.2 — Apply few-shot prompting to improve output consistency and quality
Explicit criteria fix what to report, and they leave a second failure untouched: the criteria are written, they are detailed, and the output still comes back in a different shape every run, or comes back empty on the case nobody thought to list.
The guide's answer is unambiguous, and it is the phrase to recognise in an option: few-shot examples are the most effective technique for consistently formatted, actionable output when detailed instructions alone produce inconsistent results. One example is one-shot prompting; several is multi-shot.
Two to four is the guide's own count, and the examples are targeted: they show the ambiguous cases, not the easy ones. Each carries the reasoning for why that action was chosen over a plausible alternative, because the reasoning is what transfers.
<examples>
<example>
<request>My order is broken</request>
<action>get_customer, then lookup_order, then check status</action>
<rationale>"Broken" may mean damaged goods or a failed delivery. The
details decide, so gather them before routing anywhere.</rationale>
</example>
<example>
<request>Get me a manager</request>
<action>escalate_to_human, immediately</action>
<rationale>The customer asked for a person explicitly. Attempting a
resolution first overrides a stated preference.</rationale>
</example>
</examples>
The XML tags are the course's habit and they earn their place here: examples are
interpolated content sitting next to instructions, and descriptive delimiters —
<sample_input>, <ideal_output>, <examples> — are what keeps the model from
reading an example as a request.
Examples generalise; lists enumerate. The guide's wording is that few-shot examples let the model "generalize judgment to novel patterns rather than matching only pre-specified cases", and that is the whole reason they beat a longer rulebook. Two examples of an informal measurement — "two handfuls", "a splash" — teach the conversion; a lookup table of informal phrases covers exactly the phrases in it, and the next import brings a new one.
This is also why the rationale matters more than the answer. An example that shows only the output teaches a mapping. An example that shows why teaches a rule.
The four jobs the guide gives them
| Job | What the examples show |
|---|---|
| Ambiguous cases | Which action was taken, and why not the plausible alternative |
| Output format | One filled-in record: location, issue, severity, suggested fix |
| Acceptable versus genuine | A pattern that must not be flagged, beside one that must |
| Varied document structures | The same fact extracted from an inline citation and from a bibliography entry |
The last two are the ones people forget, and each answers a specific failure.
Acceptable versus genuine is the false-positive lever of 4.1 seen from the other side: rather than enumerating every benign pattern, show one benign and one genuine, and let the distinction generalise.
{"location": "src/auth/login.ts:42",
"issue": "SQL injection in the username parameter",
"severity": "critical",
"suggested_fix": "Use a parameterised query"}
Varied document structures is the guide's extraction case. A rate reported as
the rate is 42% (Smith, 2023) and the same rate reported as The rate is 42%. [1] with a bibliography at the end are one fact in two shapes, and a model shown
only the first returns null on the second. The guide names that symptom
directly: examples are what addresses empty or null extraction of required
fields on documents whose format varies.
Two signals in the scenario, together, name this answer: the instructions are already detailed and the output is still inconsistent, and the model has to decide cases nobody listed in advance. Few-shot is the only lever that teaches a distinction instead of enumerating instances.
The distractors are all reasonable-sounding next moves. Rewriting the instructions more firmly repeats what the scenario says has already failed. Maintaining a whitelist or lookup table of accepted patterns covers what is in it and nothing else, and the scenario usually says new phrasings keep arriving. Post-processing the output with a filter cannot reconstruct a decision the model never made. Lowering sensitivity suppresses true positives at the same rate as false ones. And rejecting the hard inputs for manual handling buys correctness by refusing to do the work.
Twenty examples of the easy case. Volume is not the variable; targeting is. Two to four examples aimed at the cases that are actually ambiguous outperform a page of examples the model already handles.
Examples that show the output without the reasoning. They pin the format and teach nothing transferable, so the first unseen phrasing lands back where it started.
Test yourself on this section
Q7 Quantities nobody can enumerate
Scenario: A recipe importer turns ingredient lines into structured quantities. Informal amounts such as "a couple of handfuls" or "a splash" come back empty or wildly wrong, and fresh phrasings appear with every new import.
Question: What is the most effective fix?
A) Lower the extractor's sensitivity so it reports fewer questionable quantities and the wildest values disappear.
B) Keep a lookup table mapping every informal phrase to a quantity, extended by hand as new phrasings appear in the imports.
C) Show two to four worked examples, each one carrying the reasoning that takes an informal phrase to a converted amount rather than the amount alone.
D) Reject ingredient lines carrying no explicit unit and route them to manual entry, so no invented quantity ever reaches a published recipe.
Answer: worked examples that teach the conversion, not a dictionary of phrases
Why: two to four targeted examples are enough, and the reasoning each one carries is what does the teaching: the model reads how the phrase was resolved, not merely what it was resolved to. A rationale shown inside the example teaches a transferable rule, so the model generalizes to phrasings it has never seen — which is what a space nobody can enumerate demands. Written instructions cannot describe this mapping; examples demonstrate it.
Why the others are wrong:
- Suppresses good extractions at the same rate as bad ones.
- Covers only what is listed, and the scenario says the list is never complete.
- Buys correctness by refusing to do the work: every informal line becomes a manual task, and informal lines are most of the corpus.
3. 4.3 — Enforce structured output using tool use and JSON schemas
Ask for JSON in prose and you get JSON most of the time: a trailing comma here, a fenced code block there, a field renamed on the third document. The guide's answer is a mechanism rather than a phrasing, and it states it without hedging: tool use with JSON schemas is the most reliable approach for guaranteed schema-compliant structured output, and it eliminates JSON syntax errors.
The move is a small inversion. You declare a tool whose input_schema is the
shape you want out, you force the model to call it, and you read the extraction
from the call's input. The tool is never executed. Its whole purpose is to
be a schema the model must fill.
{
"tools": [{
"name": "extract_invoice",
"description": "Records the fields extracted from an invoice document.",
"input_schema": {
"type": "object",
"properties": {
"invoice_number": { "type": "string" },
"issue_date": { "type": "string", "description": "ISO 8601, YYYY-MM-DD" },
"total": { "type": "number" },
"company_name": { "type": ["string", "null"] },
"document_kind": { "enum": ["invoice", "credit_note", "other", "unclear"] },
"kind_detail": { "type": ["string", "null"] }
},
"required": ["invoice_number", "total", "document_kind"]
}
}],
"tool_choice": { "type": "tool", "name": "extract_invoice" }
}
The reply carries stop_reason: "tool_use" and a tool_use block whose input
is your record. No parsing of prose, no repair pass:
{ "type": "tool_use", "id": "toolu_01A", "name": "extract_invoice",
"input": { "invoice_number": "INV-2024-0912", "issue_date": "2024-09-12",
"total": 1249.90, "company_name": null,
"document_kind": "invoice", "kind_detail": null } }
tool_choice
This is the field D0 and D1 both deferred here. It decides whether calling a tool is optional, mandatory, or mandatory and named.
| Value | Behaviour | Reach for it when |
|---|---|---|
{"type": "auto"} |
Default. The model decides, and may answer in text | The tool is not always relevant |
{"type": "any"} |
The model must call a tool, and picks which | Several extraction schemas exist and you do not know the document type |
{"type": "tool", "name": "extract_metadata"} |
The model must call that tool | One schema is mandatory, or one extraction must run before the rest |
{"type": "none"} |
No tool may be called | You want prose even though tools are declared |
The two uses the guide names are worth memorising as scenarios, because that is
how they are examined. any is the answer when a pipeline receives invoices,
purchase orders and delivery notes through the same door: you have a schema for
each, you cannot say in advance which applies, and what you cannot afford is a
conversational reply instead of a call. A forced named tool is the answer
when order matters: extract_metadata must run first, and the enrichment steps
come in later turns.
Field behaviour that follows from forcing a call: under any or a forced tool,
the model does not produce a natural-language explanation before the call. If you
want both the reasoning and the call, stay on auto and ask for the tool in the
prompt: "use the extract_invoice tool in your response".
Second constraint, same family: with manually enabled extended thinking, only
auto and none are supported.
Designing the schema so it does not invite invention
A schema is a contract, and a badly drawn contract makes fabrication the compliant answer.
| Rule | Written as | Why |
|---|---|---|
| Optional means optional | Leave the field out of required |
A required field the source does not carry forces the model to choose between breaking the schema and inventing a value |
| Nullable | "type": ["string", "null"] |
Absence needs a representation, otherwise a gift certificate comes back with a plausible price |
enum + "other" + detail |
"enum": [..., "other"] plus a free-text field |
A closed enum produces a wrong category rather than an admission that none fits |
"unclear" |
An enum value of its own | Lets the model say the document does not settle it |
The guide's last skill on this task statement is easy to skip and cheap to apply: format normalization rules go in the prompt, next to the strict schema. The schema says a field is a string; only the prompt can say which string.
Normalise before you fill the schema:
dates → ISO 8601 (YYYY-MM-DD); resolve "yesterday" to an absolute date
money → numeric amount plus a separate ISO 4217 code; "five bucks" → 5, USD
percentages → decimal fraction; "half" → 0.5
absent → null, never a placeholder, never a guess
The guide is the law here. Its TS 4.3 tests tool_use with a JSON schema and
tool_choice. If an option names output_config.format or
client.messages.parse(), it is not the expected answer on this exam. The same
verdict falls on message prefilling paired with a stop sequence: the API course
teaches it as a route to structured output, and this corpus's out-of-scope list
sets it aside on the strength of this task statement. This section is the
promise that entry makes. It is also incompatible with the newer mechanism
below, so the two are never combined in practice either.
Field nuance, so you are not surprised in production: the platform now carries a
dedicated JSON-outputs mechanism (output_config.format, and messages.parse()
with a Pydantic model in the Python SDK) that constrains the response text, plus
strict: true on a tool, which constrains tool arguments by grammar-constrained
sampling rather than by adherence. Same goal, newer surfaces, absent from the
guide.
On that strict path the platform accepts only a subset of JSON Schema, and the
part worth carrying back to ordinary tool use is the failure mode: constraints
like minimum, maxLength and pattern are dropped silently rather than
rejected. A bound you wrote is not a bound that was enforced. If a number has to
stay in range, your code checks it, which is 4.4.
Three phrasings, three answers. "Guarantee a schema-conformant payload" →
tool_use with a JSON schema. "The document type is unknown and several schemas
exist" → tool_choice: {"type": "any"}. "This extraction must run before the
enrichment steps" → force it by name.
The distractors are consistent across questions. tool_choice: "auto" leaves the
model free to answer in prose, so the extraction stops being guaranteed. It is
the closest wrong answer and the one to read for. Describing the schema in the
prompt and asking for JSON reintroduces exactly the syntax errors the mechanism
removes. Repairing the output with a regular expression afterwards treats the
symptom and moves the fragility into your code. And raising max_tokens answers
a shape problem with length.
Believing the schema validated the content. A forced call guarantees the shape of the payload and nothing about the truth of its values: line items that do not sum to the total, a date in the wrong field, an amount that never appeared in the document. The guide says so explicitly, and 4.4 is where that is handled.
Marking every field required for a tidy downstream type. The tidiness is
bought with invented values, which is the more expensive defect.
Test yourself on this section
Q4 A gift with no price on it
Scenario: A museum catalogs donation paperwork into a schema where acquisition_price is a compulsory string. Many pieces arrive as gifts and carry no amount anywhere on the paperwork, yet those records come back with tidy round figures nobody can trace.
Question: What should change?
A) Fill the field with zero when the paperwork names no amount, so every record can satisfy the schema without invention.
B) Let the field accept null and take it out of the compulsory set.
C) Split the intake in two: route gift paperwork to a second schema without the field, and keep purchases on the current one.
D) Have the model report how sure it is of each amount, so catalogers can set aside the values it is least sure of.
Answer: let absence be a legal value, and stop demanding one
Why: as long as the field is compulsory, the model chooses between breaking the schema and producing an amount, and it produces one. A field the source may leave empty needs a representation for empty; a gift then comes back with no price instead of a convincing one.
Why the others are wrong:
- Zero is an amount, and downstream nothing separates a gift from a bargain.
- Whoever routes the paperwork faces the same missing information first.
- The invented figure stays in the catalog, and the field still demands one.
4. 4.4 — Implement validation, retry, and feedback loops for extraction quality
The payload conforms and the values are wrong. Forcing a call through a JSON
schema guarantees the shape of the record and nothing about what is in it:
total reads 150 while the line items sum to 145; a supplier name sits in the
customer field. Nothing in the schema can see it, because nothing in the schema
knows arithmetic.
The mechanism is a loop, and the guide tests its four steps:
- The model produces an extraction.
- Your code validates it, first with JSON Schema for structure, then with business rules for meaning. Pydantic is the appendix's named tool, and it does both: types and required fields in the model, cross-field rules in a validator.
- On failure, you send a follow-up request carrying three things: the original document, the failed extraction, and the specific validation error.
- You track which errors retries actually fix.
Step 3 is the whole of "retry with error feedback", and its adjective is the substance. "There is an error, try again" re-rolls the same dice. The error has to name the field and the gap.
try:
invoice = Invoice.model_validate(tool_use.input) # types, then @model_validator
return invoice
except ValidationError as e:
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": [{
"type": "tool_result",
"tool_use_id": tool_use.id,
"is_error": True,
"content": (
"total is 150.00 but the line_items sum to 145.00. "
"Re-read the line items and resubmit; do not adjust total to fit."
),
}]})
The follow-up rides the ordinary conversation: the assistant message holding the
tool_use block goes back in the history, and the answer to it is a
tool_result block carrying the same tool_use_id, with is_error: true. The
document is already in the history, so the model sees all three of the guide's
ingredients at once.
When a retry cannot work
This is the judgement the task statement is built around, and it is a question about the source document, never about the number of attempts.
| Retry will fix it | Retry cannot fix it |
|---|---|
Format errors — a date written 09/12/24 where the schema wants ISO 8601 |
The information is not in the document at all |
| Structural errors — a value in the wrong field, a list where an object belongs | The figure lives in an appendix or an external document you did not provide |
| Arithmetic inconsistencies the document itself can settle |
The left column shares one property: the model had the information and shaped it
badly. Feeding back the exact error is enough. The right column shares the
opposite one, and a retry there can only produce a more confident invention. The
answer is a nullable field and a null, which is 4.3's design rule arriving as
an operational decision.
Two fields that make the pipeline self-diagnosing
The guide names both, and both are cheap additions to a schema.
stated_total beside calculated_total, plus conflict_detected. Have the
model extract the figure the document declares and the sum it computes from the
lines, then flag the disagreement instead of silently picking one. The conflict
survives into your code, where a human can arbitrate.
{"stated_total": 150.00, "calculated_total": 145.00,
"conflict_detected": true, "line_items": [ … ]}
detected_pattern on every finding. A code-review pipeline that records
which construct triggered each finding turns dismissals into data: when
developers reject findings, you can group the rejections by pattern and see that
one construct produces most of the noise, which is precisely the category 4.1
tells you to disable and rewrite.
{"location": "src/auth/login.ts:42",
"issue": "Possible null dereference",
"severity": "medium",
"detected_pattern": "no optional chaining on user.profile.email",
"suggested_fix": "user?.profile?.email"}
Two failure signals your validation code should not confuse with a bad
extraction. stop_reason: "refusal" means the model declined on safety grounds,
and it comes back with HTTP 200, so code that only checks the status treats a
refusal as a success and parses a non-conforming reply. stop_reason: "max_tokens" means the output was cut off mid-generation; the retry there is a
larger budget, not a better prompt. The full set of stop_reason values is D0's.
The last floor of the pattern, and a guide skill: have the model report a confidence per field, route low-confidence extractions to human review, and measure accuracy by document type and by field. An acceptable overall accuracy can hide one document type that fails systematically.
Read for where the missing information is. "The value appears nowhere in the document" or "that figure is in the contract, which was not provided" → the retry is pointless; make the field nullable and record the gap. "The date came back in the wrong format", "the values do not sum" → the retry works, provided the feedback names the field and the discrepancy.
The distractors on validation questions are neighbouring good practices applied
to the wrong failure. More few-shot examples improve the extraction upstream and
do nothing about a runtime validation failure. A stricter schema removes
malformed payloads, not disagreements between fields. A custom post-processing
parser reimplements what the validator already does. Raising max_tokens bounds
length, not correctness. And retrying with the same prompt, harder, is the
control condition.
Retrying without the error. Resending the document and hoping is a lottery, and it is billed per ticket.
Letting the validator repair the value quietly. Overwriting total with the
computed sum makes the pipeline green and destroys the signal that the document
or the extraction was wrong. Flag the conflict; decide it elsewhere.
Test yourself on this section
Q5 Hours that do not add up (Select the 2 correct answers.)
Scenario: A payroll extractor returns a timesheet whose reported_hours reads 38, while the daily entries sum to 41. The payload parses and matches the schema.
Question: Which two statements fit this situation?
A) It is a syntax error; enforcing the tool schema more strictly makes the payload conform.
B) It is a semantic error; validate programmatically and resubmit on failure with the exact discrepancy named.
C) It is a syntax error; a few more worked examples keep the arithmetic from drifting.
D) It is a semantic error; carry computed_hours beside reported_hours with a conflict_detected flag.
E) It is a syntax error; a higher output token cap leaves room to finish the arithmetic.
Answers: a semantic error, handled by a validate-retry loop and by an explicit conflict flag
Why: a forced tool call guarantees the shape of the payload, never the truth of its values. The retry carries the document, the rejected extraction and the named gap, which lets the model correct itself; carrying both figures plus a flag keeps the conflict visible to downstream code.
Why the others are wrong:
- Schema enforcement removes malformed payloads, not disagreement between fields.
- Confuses shape with meaning; examples do not repair a systematic arithmetic gap.
- The token cap bounds output length, not the correctness of a sum.
5. 4.5 — Design efficient batch processing strategies
The Message Batches API processes large volumes of Messages API requests asynchronously. Three properties decide every question about it:
| Property | Value |
|---|---|
| Cost | 50 % off standard rates |
| Processing window | Up to 24 hours; most batches finish in under an hour |
| Latency SLA | None |
| Correlation | custom_id, set by you on each request |
| Results order | Not guaranteed |
| Multi-turn tool calling in one request | Not supported |
The third row is the one that decides fit, and it is the one distractors attack. "Most batches finish within the hour" is an observation about queues; the twenty-four hours bound how long processing may run. Neither is a promise about any single submission.
POST /v1/messages/batches
{ "requests": [
{ "custom_id": "invoice-10231",
"params": { "model": "claude-sonnet-5", "max_tokens": 1024,
"messages": [{ "role": "user", "content": "Extract: …" }] } },
{ "custom_id": "invoice-10232",
"params": { "model": "claude-sonnet-5", "max_tokens": 1024,
"messages": [{ "role": "user", "content": "Extract: …" }] } }
] }
// Results come back unordered; pair by custom_id, never by position.
{ "custom_id": "invoice-10232", "result": { "type": "succeeded", "message": { … } } }
{ "custom_id": "invoice-10231", "result": { "type": "errored", "error": { … } } }
// → resubmit invoice-10231 alone.
custom_id is not decoration. Because results are unordered, it is the only
reliable way to say which response belongs to which document, which is also what
makes selective resubmission possible.
Which workload goes where
| Workload | API | Why |
|---|---|---|
| Pre-merge check, a developer waiting | Synchronous | Blocking; twenty-four hours is not a possible answer |
| Interactive code review | Synchronous | Immediate response required |
| Overnight technical-debt report | Batch | Wanted by morning; half price |
| Weekly security audit | Batch | Latency-tolerant, high volume |
| Nightly test generation | Batch | Non-blocking by construction |
The criterion is a single question: is anyone waiting? Cost savings never promote a blocking workflow into a batch.
A batch request buys one pass and one answer. You may declare tools and send a multi-turn history, but nothing executes a tool in the middle of a batch request and hands the result back to the model for another turn. Any step whose design requires the model to ask for a file, receive it, and continue does not fit. No amount of polling changes that, because the limitation is structural rather than temporal.
That is the fact behind an entire family of questions: an agentic loop belongs to the synchronous API, and D1's TS 1.1 is where that loop is taught.
Submission cadence under an SLA
The guide asks you to compute one. The arithmetic first: with a 30-hour
commitment and processing that may take up to 24 hours, a submission can wait
at most 30 − 24 = 6 hours in the queue before it is sent.
Six hours is the ceiling, not the answer. A batch submitted exactly at the six hour mark and processed at the cap consumes the entire budget, leaving nothing for a document that failed and must be resubmitted. The guide's own example answers 4-hour windows for a 30-hour SLA, the ceiling minus a margin.
Note which number entered that calculation. Most batches finish in under an hour, and that observation has no place in the arithmetic: you size on the 24 hours the platform bounds, never on the hour you have been getting.
The trap on cadence questions is the arithmetically exact number. Six hours is
correct arithmetic and a wrong operational answer, which is what makes it
attractive. Compute commitment − processing cap to find the ceiling, then take
the option below it.
Handling failures
Four result types come back: succeeded, errored, canceled, expired, and
only succeeded is billed. One failed request does not affect the others.
- Identify failures by
custom_idand resubmit only those. Re-running the whole batch pays again for everything that worked. - Resubmit with a modification, not identically, when the cause is in the
request: a document that exceeded the context limit goes back as chunks. An
invalid_request_errormeans the body must be corrected before it is sent again; a server error can be retried as it stands. - Refine the prompt on a sample first. Dry-run one representative request through the synchronous Messages API, get the extraction right, and only then submit the thousand; otherwise you discover a prompt defect after paying for every document in the batch.
Two operational numbers worth carrying: a batch holds up to 100 000 requests or
256 MB, whichever comes first, and results stay retrievable for 29 days. And a
short list of parameters a batch refuses, each for the same reason, since they
are synchronous-only settings: stream: true (results come back as files),
speed, store and previous_thread_event_id, cache_hint and
context_hint, and max_tokens: 0.
Moving a blocking check to the batch API for the discount. The saving is real, the developer is still waiting, and no latency is promised. A distractor often softens it with status polling; polling reads a state, it does not create a guarantee.
Pairing responses to requests by position. Results are unordered. Code that zips two lists together will mis-attribute extractions silently, and the defect surfaces as data that is subtly wrong rather than as an error.
Test yourself on this section
Q2 A nightly job's clock, borrowed by a live feature
Scenario: A retailer re-translates its whole product catalog every night through the Message Batches API, and the job has come back within twenty minutes every night this quarter. A product manager cites that record to move the storefront's live "translate this review" button onto the same endpoint.
Question: Where does the argument break?
A) A live button submits one request at a time, and the endpoint prices a submission as a unit rather than per token, so a batch holding a single review would forfeit the discount that motivates the move.
B) The nightly run clears in twenty minutes because it lands in the emptiest hours of the queue, while the button would submit all day long, when the same endpoint runs slower.
C) Twenty minutes is an observation, not a commitment: this endpoint publishes no latency guarantee at all, and the twenty-four hours it quotes bound how long processing may run rather than promise when any one submission will be handed back.
D) The endpoint draws no line between a scheduled caller and an interactive one, so the button is entitled to whatever treatment the nightly job already gets.
Answer: a measured time is not a service commitment
Why: no latency is promised here at all. The twenty-four hours describe how long processing may run, not what a given run will take, and last quarter's twenty minutes records the queue those particular nights happened to meet. A shopper waiting in front of a button needs a bound this endpoint never offers.
Why the others are wrong:
- The discount applies per token, so a submission of one still earns it.
- Makes the flaw a matter of which hour it is, when no hour comes with a promised turnaround.
- True of the endpoint, and beside the point: the requirement lives in the workload.
Q6 What the batch endpoint will and will not do (Select the 2 correct answers.)
Scenario: A media team scores the sentiment of 40,000 archived reviews every Sunday night through the Message Batches API.
Question: Which two statements about that run are correct?
A) Results come back in submission order, so a response position is a safe key to its review.
B) Results come back unordered, so each one is paired with its request through the custom_id the team set.
C) A batch request buys one pass and one answer, so a multi-turn tool loop does not fit.
D) The endpoint halves latency along with price, so the weekly job finishes well inside its window.
E) Results are discarded once the batch completes, so the run must stream them out as they land.
Answers: pair results by their correlation key, and expect a single pass per request
Why: the two structural facts that decide whether a workload fits. The correlation key ties each response to its document and lets the team resubmit only the failures; a batch request buys one turn, so a step whose code must execute a tool and hand the result back to the model has nowhere to run.
Why the others are wrong:
- Position is exactly what must not be trusted; results are unordered.
- Confuses price with latency: the discount is real, the latency is not guaranteed.
- Results stay retrievable for 29 days; no streaming workaround is needed.
6. 4.6 — Design multi-instance and multi-pass review architectures
Ask the session that just wrote the code to review it, and it will find little. Not because it is careless, but because of what it carries: the model retains its reasoning context from generation, which makes it less likely to question its own decisions in the same session. It already considered the alternative and rejected it. Re-reading confirms; it does not test.
The guide's remedy is architectural rather than instructional, and the comparison it draws is the point: an independent review instance, with no prior reasoning context, catches subtle issues better than a self-review instruction or extended thinking. Telling the model to be critical of itself, or giving it more room to deliberate, both operate inside the transcript that produced the bias.
A different instance, not a better instruction. The fixes the guide rules out all stay inside the transcript that produced the code: instructing the same session to review its work critically, or giving it extended thinking to deliberate further. Both add effort to a reader that already holds the reasoning it would have to question. The remedy is architectural, so read each option for which instance receives the artifact, not for how hard that instance is asked to look.
"Independent" has an operational meaning: a separate session, or a subagent, which receives the artifact and the criteria and not the conversation that produced them. In Claude Code that is a reviewer subagent, running in its own context window, restricted to read-only tools, and checked into the repository so the team shares one reviewer. D3's TS 3.6 puts the same instance in a pipeline; what belongs here is why the instance has to be a different one.
Splitting a review that is too large for one pass
A pull request touching thirty files reviewed in a single pass produces inconsistent output: a pattern flagged in one file and approved in another, findings that thin out towards the end. The cause the guide names is attention dilution, and the remedy is to shrink each pass rather than to enlarge the context.
| Pass | Scope | Finds |
|---|---|---|
| Per-file | One file at a time | Local bugs, security issues, quality |
| Integration | Signatures and call sites across the changed set | Type mismatches across module boundaries, circular dependencies, broken data flow |
# Pass 1 — local, one invocation per file.
for file in $(git diff --name-only main...HEAD); do
review --scope="$file" \
--prompt="Local issues only: bugs, security, quality in THIS file."
done
# Pass 2 — integration, one invocation over the changed set.
review --scope="$(git diff --name-only main...HEAD)" \
--prompt="Cross-file only: type mismatches across module boundaries,
circular dependencies, broken data flow. Ignore local issues."
Splitting a task into focused passes is D1's task decomposition (TS 1.6) applied to review; what is specific here is the cut, local against cross-file, because those are the two things one undifferentiated pass does worst.
Confidence, self-reported
The third skill is a verification pass in which the model reports a confidence alongside each finding. The value is not the number itself but the routing it enables: high-confidence findings go straight to the developer, low-confidence ones to a human check, and the threshold moves as you observe which band is worth reading.
{"file": "src/billing/refund.ts", "line": 88,
"issue": "Refund may exceed the original charge when a partial refund precedes it",
"confidence": 0.42,
"detected_pattern": "no accumulated-refund check before subtraction"}
Two cautions keep this honest. A confidence reported from inside the generating context inherits that context's blind spot, so the verification pass is worth running in the independent instance. And a self-reported number is a signal for routing, not a measurement: what turns it into one is comparing it against findings whose verdict you already know.
Three symptoms, three answers. Generated code reviewed clean by the session that wrote it → a second instance that never saw the generation. Contradictory findings across files in one large review → split into per-file passes plus an integration pass; name the cause as attention dilution, not missing context. Too many findings for humans to triage → a self-reported confidence per finding, used to route.
The distractors are seductive because each names a real capability. Moving to a model with a larger context window answers dilution with capacity, confusing how much can be held with how evenly it is attended to. Instructing the same session to "review your work critically" or enabling extended thinking on it both stay inside the transcript that caused the bias. The guide rules out exactly these two by name. Three full passes with a consensus vote sounds rigorous and discards every genuine bug that only one pass saw. And asking developers to split their pull requests moves the work onto people without changing the system.
Handing the reviewer the generation transcript "for context". It is the one thing that must not travel: the transcript is the reasoning you want the reviewer not to share. Pass the artifact and the criteria.
Reviewing the diff in the generating session to save a call. The saving is real and the review is worth less than the call it saved.
Test yourself on this section
Q3 Cases graded by whoever invented them
Scenario: A pipeline invents synthetic test cases for a billing rules engine, then asks the same conversation to mark which ones actually break the stated rule. The marking comes back almost entirely clean, yet a QA engineer opens the file and finds cases that break nothing at all.
Question: What restores the marking?
A) Score each generated case for confidence, and review only the weak ones.
B) Keep the marking in the same conversation, but restate the rule in a fresh turn before the model decides, so it judges against the wording rather than its memory.
C) Hand the rule and the bare cases to a second instance that never saw the generation turn.
D) Ask for the marking inline, one verdict written under each case as it is produced, so no case is judged far from the rule that prompted it.
Answer: a second instance carrying none of the generation context
Why: the generator still holds the reasoning that produced each case, so rereading confirms those choices instead of testing them. An instance handed only the rule and the artifact reads them the way an outside grader would.
Why the others are wrong:
- Confidence reported from inside the same context inherits the same blind spot.
- Restating the rule leaves the anchoring transcript exactly where it was.
- Ties the verdict tighter still to the moment of generation.