Practice Questions — CCA-F

112 QCM corrigés · énoncés, options et corrigés en anglais, comme à l'examen · les 6 scénarios officiels et les 30 task statements sont couverts

Pourquoi cette page est en anglais. L'examen CCAR-F est administré en anglais. Traduire l'entraînement, c'est s'entraîner à un exercice que vous ne passerez pas : le jour J, le coût de lecture porte sur des formulations anglaises précises — most effective first step, root cause, deterministic, proportionate — qui portent à elles seules la moitié de la discrimination entre options. Les commentaires de méthode restent en français.

Table des matières
  1. Méthode et distracteurs récurrents
  2. Questions officielles du guide (Q1–Q12)
  3. Scénario 1 — Customer Support Resolution Agent (Q13–Q28)
  4. Scénario 2 — Code Generation with Claude Code (Q29–Q42)
  5. Scénario 3 — Multi-Agent Research System (Q43–Q58)
  6. Scénario 4 — Developer Productivity with Claude (Q59–Q72)
  7. Scénario 5 — Claude Code for Continuous Integration (Q73–Q86)
  8. Scénario 6 — Structured Data Extraction (Q87–Q100)
  9. Questions transversales et multi-réponses (Q101–Q112)
  10. Couverture — les 30 task statements

1. Méthode et distracteurs récurrents

Comment lire un item

  1. Lire le scénario avant les options. Formulez mentalement votre réponse, puis cherchez-la dans la liste. Lire les options d'abord, c'est se laisser guider par le distracteur le plus séduisant.
  2. Repérer le qualificateur de la question. Most effective, first step, root cause, most maintainable : ces mots tranchent entre deux options toutes deux « correctes ». Une solution juste mais disproportionnée perd contre une solution juste et proportionnée quand la question dit first step.
  3. Compter les réponses attendues. Chaque item indique combien en sélectionner.
  4. Ne jamais laisser blanc — il n'y a pas de pénalité.

Les sept distracteurs qui reviennent

Formulation type Pourquoi c'est faux
« Strengthen the system prompt to state that X is mandatory » Une consigne en langage naturel a un taux d'échec non nul. Dès qu'il y a une conséquence financière ou réglementaire, la réponse est un programmatic enforcement (hook, gate de prérequis).
« Deploy a classifier trained on historical data » Sur-ingénierie : on ajoute de l'infrastructure et un besoin de données annotées avant d'avoir essayé la correction proportionnée.
« Have the model self-report a confidence score, then route below a threshold » Proxy non calibré : le modèle est confiant précisément là où il se trompe. Attention — field-level confidence calibrée sur un jeu étiqueté est, elle, la bonne réponse pour router une relecture.
« Detect customer frustration and escalate on negative sentiment » Le sentiment ne corrèle pas avec la complexité du cas.
« Terminate the workflow when a subagent fails » On perd les résultats partiels. La réponse est la dégradation gracieuse avec annotation des lacunes.
« Switch to a model with a larger context window » Une fenêtre plus grande ne corrige ni la dilution d'attention, ni la dégradation de contexte, ni le lost-in-the-middle.
« Move the blocking check to the Message Batches API » Jusqu'à 24 h de traitement, aucune garantie de latence : incompatible avec un workflow bloquant.

2. Questions officielles du guide (Q1–Q12)

Note de fidélité : ces douze items reproduisent les questions d'exemple du guide officiel — énoncés, options et bonnes réponses. Leurs bonnes réponses sont majoritairement « A » : c'est un artefact du guide, pas un biais de conception. Ne devinez jamais à la lettre, travaillez le raisonnement.

Scenario: Customer Support Resolution Agent

Q1 Deterministic tool ordering

Scenario: Production data shows that in 12% of cases, your agent skips get_customer entirely and calls lookup_order using only the customer's stated name, occasionally leading to misidentified accounts and incorrect refunds.

Question: What change would most effectively address this reliability issue?

A) Add a programmatic prerequisite that blocks lookup_order and process_refund calls until get_customer has returned a verified customer ID.

B) Enhance the system prompt to state that customer verification via get_customer is mandatory before any order operations.

C) Add few-shot examples showing the agent always calling get_customer first, even when customers volunteer order details.

D) Implement a routing classifier that analyzes each request and enables only the subset of tools appropriate for that request type.


Answer: Add a programmatic prerequisite that blocks the downstream calls

Why: when a specific tool sequence is required for critical business logic — verifying identity before moving money — programmatic enforcement provides a deterministic guarantee that prompt-based approaches cannot. This is Task 1.4: prerequisite gates versus prompt-based guidance.

Why the others are wrong:

  • Relies on probabilistic LLM compliance. A 12% failure rate with financial consequences is exactly the case where "usually follows the instruction" is not good enough.
  • Also relies on probabilistic LLM compliance, for the same reason: a 12% failure rate with financial consequences is exactly the case where "usually follows the instruction" is not good enough.
  • Addresses tool availability, not tool ordering. The agent has the right tools; it calls them in the wrong order.
Q2 Minimal tool descriptions

Scenario: Production logs show the agent frequently calls get_customer when users ask about orders (e.g. "check my order #12345"), instead of calling lookup_order. Both tools have minimal descriptions ("Retrieves customer information" / "Retrieves order details") and accept similar identifier formats.

Question: What’s the most effective first step to improve tool selection reliability?

A) Add few-shot examples to the system prompt demonstrating correct tool selection patterns, with 5–8 examples showing order-related queries routing to lookup_order.

B) Expand each tool's description to include input formats it handles, example queries, edge cases, and boundaries explaining when to use it versus similar tools.

C) Implement a routing layer that parses user input before each turn and pre-selects the appropriate tool based on detected keywords and identifier patterns.

D) Consolidate both tools into a single lookup_entity tool that accepts any identifier and internally determines which backend to query.


Answer: Expand each tool's description

Why: tool descriptions are the primary mechanism an LLM uses for tool selection. Minimal descriptions leave the model without the information it needs to tell two similar tools apart. Option B addresses that root cause with a low-effort, high-leverage fix (Task 2.1).

Why the others are wrong:

  • Adds token overhead without fixing the underlying ambiguity — the descriptions stay uninformative.
  • Is over-engineered and bypasses the model's own language understanding.
  • Is a defensible architectural choice, but it is far more work than a first step warrants when the immediate problem is inadequate descriptions.
Q3 Escalation calibration

Scenario: Your agent achieves 55% first-contact resolution, well below the 80% target. Logs show it escalates straightforward cases (standard damage replacements with photo evidence) while attempting to autonomously handle complex situations requiring policy exceptions.

Question: What’s the most effective way to improve escalation calibration?

A) Add explicit escalation criteria to your system prompt with few-shot examples demonstrating when to escalate versus resolve autonomously.

B) Have the agent self-report a confidence score (1–10) before each response and automatically route requests to humans when confidence falls below a threshold.

C) Deploy a separate classifier model trained on historical tickets to predict which requests need escalation before the main agent begins processing.

D) Implement sentiment analysis to detect customer frustration levels and automatically escalate when negative sentiment exceeds a threshold.


Answer: Explicit escalation criteria plus few-shot examples

Why: the root cause is unclear decision boundaries, and the proportionate fix is to state the boundaries and demonstrate them on ambiguous cases. Criteria alone leave the ambiguous cases open; examples alone leave the rule implicit — the guide's answer is the combination (Task 5.2).

Why the others are wrong:

  • Fails on its own premise: the agent is already incorrectly confident on hard cases, so its self-rated confidence is exactly the signal you cannot trust here.
  • Is over-engineered — labelled data and ML infrastructure before prompt optimisation has been tried.
  • Solves a different problem. Sentiment does not correlate with case complexity, which is the actual issue.

Scenario: Code Generation with Claude Code

Q4 Where a shared slash command lives

Scenario: You want to create a custom /review slash command that runs your team's standard code review checklist. This command should be available to every developer when they clone or pull the repository.

Question: Where should you create this command file?

A) In the .claude/commands/ directory in the project repository

B) In ~/.claude/commands/ in each developer's home directory

C) In the CLAUDE.md file at the project root

D) In a .claude/config.json file with a commands array


Answer: .claude/commands/ in the project repository

Why: project-scoped commands live in the repository, so they are version-controlled and automatically available to everyone who clones or pulls (Task 3.2). The distinction project-scope versus user-scope is the whole point of the item.

Why the others are wrong:

  • ~/.claude/commands/ is the personal scope — not shared through version control, so a new teammate would not get it.
  • CLAUDE.md carries project instructions and context, not command definitions.
  • Describes a configuration mechanism that does not exist in Claude Code.
Q5 Plan mode versus direct execution

Scenario: You’ve been assigned to restructure the team's monolithic application into microservices. This will involve changes across dozens of files and requires decisions about service boundaries and module dependencies.

Question: Which approach should you take?

A) Enter plan mode to explore the codebase, understand dependencies, and design an implementation approach before making changes.

B) Start with direct execution and make changes incrementally, letting the implementation reveal the natural service boundaries.

C) Use direct execution with comprehensive upfront instructions detailing exactly how each service should be structured.

D) Begin in direct execution mode and only switch to plan mode if you encounter unexpected complexity during implementation.


Answer: Enter plan mode

Why: plan mode exists for exactly this profile — large-scale change, multiple valid approaches, architectural decisions, many files. It allows safe exploration and design before committing to changes (Task 3.4).

Why the others are wrong:

  • Risks expensive rework when dependencies surface late.
  • Assumes you already know the right structure without having explored the code.
  • Treats the complexity as hypothetical. It is stated in the requirements, not something that might emerge.
Q6 Conventions for files spread across directories

Scenario: Your codebase has distinct areas with different coding conventions: React components use functional style with hooks, API handlers use async/await with specific error handling, and database models follow a repository pattern. Test files are spread throughout the codebase alongside the code they test (e.g. Button.test.tsx next to Button.tsx), and you want all tests to follow the same conventions regardless of location.

Question: What’s the most maintainable way to ensure Claude automatically applies the correct conventions when generating code?

A) Create rule files in .claude/rules/ with YAML frontmatter specifying glob patterns to conditionally apply conventions based on file paths

B) Consolidate all conventions in the root CLAUDE.md file under headers for each area, relying on Claude to infer which section applies

C) Create skills in .claude/skills/ for each code type that include the relevant conventions in their SKILL.md files

D) Place a separate CLAUDE.md file in each subdirectory containing that area's specific conventions


Answer: .claude/rules/ with glob patterns in frontmatter

Why: a rule scoped with paths: ["**/*.test.tsx"] loads on the file being edited, wherever it sits in the tree. That is precisely the requirement — conventions that follow a file type, not a directory (Task 3.3).

Why the others are wrong:

  • Relies on the model inferring which section applies, rather than on explicit matching.
  • Skills are invoked on demand; the requirement says "automatically", which rules out anything requiring a deliberate invocation.
  • Directory-bound CLAUDE.md files cannot cleanly cover files scattered across many directories — you would duplicate the same conventions everywhere.

Scenario: Multi-Agent Research System

Q7 Reports that cover only part of the topic

Scenario: After running the system on the topic "impact of AI on creative industries," you observe that each subagent completes successfully: the web search agent finds relevant articles, the document analysis agent summarizes papers correctly, and the synthesis agent produces coherent output. However, the final reports cover only visual arts, completely missing music, writing, and film production. When you examine the coordinator's logs, you see it decomposed the topic into three subtasks: "AI in digital art creation," "AI in graphic design," and "AI in photography."

Question: What is the most likely root cause?

A) The synthesis agent lacks instructions for identifying coverage gaps in the findings it receives from other agents.

B) The coordinator agent's task decomposition is too narrow, resulting in subagent assignments that don't cover all relevant domains of the topic.

C) The web search agent's queries are not comprehensive enough and need to be expanded to cover more creative industry sectors.

D) The document analysis agent is filtering out sources related to non-visual creative industries due to overly restrictive relevance criteria.


Answer: The coordinator's task decomposition is too narrow

Why: the logs give the answer directly — three subtasks, all visual. The subagents executed their assignments correctly; the defect is in what they were assigned (Task 1.2, risks of overly narrow decomposition).

Why the others are wrong:

  • Blames the synthesis agent, which is working correctly within the scope it was given: it can only report gaps between the findings it receives, and no gap-detection instruction recovers music, writing and film production when no subagent was ever asked about them.
  • Blames the search agent, which also worked correctly within its assignment. Broader queries cannot widen three subtasks that are all about visual arts — the scope was fixed before any query was issued.
  • Blames the document analysis agent, which summarised its papers correctly. There were no non-visual sources for it to filter out, because nothing upstream ever collected any.
Q8 A subagent times out

Scenario: The web search subagent times out while researching a complex topic. You need to design how this failure information flows back to the coordinator agent.

Question: Which error propagation approach best enables intelligent recovery?

A) Return structured error context to the coordinator including the failure type, the attempted query, any partial results, and potential alternative approaches.

B) Implement automatic retry logic with exponential backoff within the subagent, returning a generic "search unavailable" status only after all retries are exhausted.

C) Catch the timeout within the subagent and return an empty result set marked as successful.

D) Propagate the timeout exception directly to a top-level handler that terminates the entire research workflow.


Answer: Structured error context

Why: the coordinator can only make a good recovery decision — retry with a narrower query, switch source, or proceed with partial results and annotate the gap — if it receives the material to decide with (Task 5.3).

Why the others are wrong:

  • The generic status throws away exactly the context the coordinator needs. Local retries are fine; collapsing the outcome to "unavailable" is not.
  • Silent suppression — a failure reported as success. The coordinator concludes there is nothing to find and the gap never surfaces.
  • Terminates the whole workflow over one failed branch, discarding every result already obtained.
Q9 Round-trips for fact verification

Scenario: During testing, you observe that the synthesis agent frequently needs to verify specific claims while combining findings. Currently, when verification is needed, the synthesis agent returns control to the coordinator, which invokes the web search agent, then re-invokes synthesis with results. This adds 2-3 round trips per task and increases latency by 40%. Your evaluation shows that 85% of these verifications are simple fact-checks (dates, names, statistics) while 15% require deeper investigation.

Question: What's the most effective approach to reduce overhead while maintaining system reliability?

A) Give the synthesis agent a scoped verify_fact tool for simple lookups, while complex verifications continue delegating to the web search agent through the coordinator.

B) Have the synthesis agent accumulate all verification needs and return them as a batch to the coordinator at the end of its pass, which then sends them all to the web search agent at once.

C) Give the synthesis agent access to all web search tools so it can handle any verification need directly without round-trips through the coordinator.

D) Have the web search agent proactively cache extra context around each source during initial research, anticipating what the synthesis agent might need to verify.


Answer: A scoped verify_fact tool for the common case

Why: least privilege applied to tool distribution — the synthesis agent gets exactly what covers the 85% case, and the existing coordination path is preserved for the 15% that genuinely needs it (Task 2.3).

Why the others are wrong:

  • Batching creates blocking dependencies: later synthesis steps often depend on facts verified earlier.
  • Over-provisions the synthesis agent with tools outside its specialisation, which is the documented cause of tool misuse.
  • Speculative caching cannot reliably predict which claims will need checking.

Scenario: Claude Code for Continuous Integration

Q10 The pipeline job hangs

Scenario: Your pipeline script runs claude "Analyze this pull request for security issues" but the job hangs indefinitely. Logs indicate Claude Code is waiting for interactive input.

Question: What’s the correct approach to run Claude Code in an automated pipeline?

A) Add the -p flag: claude -p "Analyze this pull request for security issues"

B) Set the environment variable CLAUDE_HEADLESS=true before running the command

C) Redirect stdin from /dev/null: claude "Analyze this pull request for security issues" < /dev/null

D) Add the --batch flag: claude --batch "Analyze this pull request for security issues"


Answer: The -p (or --print) flag

Why: -p is the documented non-interactive mode: it processes the prompt, writes the result to stdout, and exits — which is exactly what a CI job needs (Task 3.6).

Why the others are wrong:

  • References a feature that does not exist: there is no CLAUDE_HEADLESS variable, so the command runs unchanged and the job still waits for input.
  • Is a Unix workaround that does not address the command's own mode of operation.
  • References a feature that does not exist either: there is no --batch flag on the CLI, and an unrecognised option does not switch the command out of interactive mode.
Q11 Which workflows can move to the Batches API

Scenario: Your team wants to reduce API costs for automated analysis. Currently, real-time Claude calls power two workflows: (1) a blocking pre-merge check that must complete before developers can merge, and (2) a technical debt report generated overnight for review the next morning. Your manager proposes switching both to the Message Batches API for its 50% cost savings.

Question: How should you evaluate this proposal?

A) Use batch processing for the technical debt reports only; keep real-time calls for pre-merge checks.

B) Switch both workflows to batch processing with status polling to check for completion.

C) Keep real-time calls for both workflows to avoid batch result ordering issues.

D) Switch both to batch processing with a timeout fallback to real-time if batches take too long.


Answer: Batch for the overnight report only

Why: the Batches API gives 50% savings but processes in up to 24 hours with no latency SLA. That is disqualifying for a blocking pre-merge check and ideal for an overnight report (Task 4.5).

Why the others are wrong:

  • "batches are often faster than the ceiling" is not a guarantee, and a blocking workflow needs a guarantee.
  • Rests on a misconception — batch responses are correlated with custom_id, so ordering is a solved problem.
  • Adds a dual code path and a fallback to solve a problem that disappears by matching each API to its workload.
Q12 Inconsistent review of a 14-file pull request

Scenario: A pull request modifies 14 files across the stock tracking module. Your single-pass review analyzing all files together produces inconsistent results: detailed feedback for some files but superficial comments for others, obvious bugs missed, and contradictory feedback—flagging a pattern as problematic in one file while approving identical code elsewhere in the same PR.

Question: How should you restructure the review?

A) Split into focused passes: analyse each file individually for local issues, then run a separate integration-focused pass examining cross-file data flow.

B) Require developers to split large PRs into smaller submissions of 3–4 files before the automated review runs.

C) Switch to a higher-tier model with a larger context window to give all 14 files adequate attention in one pass.

D) Run three independent review passes on the full PR and only flag issues that appear in at least two of the three runs.


Answer: Per-file passes plus a separate integration pass

Why: the root cause is attention dilution across many files at once. Per-file passes give consistent depth; a dedicated integration pass still catches the cross-file issues (Tasks 1.6 and 4.6).

Why the others are wrong:

  • Moves the burden onto developers without improving the system.
  • Confuses context capacity with attention quality — fitting the files was never the problem.
  • Consensus voting suppresses real bugs, since a genuine issue may be caught in only one run.

3. Scénario 1 — Customer Support Resolution Agent (Q13–Q28)

Contexte du scénario. Agent de résolution support bâti sur le Claude Agent SDK. Il traite des demandes à forte ambiguïté — retours, litiges de facturation, problèmes de compte — via des outils MCP (get_customer, lookup_order, process_refund, escalate_to_human). Objectif : 80 % de résolution au premier contact, tout en sachant quand escalader. Domaines dominants : D1, D2, D5.

Q13 Terminating the agentic loop

Scenario: Your agentic loop occasionally stops before the work is done: the agent calls lookup_order, receives the result, and the loop exits before the refund is processed. Your implementation exits as soon as the response contains any text content.

Question: What is the correct loop control condition?

A) Continue until the model's text contains a closing phrase such as "Is there anything else I can help with?"

B) Continue for a fixed maximum of 10 iterations, then return the last response.

C) Continue while stop_reason is "tool_use"; exit when it is "end_turn".

D) Continue until the assistant produces no text content in its response.


Answer: Drive the loop from stop_reason

Why: stop_reason is the explicit control signal. "tool_use" means the model is waiting on a tool result: execute it, append the tool_result, call again. "end_turn" means the turn is finished (Task 1.1). The bug described is exactly the anti-pattern of treating text content as a completion indicator — a model can emit text and request a tool in the same response.

Why the others are wrong:

  • Parsing natural-language signals for termination is the anti-pattern the guide names most explicitly.
  • An iteration cap is a safety net, never the primary stopping mechanism: it truncates legitimate work and lets broken loops run to the cap.
  • Inverts the reported bug without escaping it: an "end_turn" response almost always carries text, so the loop would never stop, while a "tool_use" response that carries none would end it immediately. The presence of text, like its absence, is never the termination signal.
Q14 What to send on the next iteration

Scenario: After executing lookup_order, your code calls the API again with only the original user message and the tool's output as a plain string appended to it. The agent then re-requests the same tool call.

Question: What is the correct way to return the result to the model?

A) Send only the tool output as a new user message; the model retains the rest of the conversation server-side.

B) Append the assistant's tool_use block and a user message containing a tool_result block with the matching tool_use_id, keeping the full prior history.

C) Send the tool output in the system field so it takes priority over the conversation.

D) Replace the conversation with a summary of what happened so far plus the tool output, to keep the context small.


Answer: Full history, with the tool_use / tool_result pair intact

Why: the API is stateless. Tool results are appended to conversation history as tool_result blocks linked by tool_use_id, so the model can see that the call it requested has been satisfied and reason about the next action (Task 1.1). Losing that link is why the agent re-requests the same call.

Why the others are wrong:

  • There is no server-side conversation state to retain — that is the definition of a stateless API.
  • The system prompt defines behaviour; it is not where turn-level results belong.
  • Summarising mid-loop discards the structured link between request and result, and drops exactly the transactional details you will need later.
Q15 Heterogeneous date formats across MCP tools

Scenario: Your three MCP tools return dates differently: get_customer emits Unix timestamps, lookup_order emits ISO 8601 strings, and the legacy billing tool emits numeric status codes. The agent regularly miscomputes whether an order is still inside the 30-day return window.

Question: What is the most reliable way to fix this?

A) Add few-shot examples showing correct conversions from each of the three formats, so the model has a worked pattern to follow before it judges any return window.

B) Ask the model to call a convert_date tool whenever it meets an unrecognised timestamp, concentrating conversion in one maintained implementation.

C) Implement a PostToolUse hook that normalises the tool results into a single date and status representation before the model sees them.

D) Add a section to the system prompt explaining each tool's date format and asking the model to convert carefully.


Answer: A PostToolUse hook that normalises the data

Why: data normalisation is a deterministic transformation, so it belongs in code that runs on every result rather than in an instruction the model may or may not apply. A PostToolUse hook intercepts results before the model processes them, which is its documented purpose (Task 1.5).

Why the others are wrong:

  • Examples make the conversion probable rather than exact: the model still performs the arithmetic itself, on a problem that has one right answer, and the examples occupy context on every request whether or not a date appears.
  • The logic does end up in one place, but reaching it depends on the model noticing that a conversion is needed, and every date it does notice costs an extra round trip.
  • Prompt text describing three formats leaves the same conversion to the model on every tool result, and pays for the description on every request — an instruction where a transformation is what is required.
Q16 A refund ceiling that must hold

Scenario: Company policy caps agent-issued refunds at $500; anything above must go to a human. The system prompt states the rule, but audit logs show four refunds above the cap were processed last quarter.

Question: What change guarantees the policy holds?

A) A hook that intercepts the outgoing process_refund call, blocks it when the amount exceeds $500, and redirects to the escalation workflow.

B) A post-hoc audit job that flags refunds above $500 the next morning, so finance can reverse each breach before the customer has spent the money.

C) A stronger system prompt that repeats the $500 limit in capital letters and warns about compliance consequences.

D) Few-shot examples showing the agent escalating a $750 refund and a $1,200 refund, so escalation becomes the behaviour it imitates on every request above the cap.


Answer: Intercept the tool call and block it

Why: business rules that must hold every time require deterministic enforcement. A tool-call interception hook evaluates the actual arguments and can refuse the action, then route to escalation — prompt text cannot (Task 1.5).

Why the others are wrong:

  • Catching every breach is not the same as preventing one: the money has already left, reversal depends on the customer co-operating, and the policy is about not issuing the refund in the first place.
  • Louder wording lowers the failure rate without removing it. The rule was already stated in the system prompt, and the breaches recorded last quarter are the evidence that "usually complies" is the wrong standard here.
  • Demonstrations steer the agent towards escalating without ever making the cap impossible to cross; a demonstrated behaviour remains probabilistic, and this policy needs a guarantee.
Q17 Every failure looks the same

Scenario: All four MCP tools return {"error": "Operation failed"} with isError: true on any failure — a database timeout, an invalid order number, a refund refused by policy, and a permissions problem all look identical. The agent retries everything three times, then apologises.

Question: What should the tools return instead?

A) Structured error metadata including an errorCategory (transient / validation / business / permission), an isRetryable boolean, and a human-readable description.

B) A success response with an empty result set, so the agent moves on without retrying.

C) The raw stack trace from the backend service, so the agent has the maximum information available to it.

D) A numeric error code the agent can resolve against a recovery table carried in the system prompt.


Answer: Structured error metadata with a category and a retryable flag

Why: the agent's recovery decision depends on which kind of failure occurred. A category plus an explicit isRetryable flag lets it retry a timeout, correct an invalid input, explain a policy refusal to the customer, and escalate a permissions problem — four different behaviours from one uniform signal today (Task 2.2).

Why the others are wrong:

  • Is silent suppression — reporting failure as success, which prevents any recovery at all.
  • Dumps implementation detail into the context without answering the only question that matters: is this worth retrying?
  • Pushes the semantics into a lookup table the model must apply correctly, when the tool can simply return them.
Q18 A refund the policy forbids

Scenario: A customer requests a refund on an item bought 60 days ago; the return window is 30 days. Your process_refund tool currently returns {"isError": true, "error": "Refund rejected"}, and the agent retries the call twice before telling the customer that "the system is unavailable".

Question: How should the tool report this?

A) A successful response with refund_amount: 0, letting the agent infer that the refund was declined.

B) The same error, plus instructions in the system prompt telling the agent never to retry refund errors.

C) A transient error, so the agent retries and eventually escalates to a human who can override the window.

D) A structured business error marked retriable: false, with a customer-friendly explanation.


Answer: A business error, explicitly non-retryable, with a usable explanation

Why: a policy refusal is a definitive outcome, not a fault. Marking it non-retryable stops the wasted attempts, and the explanation carried with the error — the item fell outside the 30-day window — gives the agent something accurate to say, and an alternative to offer, instead of inventing an outage (Task 2.2).

Why the others are wrong:

  • Makes a refusal indistinguishable from a zero-value success, which is precisely the ambiguity the isError flag exists to remove.
  • Leaves the tool's contract broken and patches it with a prompt rule that applies to all refund errors, including the genuinely transient ones.
  • Mislabels a deterministic refusal as a fault — retrying will fail identically every time, and the escalation arrives only after the customer has watched the agent flounder.
Q19 A number lost twenty turns back

Scenario: In a long billing dispute, the customer agreed early on to a partial refund of $89.99 on order ORD-67890. Thirty turns later, after the history has been summarised twice, the agent offers "a partial refund" without an amount and cannot recall the order number.

Question: What is the correct fix?

A) Raise the summarisation threshold so history is compressed less often and the agreed terms survive deeper into a dispute that keeps growing.

B) Instruct the summariser to preserve every identifier and amount verbatim.

C) Maintain a case facts block — customer ID, order ID, amount, status — re-injected into every prompt.

D) Move to a model with a larger context window, so summarisation never triggers and the agreed terms stay in the transcript.


Answer: A case facts block outside the summarised history

Why: progressive summarisation compresses by nature, and amounts, dates and stated expectations are exactly what it dilutes. Extracting those facts into a persistent block that is re-injected every turn puts them out of reach of compression (Task 5.1).

Why the others are wrong:

  • Postpones the compression instead of protecting the facts: a longer dispute reaches the raised threshold and dilutes the same amount and order number.
  • Relies on the summariser complying every time, on every fact — the same probabilistic guarantee that already failed twice here.
  • Also postpones rather than protects, and buys the delay at a price: a larger window leaves the agreed amount buried far back in a long transcript instead of restating it each turn.
Q20 Forty fields per order lookup

Scenario: lookup_order returns 40+ fields per order — warehouse routing codes, carrier metadata, internal audit flags — of which the agent needs five for return eligibility. Sessions handling three or four orders run out of context before the case is resolved.

Question: What is the most effective remedy?

A) Split lookup_order into 40 single-field tools so the agent asks for exactly the fields a return decision needs.

B) Instruct the agent in the system prompt to ignore the routing and audit fields a return decision never uses.

C) Trim the tool output to the return-relevant fields before it enters the model's context.

D) Run /compact as soon as the context passes 70% utilisation.


Answer: Trim verbose tool output before it accumulates

Why: tool results consume context in proportion to their size, not their relevance. Filtering at the boundary means the 35 irrelevant fields never cost anything, in this turn or any later one (Task 5.1).

Why the others are wrong:

  • Explodes the tool count, which degrades selection reliability, and multiplies round trips.
  • Telling the model to ignore text does not remove the text — the tokens are already spent.
  • Compresses after the waste has been paid for, and compression is itself lossy.
Q21 "Just give me a manager"

Scenario: A customer's second message reads: "I don't want to go through this again with a bot. Put me through to a manager." The underlying issue is a straightforward damaged-item replacement that the agent could resolve in two tool calls.

Question: What should the agent do?

A) Resolve the replacement first, since it is within its capability, and mention that a human is available if the customer is still unhappy.

B) Score the request for frustration and escalate only if the score exceeds the configured threshold.

C) Explain that it can handle this case faster than a human and ask the customer to confirm they still want to escalate.

D) Call escalate_to_human immediately, with a structured handoff summary.


Answer: Escalate immediately

Why: an explicit request for a human is one of the three reliable escalation triggers, and it is honoured without first attempting to resolve. Capability is not the criterion here — the customer's stated preference is (Task 5.2).

Why the others are wrong:

  • Overrides an explicit request by acting first and offering the human afterwards. It answers the wrong question: whether the agent can resolve the case, when what was asked was who should handle it.
  • Substitutes a sentiment proxy for an unambiguous instruction that is already in the transcript.
  • Overrides the request too, and is the subtler trap: asking the customer to re-justify something they already stated is a negotiation, not a resolution.
Q22 "This is outrageous!"

Scenario: A customer writes: "This is outrageous, I've been a customer for six years and I'm very unhappy!" They have not asked for a human. Their issue is a duplicate charge that the agent can reverse.

Question: What is the correct behaviour?

A) Acknowledge the frustration, offer the resolution, and escalate only if the customer then asks for a human.

B) Ask the customer to rate their frustration from 1 to 10 to decide whether to escalate.

C) Escalate immediately — strong negative sentiment predicts an unsatisfactory outcome.

D) Ignore the emotional content and process the reversal at once, since the duplicate charge is what the customer actually wants undone.


Answer: Acknowledge, resolve, escalate only on reiteration

Why: frustration is not a request for a human, and it does not correlate with case complexity. The nuanced pattern respects the customer's experience while still meeting the first-contact resolution target (Task 5.2).

Why the others are wrong:

  • Turns an emotional signal into a numeric one without making it any more predictive.
  • Is sentiment-based escalation — the guide's canonical unreliable trigger, and it is what drove first-contact resolution down in the first place.
  • Reverses the charge the customer wants undone and says nothing to the person who has just told you they are unhappy; acknowledgement costs one sentence and changes how the same resolution lands.
Q23 Three customers, one name

Scenario: The customer gives their name as "J. Martin". get_customer returns three matching records with different addresses and order histories. Two of the three have an open return.

Question: What should the agent do?

A) Ask the customer for one more identifier before acting.

B) Escalate to a human, because ambiguous identity is beyond the agent's remit.

C) Select the record with the most recent activity, as it is the most likely match.

D) Select the record with an open return, since that matches the customer's stated problem.


Answer: Request an additional identifier

Why: multiple matches call for clarification, not heuristic selection. An order number, an email address or a postcode settles the ambiguity in one turn, and the customer can supply any of them. Guessing wrong here means acting on the wrong person's account — the exact failure mode that identity verification exists to prevent (Task 5.2).

Why the others are wrong:

  • Escalates a case the agent can resolve with a single clarifying question, which is what drives first-contact resolution down.
  • "Most recent activity" is a heuristic dressed as reasoning: plausible, unverified, and just as capable of selecting the wrong account as any other tie-break.
  • "Has an open return" is a heuristic too, and the scenario shows why it fails — two of the three records carry one, so the criterion does not even single out a record.
Q24 The policy is silent

Scenario: A customer asks you to match a competitor's lower price on an item they bought last week. Your price-adjustment policy covers price drops on your own site and says nothing about competitors' prices — neither allowing nor forbidding a match.

Question: What should the agent do?

A) Apply the own-site adjustment rule by analogy and grant the match, since the customer is asking for the same discount either way.

B) Escalate: the policy is silent on this request, which is a policy gap.

C) Decline, since anything the policy does not explicitly permit is forbidden.

D) Ask the customer to provide the competitor's URL and then decide autonomously.


Answer: Escalate — this is a policy gap

Why: policy exceptions and gaps are a named escalation trigger. The distinction that matters is not "is this case hard?" but "does policy actually cover it?" — here it does not, so the decision is not the agent's to make (Task 5.2).

Why the others are wrong:

  • Invents a permissive policy by analogy: the own-site rule was written for one situation, and stretching it to competitors' prices is a commercial decision the agent has not been granted.
  • Invents a restrictive policy in the opposite direction. Silence is not a refusal any more than it is permission, and turning the customer away on that basis is still the agent deciding a case it was never given.
  • Gathers more evidence about a question the agent has no mandate to answer.
Q25 Handing off mid-case

Scenario: After verifying the customer, confirming the order, and offering a replacement that was refused, the agent escalates. Human agents receive only a notification saying "Customer requests human assistance — refund issue" and do not have access to the conversation transcript. They routinely re-ask questions the customer already answered.

Question: What should the handoff include?

A) The customer ID and a link to the account, letting the human open the record and work the case out for themselves.

B) A structured handoff summary: identifiers, root cause, actions completed, recommendation.

C) The agent's confidence score for the case, so the human knows how much of the work to re-verify.

D) The full raw transcript, so the human loses nothing and can judge the case from the customer's own wording.


Answer: A structured, self-sufficient handoff summary

Why: the receiving human has no transcript, so the summary is the context. It must carry the facts — customer ID, order ID, the refund amount discussed — the diagnosis, the actions already taken, the reason for escalating and a recommended next step: enough to act without re-interviewing the customer (Task 1.4).

Why the others are wrong:

  • Discards everything the agent established and guarantees the duplicate questioning described.
  • A confidence number tells the human nothing about what happened or what to do.
  • A raw transcript transfers the reading burden instead of the conclusion; the human still has to reconstruct the root cause.
Q26 Three problems in one message

Scenario: A customer writes a single message containing three issues: a damaged item on order A, a duplicate charge on order B, and a request to change the delivery address on order C. The agent addresses the damaged item, then replies as though the conversation were finished.

Question: What is the right handling pattern?

A) Escalate the whole conversation, since a message carrying three separate issues exceeds what a single agent should be resolving on its own.

B) Handle the first issue, then ask the customer to resend the remaining two separately, so each one gets a clean thread.

C) Decompose the message into three items, investigate each in parallel where tools allow using shared customer context, then synthesise one response covering all three.

D) Address whichever issue has the highest monetary value first and note the other two for follow-up on the next contact.


Answer: Decompose, investigate each, synthesise one answer

Why: multi-concern requests are decomposed into distinct items that share the already-established customer context, and resolved together. Splitting the investigation while unifying the response is what keeps first-contact resolution high (Task 1.4).

Why the others are wrong:

  • The number of issues does not make any of them complex; nothing here matches an escalation trigger.
  • Pushes work back to the customer and turns one contact into three.
  • Silently drops two legitimate requests.
Q27 An instruction that pulls a tool

Scenario: Your system prompt opens with "Always verify the customer's identity." You observe get_customer being called on every turn, including when the customer only asks about store opening hours. The tool descriptions are detailed and well differentiated.

Question: What is the root cause?

A) The agent has too many tools and falls back to the first one in the list whenever it hesitates.

B) The system prompt wording creates an unintended association with get_customer; the fix is to state when verification is actually required.

C) get_customer's description is too broad, so the model reaches for it on requests that never involve an account and pays for a lookup nobody needed.

D) tool_choice is set to "any", forcing a tool call on every turn even when the question needs none.


Answer: Keyword-sensitive system prompt wording

Why: the word "always" reads as an unconditional trigger, and the model attaches it to the semantically closest tool. Keyword-sensitive instructions can override well-written tool descriptions — which is why the scenario is careful to tell you the descriptions are good. The repair is to say when verification applies rather than "always" (Task 2.1).

Why the others are wrong:

  • There is no positional fallback to hesitate into: selection is driven by the tool descriptions and the system prompt, never by a tool's rank in the list.
  • Contradicts the stated facts: the descriptions are detailed and well differentiated, so breadth is not what drags the tool onto turns about opening hours.
  • Would indeed force a call on every turn, but nothing in the scenario points to a forced tool_choice, and the default is "auto".
Q28 Guaranteeing a tool call

Scenario: You are building a triage step that must always produce a structured classification of the incoming message. You have three classification tools, one per request family, and the request type is not known in advance. Today the model sometimes answers in prose instead of calling any of them.

Question: Which configuration fits?

A) tool_choice: {"type": "auto"}

B) tool_choice: {"type": "tool", "name": "classify_billing"}

C) tool_choice: {"type": "none"} with a prompt instructing the model to output JSON

D) tool_choice: {"type": "any"}


Answer: {"type": "any"}

Why: "any" guarantees a tool call while leaving the choice of tool to the model — exactly right when several schemas coexist and the input type is unknown (Task 2.3).

Why the others are wrong:

  • "auto" is the current behaviour: the model may answer in text, which is the reported problem.
  • Forcing one named tool works when you know which one should run; here it would misclassify every non-billing request.
  • A prompt asking for JSON is guidance, not a guarantee — the reported failure over again. "none" is API vocabulary the guide never lists: its enumeration stops at "auto", "any" and forced selection. The API reference defines it as preventing the model from using any tool, so this option forbids the very call the step depends on.

4. Scénario 2 — Code Generation with Claude Code (Q29–Q42)

Contexte du scénario. Votre équipe utilise Claude Code pour la génération, le refactoring, le débogage et la documentation. Il faut l'intégrer au workflow de développement : slash commands personnalisées, configurations CLAUDE.md, et savoir quand utiliser le plan mode plutôt que l'exécution directe. Domaines dominants : D3, D5.

Q29 The new hire gets different behaviour

Scenario: Your team's conventions — error handling style, test naming, import ordering — are respected on every existing developer's machine. A developer who joined last week reports that Claude Code ignores all of them. The conventions were written by the tech lead in ~/.claude/CLAUDE.md.

Question: What is the diagnosis and fix?

A) The new developer's Claude Code version is older and does not yet read the convention file; ask them to upgrade.

B) The file is too long and is silently truncated; split it.

C) The new developer must run /init to generate a configuration from the codebase, which will infer the conventions from existing code.

D) The conventions live at user level; move them to a project-level CLAUDE.md in the repository.


Answer: User-level configuration is personal by definition

Why: ~/.claude/CLAUDE.md applies to one user on one machine and is never distributed with the repository. Anything the whole team must follow belongs at project level, in version control (Task 3.1). Note the tell in the scenario: it works for everyone who was there when the file was written, and only for them.

Why the others are wrong:

  • No release of the tool can read a file that is not on the machine, and a version difference would not split cleanly between the long-standing developers and the one who joined last week.
  • Truncation would degrade everyone equally rather than one person, and here nothing is being truncated: the file is not on that machine at all.
  • /init writes a fresh local file from whatever the codebase happens to show, which is not the same as the conventions the tech lead wrote down, and nothing it produces reaches the rest of the team.
Q30 One monorepo, four sets of standards

Scenario: Your monorepo holds four packages, each with its own standards documents already maintained by their owners (docs/api-standards.md, docs/ui-standards.md, and so on). You want each package's CLAUDE.md to carry the relevant standards without copying them, so that updates to the source documents propagate.

Question: Which mechanism fits?

A) Use @import in each package's CLAUDE.md to reference the relevant standards files.

B) Concatenate all four standards documents into the root CLAUDE.md so every session sees them.

C) Create a skill per package whose SKILL.md carries that package's standards, loaded only when invoked.

D) Symlink each standards file into the package directory as CLAUDE.md, so source edits propagate automatically.


Answer: @import the relevant standards files

Why: @import exists to keep CLAUDE.md modular by referencing external files rather than duplicating them, and it lets each package's maintainer include only what is relevant to their domain (Task 3.1).

Why the others are wrong:

  • Every session does see them, which is the cost: work touching one package pays for the other three, and the owners lose the file they maintain.
  • Loading only on invocation is the wrong shape here — standards that govern every line written in a package must apply without anyone remembering to ask for them.
  • Edits would propagate, but the symlink consumes the whole file: there is no room for package-specific content beside it, and it breaks on platforms without symlink support.
Q31 Inconsistent behaviour between sessions

Scenario: Claude Code applies your testing conventions in some sessions and not others, on the same machine and the same repository. You suspect a configuration file is not always being picked up.

Question: What is the fastest way to confirm?

A) Add a sentinel instruction ("always start your reply with OK") and check whether it appears.

B) Compare the output of two sessions and infer which rules were active.

C) Run /memory to see which memory files are actually loaded in the current session.

D) Delete all configuration files and recreate them one by one.


Answer: /memory

Why: /memory reports which memory files are loaded, which turns a guess about configuration hierarchy into an observation (Task 3.1).

Why the others are wrong:

  • A workable hack, but it only tells you whether one file loaded, not which ones did.
  • Inference from output is exactly the unreliable method that produced the confusion.
  • Destructive and slow, for information a single command already provides.
Q32 A 900-line CLAUDE.md

Scenario: Your root CLAUDE.md has grown to 900 lines covering testing, API conventions, deployment, database migrations and code review. Developers report that Claude follows the sections near the top reliably and the rest inconsistently, and every session pays for the whole file.

Question: What is the most maintainable restructuring?

A) Cut the file down to the twenty rules that developers break most often, so all that is left is read.

B) Reorder the file so the most important sections are at the top and the least important at the bottom.

C) Move the whole content into a single skill, so a session pays for the standards only on the tasks that actually need them.

D) Split it into topic-specific files under .claude/rules/, path-scoped where a topic maps to a file pattern.


Answer: Split into topic-specific rule files

Why: .claude/rules/ exists as the alternative to a monolithic CLAUDE.md. One file per topic — testing, API conventions, deployment, migrations, review — is separately maintainable, and a topic that maps to a file pattern can be path-scoped so it loads only when relevant (Tasks 3.1 and 3.3).

Why the others are wrong:

  • What survives would indeed be read reliably, but the method is to delete real conventions to fit a budget — a loss, when scoping keeps all of them and still loads only what is relevant.
  • Reordering manages the symptom — position effects — while leaving 900 lines loaded every time.
  • The saving is real and so is the price: a skill only runs when someone invokes it, and universal standards must apply without anyone remembering to ask.
Q33 A skill that floods the conversation

Scenario: Your /analyze-architecture skill walks the dependency graph and prints hundreds of lines of intermediate findings. Developers like the final summary, but after running it the main session is saturated and subsequent answers degrade.

Question: Which frontmatter option addresses this?

A) Wrapping the skill body in an instruction to "be concise and summarise aggressively", so it prints less.

B) argument-hint, prompting the developer to narrow the scope so the walk produces less output.

C) context: fork, so the skill runs in an isolated subagent context and returns only its summary.

D) allowed-tools, restricting the skill to read-only tools so it gathers less and prints less.


Answer: context: fork

Why: context: fork runs the skill in an isolated sub-agent context precisely so verbose output does not pollute the main session — the documented use case being exactly this kind of codebase analysis (Task 3.2).

Why the others are wrong:

  • A prompt instruction gives no guarantee, and the intermediate findings are useful — they just belong elsewhere.
  • A narrower walk does produce less, but the remainder still lands in the main session, and the judgement of how far to narrow is pushed onto the developer every time.
  • Read-only tools still read, and reading is what generates the volume here: allowed-tools governs what the skill may do, never where its output lands.
Q34 A skill that must not delete anything

Scenario: You are writing a /scaffold-component skill that creates new files from templates. During testing it occasionally runs shell commands and once removed a directory it had just created. You want the skill to be structurally incapable of destructive actions.

Question: What do you configure?

A) context: fork, so any damage stays in the isolated context.

B) A PostToolUse hook that reverts file deletions after they occur.

C) An instruction at the top of SKILL.md forbidding the use of Bash.

D) allowed-tools in the skill's frontmatter, limited to the file-writing operations the skill actually needs.


Answer: allowed-tools in the frontmatter

Why: allowed-tools restricts tool access during skill execution. Limiting the skill to the write operations it needs removes the capability rather than discouraging its use (Task 3.2).

Why the others are wrong:

  • Forking isolates context, not filesystem effects — a subagent writes to the same disk.
  • Reverting after the fact is not prevention, and some deletions are not recoverable.
  • A prose prohibition is probabilistic, and this is a destructive action.
Q35 A skill invoked with no arguments

Scenario: Your /new-endpoint skill needs a resource name and an HTTP method. Developers frequently type /new-endpoint with nothing after it, and the skill either guesses or produces a generic stub.

Question: Which frontmatter field addresses this?

A) context: fork, so a wrong guess is thrown away with the isolated context instead of persisting into the developer's branch.

B) argument-hint, which asks the developer for the values the skill needs before it runs.

C) description, rewritten to spell out the expected arguments before the developer types them.

D) allowed-tools, restricting the skill until arguments are supplied.


Answer: argument-hint

Why: argument-hint exists to surface the required parameters at invocation time (Task 3.2). It converts a silent guess into an explicit prompt.

Why the others are wrong:

  • Addresses context isolation: discarding a bad guess is not the same as obtaining the resource name and method, so the developer still gets a generic stub, merely a tidier one.
  • The description is read by the model when it decides whether to use the skill; nobody shows it to the developer as they type, so it cannot ask for what is missing.
  • Addresses permissions: allowed-tools restricts which tools the skill may call, and it has no way to condition execution on the presence of arguments.
Q36 Skill or CLAUDE.md?

Scenario: You need to encode two things: (1) the team's naming and error-handling standards, which apply to every piece of code anyone writes, and (2) a twelve-step release-preparation procedure that runs roughly once a month.

Question: Where does each belong?

A) Both in CLAUDE.md, so no one has to remember to invoke a skill when a release is due.

B) Both as skills, to keep CLAUDE.md small and every session cheap.

C) Standards in CLAUDE.md (always loaded, universal), release procedure as a skill (invoked on demand).

D) Standards as a skill, release procedure in CLAUDE.md, so the twelve steps are in context on the day they matter most.


Answer: Standards in CLAUDE.md, procedure as a skill

Why: the division is between always-loaded universal standards and on-demand task-specific workflows (Task 3.2). Applying the criterion — "does this apply to every session, or to one kind of task?" — separates the two cases cleanly.

Why the others are wrong:

  • The procedure is then loaded on release day and on every other day too: a twelve-step monthly workflow paid for in every session, for the sake of one afternoon a month.
  • The base context does get cheaper, at the price of making universal standards depend on invocation, so ordinary code generation silently ignores them.
  • Inverts both criteria at once: the release procedure sits in context on release day and on every other day too, while the standards that govern every line wait for someone to invoke them.
Q37 Customising a shared skill for yourself

Scenario: The team's /commit skill in .claude/skills/ enforces a commit message format you find too verbose for your own experimental branches. You want your own variant without changing your teammates' behaviour.

Question: What do you do?

A) Add a conditional to the shared skill that checks the branch name, so experimental work gets the shorter format for everyone.

B) Create a personal variant in ~/.claude/skills/ under a different name.

C) Add your commit preferences to ~/.claude/CLAUDE.md.

D) Edit the team's SKILL.md and avoid committing the change.


Answer: A personal variant, in the user scope, under a different name

Why: personal customisation lives in ~/.claude/skills/, and using a distinct name avoids shadowing or colliding with the team's skill (Task 3.2).

Why the others are wrong:

  • Encodes one person's preference into shared team tooling: every teammate on an experimental branch now inherits a format nobody asked them about, and the shared skill grows a rule that only makes sense to you.
  • CLAUDE.md carries instructions, not skill definitions, and cannot replace a skill's behaviour.
  • An uncommitted edit to a tracked file is a permanent source of accidental commits and merge noise.
Q38 A stack trace and a one-line fix

Scenario: A production stack trace points to a null dereference in formatInvoiceDate() when the issued_at column is null. The fix is a guard clause in one function, in one file.

Question: Plan mode or direct execution?

A) Plan mode — every production fix warrants a design review first.

B) Direct execution, but only after an Explore subagent maps the call sites.

C) Plan mode, then direct execution, to document the reasoning for the postmortem.

D) Direct execution — the scope and the fix are clear.


Answer: Direct execution

Why: plan mode earns its cost on large-scale changes, competing approaches and architectural decisions. A single-file fix whose scope is bounded by the stack trace and whose change is already well understood is the guide's own example of appropriate direct execution (Task 3.4).

Why the others are wrong:

  • Applies a heavy process by rule rather than by reading the situation. "Every production fix warrants a design review" is a policy claim, and a guard clause in one function contains no design decision for plan mode to weigh.
  • Exploration is for unclear scope; here the stack trace has already localised the defect.
  • Applies the same heavy process for a different reason — documentation — and pays plan mode's cost on a change with no competing approaches to record.
Q39 Discovery that eats the session

Scenario: You are migrating a library across 45 files. The discovery phase — locating every call site, reading each one, checking which use the deprecated signature — produces a large volume of output, and by the time you reach implementation the session is degraded.

Question: How should you structure this?

A) Run /compact after each file so the context never grows past what the next file needs.

B) Run the whole migration in direct execution, file by file, restarting the session whenever the answers start to degrade.

C) Use plan mode for the investigation, delegating discovery to the Explore subagent, then execute directly.

D) Do the discovery outside Claude Code and paste the list of affected files in.


Answer: Plan mode plus Explore for discovery, then direct execution

Why: this combines the two mechanisms the guide pairs for exactly this shape of task — plan mode for a change with architectural implications across 45 files, and the Explore subagent to isolate verbose discovery output and return summaries (Task 3.4).

Why the others are wrong:

  • Compaction is lossy and, applied after every file, will erode the migration plan itself.
  • Each restart does clear the degradation, and also the accumulated understanding, so the discovery cost is paid again from scratch every time.
  • Keeps the session clean by moving the expensive part onto you, and discards the tooling that makes discovery reliable and repeatable — a hand-built list is exactly the thing nobody will redo when the migration is revisited.
Q40 A transformation described three ways

Scenario: You ask Claude to "normalise the legacy customer records into the new schema". Each run interprets "normalise" differently: one flattens nested addresses, another keeps them; one uppercases country codes, another does not. Your prose description has been rewritten twice without converging.

Question: What is the most effective next step?

A) Provide two or three concrete input/output example pairs showing exactly what a normalised record looks like.

B) Lower the temperature so every run settles on the same interpretation of the instruction.

C) Rewrite the instruction once more, longer and more precise than the two attempts before it.

D) Ask the model to explain its interpretation before each run, and correct it each time.


Answer: Concrete input/output examples

Why: when prose descriptions are interpreted inconsistently, worked examples are the most effective way to communicate the expected transformation — they pin down the cases the prose left open instead of describing them again (Task 3.5).

Why the others are wrong:

  • Temperature is out of scope for this exam, and determinism would only make one arbitrary interpretation repeatable.
  • Is the approach that has already failed twice; more words about an ambiguous transformation do not remove the ambiguity.
  • Turns every run into a negotiation instead of fixing the specification.
Q41 Building in an unfamiliar domain

Scenario: You must add a caching layer to a service you have never worked on. You know roughly what you want but suspect there are considerations — invalidation, stampede protection, failure behaviour — that you have not thought through, and you would rather surface them before code exists.

Question: Which technique fits?

A) Ask for three complete implementations and compare them, letting the differences between them reveal the decisions that matter.

B) The interview pattern: have Claude ask you questions to surface considerations you have not anticipated, before implementing.

C) Ask for a list of caching best practices and apply the ones that seem relevant.

D) Write the implementation, then request a review pass that names everything the design overlooked.


Answer: The interview pattern

Why: having the model ask questions first surfaces design considerations the developer had not anticipated — cache invalidation and failure modes being the guide's own examples — while changing the design is still cheap (Task 3.5).

Why the others are wrong:

  • The differences show you where three answers diverge, not why: each embodies its own unstated assumptions, and comparing finished code never puts those assumptions into words you can weigh.
  • A generic checklist is not calibrated to this service, and you are the one who must judge relevance — which is what you said you cannot yet do.
  • Review after the fact is useful, but it finds problems once the design has already been committed to code.
Q42 Five issues in one review

Scenario: A generated module has five problems. Three are independent — a typo in a log message, a missing null check, an unused import. Two interact: the retry logic and the timeout configuration are inconsistent with each other, and fixing either alone changes what the correct fix for the other is.

Question: How should you feed these back?

A) Send all five one at a time, so each fix can be verified in isolation before the next.

B) Send all five in one message, so the model sees the whole picture and round trips stay low.

C) Send the two interacting issues one at a time, since they are the most delicate of the five.

D) Send the two interacting issues together in a single detailed message, and iterate on the three independent ones separately.


Answer: Interacting issues together, independent issues separately

Why: the criterion is interaction, not count or difficulty. Problems whose fixes affect one another must be seen together, or the model optimises one against a moving target; independent problems are cleaner to handle one at a time (Task 3.5).

Why the others are wrong:

  • Applies sequential iteration to all five and so splits the interacting pair, which is exactly the case that cannot be fixed one side at a time.
  • Applies one message to all five and so merges the three independent problems, which are cleaner to verify separately — round-trip count is not the criterion the scenario is built around.
  • Ranks the five by delicacy when the criterion is interaction, and so inverts the rule exactly: the pair is the one case that must not be split.

5. Scénario 3 — Multi-Agent Research System (Q43–Q58)

Contexte du scénario. Système de recherche multi-agents sur le Claude Agent SDK. Un coordinateur délègue à des subagents spécialisés : recherche web, analyse de documents, synthèse, génération de rapport. Le système produit des rapports complets et sourcés. Domaines dominants : D1, D2, D5.

Q43 The coordinator cannot spawn anything

Scenario: Your coordinator's prompt describes the four subagents and when to use each. At runtime it never delegates: it attempts the research itself with its own web tools and produces a shallow report. Its allowedTools is ["WebSearch", "Read", "Write"].

Question: What is the cause?

A) Subagents must be spawned by the SDK entry point, not by an agent at runtime.

B) The subagent definitions are missing description fields, so the coordinator cannot tell them apart.

C) The coordinator's system prompt describes delegation but does not command it imperatively enough.

D) allowedTools does not include "Task", which is the mechanism for spawning subagents.


Answer: "Task" is missing from allowedTools

Why: the Task tool is how a coordinator spawns subagents, and it must appear in allowedTools for the coordinator to invoke them. Without it, no amount of prompt instruction can produce a delegation (Task 1.3). The symptom is characteristic — the agent falls back to doing the work itself with the tools it does have.

Why the others are wrong:

  • Subagents are spawned at runtime by the coordinator; that is the point of the pattern.
  • Missing descriptions would produce poor selection among subagents, not the complete absence of delegation.
  • Prompt phrasing cannot grant a capability that is not configured.
Q44 The synthesis agent knows nothing

Scenario: The web search agent and the document analysis agent both complete successfully. The coordinator then spawns the synthesis agent with the prompt "Synthesise the findings into a report." The synthesis agent produces a generic essay about the topic with no reference to anything the other two found.

Question: What is the fix?

A) Run all three subagents in a single shared session so context is common.

B) Increase the synthesis subagent's context window so it can hold the prior conversation.

C) Have the synthesis subagent call the other two subagents to ask what they found.

D) Pass the prior agents' findings in the synthesis prompt.


Answer: Pass the findings explicitly in the prompt

Why: subagents operate with isolated context. They do not inherit the coordinator's conversation history and do not share memory between invocations, so anything they need must be provided explicitly, and in full: the synthesis prompt has to carry the complete findings, not a pointer to them (Task 1.3).

Why the others are wrong:

  • Collapses the isolation that makes delegation useful in the first place.
  • Context size is irrelevant when nothing was passed — an empty context does not fill itself.
  • Breaks hub-and-spoke by creating direct subagent-to-subagent communication, and the coordinator already holds the results.
Q45 Four searches, one after another

Scenario: The coordinator needs four independent web searches on four distinct subtopics. It currently emits one Task call, waits for the result, emits the next, and so on. End-to-end latency is four times a single search.

Question: How do you parallelise?

A) Set disable_parallel_tool_use: false and re-run the existing sequence.

B) Emit the four Task calls across four consecutive turns without waiting for results.

C) Emit all four Task tool calls in a single coordinator response.

D) Create one subagent that internally runs the four searches in parallel.


Answer: Multiple Task calls in one response

Why: parallel subagents are spawned by emitting several Task tool calls within a single coordinator response, rather than across separate turns (Task 1.3). The four then run concurrently and their results come back together.

Why the others are wrong:

  • Removing a parallelism restriction does not restructure a loop that emits one call at a time.
  • Each turn still requires the previous response to complete — this is the sequential pattern with extra steps.
  • Hides the fan-out inside one agent, losing the per-subtopic scoping and the coordinator's visibility into each branch.
Q46 Citations that dissolve

Scenario: The web search agent returns its findings as flowing prose: "According to a recent industry report, adoption grew sharply, and a separate academic study found similar effects." By the time the synthesis agent has combined three such reports, no claim can be traced to a source.

Question: How should subagent output be designed?

A) Use a structured format that separates each claim from its source metadata, which downstream agents must preserve rather than compress.

B) Store the raw search results in a file and have the reader consult it if a claim needs checking.

C) Have the web search agent name each source inline in its prose, so attribution travels with the sentence.

D) Ask the synthesis agent to add citations at the end, matching claims to the sources the coordinator recorded.


Answer: Structured output separating content from metadata

Why: attribution survives summarisation only when it is carried as structure rather than prose. The metadata that travels with each claim — the evidence excerpt, the source name, the publication date — gives downstream agents something they can merge without inventing or losing the mapping (Tasks 1.3 and 5.6).

Why the others are wrong:

  • Shifts verification onto the reader, which is what a cited report is supposed to avoid.
  • Inline naming reads well until the next summarisation step rewrites the sentence, and the source name goes with it.
  • The coordinator's list says which sources were consulted, not which claim came from which — so the mapping gets reconstructed by guesswork.
Q47 A coordinator prompt that scripts every step

Scenario: Your coordinator prompt reads: "Step 1: call the web search agent with the exact topic string. Step 2: pass the first five results to the document analysis agent. Step 3: pass its output to synthesis. Step 4: call the report agent." On topics where the first five results are weak, the system still marches through all four steps and produces a poor report.

Question: How should the coordinator prompt be written instead?

A) Add a fifth step that restarts the pipeline whenever the finished report looks poor.

B) Increase the number of results passed from five to twenty-five, so the analysis agent has far more material to sift through before it commits.

C) Add explicit conditional branches to the prompt for every case where the results may come back weak, so each one is handled.

D) Specify the research goal and the quality criteria the output must meet, then let the coordinator choose its path.


Answer: Goals and quality criteria rather than procedural steps

Why: coordinator prompts that state research goals and quality criteria enable subagent adaptability, where step-by-step procedural instructions force the same path regardless of what is discovered. Stated that way, the coordinator decides which subagents to invoke and when coverage is good enough to stop (Task 1.3).

Why the others are wrong:

  • Restarting the whole pipeline is far more expensive than adapting mid-course, and it repeats the work that was fine.
  • More weak results do not become strong ones, and a bigger pile to sift through costs analysis budget without changing the fixed path that caused the poor report.
  • Enumerating branches for every anticipated case is precisely the pre-configured decision tree that model-driven decision-making replaces.
Q48 Subagents talking to each other

Scenario: To reduce latency, a developer wires the document analysis agent to call the web search agent directly when it needs an extra source, bypassing the coordinator. Failures in that path are now invisible in the coordinator's logs, and two agents have started requesting overlapping searches.

Question: What principle does this violate, and what is the fix?

A) Hub-and-spoke: all inter-subagent communication, error handling and information routing belong to the coordinator rather than to direct agent calls.

B) Least privilege: the document analysis agent should not have been given a search tool at all.

C) Nothing — direct communication is a valid optimisation as long as the results are correct.

D) Statelessness: the two agents should share a session so each can see what the other has already requested and stop paying twice for the same search results.


Answer: Hub-and-spoke, with the coordinator as the single routing point

Why: routing all subagent communication through the coordinator is what delivers observability, consistent error handling and controlled information flow. Both reported symptoms — invisible failures and duplicated work — follow directly from bypassing it (Task 1.2).

Why the others are wrong:

  • Overstates the rule: a narrowly scoped cross-role tool for a high-frequency need is legitimate. What is not legitimate is an unmediated channel between agents.
  • Correct results say nothing about whether a failure can be seen or duplicate work caught, which are the reasons the pattern exists.
  • A shared session would indeed stop the second agent paying for a search the first already ran, but it removes the isolation that makes subagents useful and still leaves the coordinator blind to failures.
Q49 A four-stage pipeline for a one-line question

Scenario: Every query runs through all four subagents. For "what is the current EU AI Act enforcement date?", the system performs a web search, a full document analysis pass, a synthesis pass and a report generation pass, taking ninety seconds to return a date.

Question: How should the coordinator be designed?

A) Add a rule that queries under ten words skip the document analysis stage, since short questions rarely need it.

B) Select which subagents to invoke from the query's actual requirements.

C) Add a cache so repeated factual queries return instantly instead of running all four stages again.

D) Keep the pipeline but run all four subagents in parallel to cut the ninety seconds down to the slowest stage.


Answer: Dynamic selection based on query requirements

Why: a coordinator's job includes deciding which subagents to invoke based on query complexity, not just orchestrating a fixed sequence (Task 1.2). The full pipeline stays available for the queries that warrant it; what changes is that a one-line factual lookup stops paying for four stages.

Why the others are wrong:

  • Query length is a poor proxy for complexity; a short question can be genuinely hard.
  • Caching helps on repeats and does nothing for the first occurrence of every simple question.
  • The stages are dependent — synthesis needs analysis output — so they cannot simply be parallelised, and the unnecessary work still gets done.
Q50 Three agents, one answer, three times

Scenario: The coordinator spawns three search subagents for "renewable energy storage". All three are given the same topic string. The returned findings overlap heavily — the same six articles appear in all three result sets — while several subdomains go uncovered.

Question: How should the work be partitioned?

A) Drop to a single search subagent and stop paying triple.

B) Deduplicate the merged results at the coordinator, so synthesis receives one clean list instead of three overlapping ones.

C) Assign each subagent a distinct subtopic or source type, so scopes do not overlap.

D) Have each subagent check what the others have already returned, so no article is fetched twice.


Answer: Partition the scope across subagents

Why: partitioning research scope — distinct subtopics or distinct source types per agent — is what makes parallel search additive rather than redundant (Task 1.2). It also addresses the second symptom: uncovered subdomains become visible as unassigned partitions.

Why the others are wrong:

  • The cost observation is right, but the cure throws away the parallelism instead of directing it.
  • A clean merged list still costs three redundant searches, and it says nothing about the subdomains nobody covered.
  • Requires inter-subagent visibility, which breaks isolation and serialises the very work that was meant to run in parallel.
Q51 The synthesis has holes

Scenario: The synthesis agent's first output covers two of the five subdomains well and mentions the other three only in passing, because the underlying searches returned little on them. Today the coordinator forwards this synthesis straight to report generation.

Question: What should the coordinator do instead?

A) Re-run the entire pipeline from the beginning with a broader topic string.

B) Instruct the report agent to write around the gaps so the report reads as complete.

C) Accept the synthesis as it stands and note in the report that three areas are less developed than the other two, leaving the reader to weigh them.

D) Evaluate the synthesis for gaps, re-delegate targeted queries on the weak subdomains, then re-invoke synthesis until coverage is sufficient.


Answer: An iterative refinement loop driven by the coordinator

Why: the coordinator evaluates synthesis output for gaps, re-delegates with targeted queries, and re-invokes synthesis until coverage is adequate. Targeted re-delegation is what makes this cheaper than starting over (Task 1.2).

Why the others are wrong:

  • Discards the two subdomains that were covered well and pays the full cost again.
  • Is actively harmful: it disguises missing evidence as completed research.
  • Annotating a gap is correct only once you have tried to close it; here nothing has been attempted, so the reader is asked to weigh evidence that was never gathered.
Q52 Two strategies from one baseline

Scenario: After an expensive analysis pass that mapped the source landscape for a topic, you want to compare two synthesis strategies — one organised chronologically, one organised by claim — without repeating the analysis and without letting either branch contaminate the other.

Question: Which mechanism fits?

A) Two fresh sessions, each re-running the analysis with a different synthesis instruction.

B) --resume with the analysis session name, running one strategy then the other in sequence.

C) One session that produces both strategies in a single response for comparison.

D) fork_session, creating two independent branches from the shared analysis baseline.


Answer: fork_session

Why: forking creates independent branches from a shared baseline, which is the documented way to explore divergent approaches without repeating the work that produced the baseline (Task 1.7).

Why the others are wrong:

  • Pays for the expensive analysis twice.
  • Resuming the same session runs the second strategy in a context already shaped by the first — the contamination you wanted to avoid.
  • Producing both in one response makes each aware of the other, which is not an independent comparison.
Q53 Where a transient failure should be handled

Scenario: The search subagent's first API call fails with a connection reset. It currently propagates the failure to the coordinator immediately. The coordinator, having no better information, retries the whole subagent invocation from scratch — re-running the two queries that had already succeeded.

Question: Where should this failure be handled?

A) In the coordinator: it has the full picture and is the only place a retry budget can be enforced.

B) In a top-level handler that restarts the research workflow on any failure, so one uniform recovery path covers everything.

C) In the subagent, with unlimited retries until the call finally succeeds.

D) In the subagent: recover locally, and propagate only the errors it cannot resolve.


Answer: Local recovery first, then structured propagation of what remains

Why: subagents implement local recovery for transient failures and propagate to the coordinator only the errors they cannot resolve, carrying the partial results and a record of what was attempted. That is what stops the coordinator from re-running work that already succeeded (Tasks 2.2 and 5.3).

Why the others are wrong:

  • A central retry budget is worth having, but making the coordinator the only retry point produces exactly the described waste: its sole lever is re-invoking the whole subagent.
  • One uniform path is easy to reason about and ruinous in practice — restarting everything over a single branch is the anti-pattern of terminating a workflow over one failure.
  • Sparing the coordinator every transient error sounds tidy, but a retry loop with no exit burns latency and budget and hides a genuine outage as a hang.
Q54 Zero results is not a failure

Scenario: Three source categories return different outcomes: academic databases return 15 articles, industry reports return 0 results, and the patent database returns a connection timeout. Your subagent reports the last two identically, as "source failure".

Question: How should these be reported to the coordinator?

A) Distinguish the access failure, where the query never completed, from the valid empty result, which is coverage information.

B) Aggregate the three into a single "67% source coverage" metric for the coordinator.

C) Report both as successes with empty results, so the workflow finishes instead of stalling on one source.

D) Report both as failures, so the coordinator retries both and errs on the side of caution.


Answer: Access failure and valid empty result are different things

Why: a timeout means the question was never answered and a retry decision is still owed; zero results means it was answered and the answer is "nothing here", which is coverage information. Conflating them makes the coordinator either retry pointlessly or record a gap that does not exist (Task 5.3).

Why the others are wrong:

  • One figure is easy to act on and impossible to act on correctly: the percentage destroys precisely the distinction the coordinator needs.
  • Keeps the run moving by silently converting a real failure into an apparent absence of evidence.
  • Retrying a query that legitimately returned nothing will return nothing again, at cost.
Q55 Reporting on incomplete evidence

Scenario: Two of five subdomains could not be researched: one source timed out repeatedly, one database was unavailable for the whole run. The remaining three are well covered. The deadline is now and the report must go out.

Question: How should the synthesis be structured?

A) Fill the two gaps with the model's general knowledge, clearly written in a more cautious tone.

B) Report only the three well-covered subdomains, so every claim that ships is fully sourced and verifiable.

C) Ship the report with explicit coverage annotations on each section.

D) Delay the report until all five subdomains can be covered, so the reader gets one complete picture.


Answer: Coverage annotations on the delivered report

Why: graceful degradation means proceeding with partial results while making the limits of the evidence explicit — an annotation says the gap comes from an unavailable source, not from an absence of findings. Marking which sections are well supported and which rest on a source that never answered is what lets a reader act on the report; without it, they cannot (Task 5.3).

Why the others are wrong:

  • Substitutes unsourced model output for research, in a report whose value is that it is sourced.
  • Every shipped claim is sourced, and silent omission is still indistinguishable, to the reader, from "there was nothing to say".
  • Trades a known, annotated limitation for a missed deadline, and the complete picture may never arrive.
Q56 Attribution lost at the summarisation step

Scenario: Subagents return well-attributed findings. The coordinator summarises each subagent's output before passing it to synthesis, to save context. The final report contains accurate claims with no traceable sources, and reviewers cannot verify anything.

Question: What is the fix?

A) Add a bibliography section listing every source consulted, so reviewers can reach the underlying material and check any claim against it themselves.

B) Have the report agent re-search each claim to rediscover a plausible source.

C) Require structured claim-source mappings that every downstream step, including the coordinator's summarisation, must preserve.

D) Stop summarising subagent output and pass everything through verbatim.


Answer: Claim-source mappings preserved through every step

Why: attribution is lost when findings are compressed without preserving the claim-to-source link. Making that mapping a structural requirement that each step must carry forward is what survives summarisation (Task 5.6).

Why the others are wrong:

  • Reviewers reach the material but cannot check a claim against it without re-reading every source — a global list does not say which source backs which sentence, the exact information that was lost.
  • Restores citations that merely look right: re-deriving a source for an existing claim invites confident mis-attribution.
  • Nothing is dropped and nothing fits either — it removes the compression the coordinator needs, when the mapping can be preserved and the prose compressed.
Q57 Two credible sources, two numbers

Scenario: The document analysis agent finds two figures for the same statistic: 12% from a platform's own annual report, using automated classification, and 8% from an industry survey of 500 labels. Both sources are credible and the methodologies differ.

Question: What should the analysis agent pass to synthesis?

A) The average of the two values, noted as an estimate for synthesis to use.

B) The value from the larger or more authoritative source, with the other discarded as noise.

C) Both values, attributed and annotated as a conflict.

D) Neither, since contradictory data is not reportable and publishing both would undermine the report.


Answer: Both values, annotated as a conflict, with attribution

Why: conflicting statistics from credible sources are annotated with their attribution — source, date, methodology — rather than arbitrarily resolved. The analysis agent completes its work and lets the coordinator decide; the methodological difference is itself part of the finding (Task 5.6).

Why the others are wrong:

  • Gives synthesis one number to work with, at the price of fabricating a figure no source reports.
  • "more authoritative" is a judgement the analysis agent is not positioned to make, and calling the other figure noise silently deletes valid evidence.
  • A qualified conflict is reportable; withholding the statistic impoverishes the report instead of stating the uncertainty.
Q58 A contradiction that is really a trend

Scenario: The report flags a contradiction: "Source A states adoption is at 10%, while Source B states 15%." A reviewer checks and finds Source A was published in 2023 and Source B in 2024. Several similar false conflicts appear across the report.

Question: What should be required of subagent output?

A) A final pass asking the model to remove the contradictions it finds.

B) A confidence score on each figure, so the weaker one is dropped automatically.

C) A rule that when two figures disagree, the higher one is reported, so the pipeline stays deterministic and reviewers always see the same number twice running.

D) Publication or data-collection dates in the structured output, so a temporal difference reads as change rather than conflict.


Answer: Require dates in the structured output

Why: without temporal metadata, a measurement taken a year apart looks like a disagreement. Requiring publication or collection dates turns an apparent conflict into a legible trend (Task 5.6).

Why the others are wrong:

  • A cleanup pass makes the report read smoothly while removing the real contradictions along with the false ones.
  • Automates the wrong dimension: both figures may be entirely reliable, and what separates them is time, not confidence.
  • Deterministic and arbitrary at once — a repeatable tie-break still discards half the evidence and encodes an upward bias in every figure it touches.

6. Scénario 4 — Developer Productivity with Claude (Q59–Q72)

Contexte du scénario. Outils de productivité développeur sur le Claude Agent SDK. L'agent aide les ingénieurs à explorer des codebases inconnues, comprendre des systèmes legacy, générer du boilerplate et automatiser des tâches répétitives. Il utilise les outils built-in (Read, Write, Bash, Grep, Glob) et s'intègre à des serveurs MCP. Domaines dominants : D2, D3, D1.

Q59 Finding every caller of a function

Scenario: A developer asks: "Where is calculateTax() called from?" The agent needs to locate every call site across a codebase of several thousand files.

Question: Which built-in tool is appropriate?

A) Bash running a recursive listing, then Read on the results.

B) Read on every source file, checking each for the identifier.

C) Grep, which searches file contents for a pattern.

D) Glob, which matches file paths against a pattern.


Answer: Grep

Why: Grep is the content-search tool — function names, error messages, import statements. Finding callers is a content question, not a filename question (Task 2.5).

Why the others are wrong:

  • A listing gives paths, not matches, and still leaves the search to be done.
  • Reading every file to find one identifier is the pattern the tool exists to replace, and it would exhaust the context.
  • Glob matches paths; it cannot see inside files.
Q60 Locating all the test files

Scenario: You need every React component test in the repository. They are named *.test.tsx and sit next to the components they test, scattered across dozens of directories.

Question: Which tool fits?

A) Bash with find, since Glob does not support recursive patterns.

B) Glob with the pattern **/*.test.tsx.

C) Read on the test configuration to get the list of test paths.

D) Grep for the string describe( across the repository.


Answer: Glob with **/*.test.tsx

Why: finding files by name or extension pattern is exactly what Glob is for, and ** handles the recursion regardless of directory (Task 2.5).

Why the others are wrong:

  • Rests on a false premise — Glob supports recursive patterns.
  • The configuration usually holds patterns, not an enumerated list, so you would still have to resolve them.
  • Searching for describe( finds test content, missing files that use another style and matching non-test files that happen to contain the string.
Q61 Edit cannot find a unique anchor

Scenario: The agent tries to change a return null; statement in a 400-line file. Edit fails because return null; appears eleven times in that file.

Question: What is the correct fallback?

A) Delete the file and recreate it from scratch with Write.

B) Read the full file, then Write the modified contents back.

C) Use Bash with sed and a line number.

D) Retry Edit with the same anchor text until it happens to match the right occurrence.


Answer: Read then Write

Why: when Edit cannot find unique anchor text, loading the whole file and writing it back is the documented reliable fallback (Task 2.5).

Why the others are wrong:

  • Deleting first risks losing the content if the write fails, for no benefit over an in-place rewrite.
  • Shelling out for a targeted edit is fragile and bypasses the file tooling; line numbers shift as soon as anything else changes.
  • Edit is deterministic — the same ambiguous anchor fails identically every time.
Q62 Understanding an unfamiliar service

Scenario: A developer asks the agent to explain how a payment service works. The service has about 200 files. The agent's first move is to Read all of them.

Question: What is the better strategy?

A) Build understanding incrementally: Grep for entry points, then Read along the imports that matter.

B) Ask the developer which files to read, since they know the service best.

C) Read the files whose names contain "payment", since the convention marks them.

D) Read all 200 files but summarise each in one line to save context.


Answer: Grep for entry points, then Read along the call graph

Why: incremental exploration — locate entry points by content, then follow imports to trace flows — gets to an accurate picture without loading the whole service into context (Task 2.5).

Why the others are wrong:

  • The developer is asking precisely because they do not know the service either.
  • Naming conventions miss the collaborators — the gateway client, the ledger, the retry policy — that carry the actual behaviour.
  • Still reads 200 files; summarising afterwards does not recover the context already consumed.
Q63 A function re-exported through wrappers

Scenario: validateOrder is defined in one module, re-exported by a barrel file under the same name, and re-exported again by a facade under the alias checkOrder. A search for validateOrder finds three hits and misses every caller that imports the facade.

Question: How should the agent trace real usage?

A) Read every file that imports from the facade, since that is where the alias appears in the source.

B) Identify every name the function is exported under, then search for each.

C) Rename the function so every usage becomes consistent, then search once with no aliases left to miss.

D) Search for the original name and accept the gaps.


Answer: Enumerate the exported names first, then search for each

Why: tracing usage across wrapper modules means resolving the set of names the function travels under, then searching for each of them (Task 2.5). Skipping the first step is what produced the incomplete result.

Why the others are wrong:

  • The alias is indeed visible at each import site, but reading them covers only that one hop and misses any further re-export beyond it.
  • Modifies the codebase to make a read-only question easier, which is a disproportionate side effect for an answer a second search would have given.
  • Text search is not the limit here — enumerating the aliases first makes them findable, so this accepts a knowingly incomplete answer to a question about completeness.
Q64 Where two MCP servers belong

Scenario: Your team needs a Jira MCP server available to every developer on the project. Separately, you personally want to try an experimental local server that indexes your own notes; nobody else should get it.

Question: Where do you configure each?

A) Jira in .mcp.json; the experimental server in ~/.claude.json.

B) Both in .mcp.json, with the experimental one commented out so teammates never load it.

C) Both in ~/.claude.json, and ask each teammate to add the Jira entry themselves.

D) Jira in CLAUDE.md as documentation; the experimental server in .mcp.json.


Answer: Project scope for shared tooling, user scope for personal servers

Why: .mcp.json is the project-level, version-controlled scope for shared team tooling; ~/.claude.json is the user-level scope for personal or experimental servers. Both are discovered at connection time and available simultaneously (Task 2.4).

Why the others are wrong:

  • Teammates do not load it today; a commented-out entry in a shared file is a change waiting to be uncommented by accident.
  • Makes shared tooling depend on every teammate configuring it by hand — the failure mode project scope exists to prevent.
  • CLAUDE.md does not configure servers, and it inverts the two scopes.
Q65 A token in a committed file

Scenario: The GitHub MCP server needs an authentication token. .mcp.json is committed to the repository so every developer gets the server configuration.

Question: How do you supply the credential?

A) Move the whole server to ~/.claude.json on each machine.

B) Put a shared team token in .mcp.json and restrict repository access, so onboarding stays effortless and nobody has to configure anything.

C) Reference it with environment variable expansion — ${GITHUB_TOKEN} — so each developer supplies their own value.

D) Store the token in CLAUDE.md, documentation rather than configuration, so no scanner flags it.


Answer: Environment variable expansion in .mcp.json

Why: ${VAR} expansion in .mcp.json is the documented mechanism for credential management without committing secrets — the configuration is shared, the value is not (Task 2.4).

Why the others are wrong:

  • Keeps the token off the repository by giving up the shared configuration, which is the thing project scope provides.
  • Onboarding with nothing to configure does not undo a committed secret: repository permissions do not remove it, and it lands in the history permanently.
  • Being unflagged is not being safe — it puts a credential in a file that is both committed and loaded into every prompt.
Q66 The agent ignores your MCP tool

Scenario: You built an MCP tool that queries a pre-built code index and returns callers, definitions and type information in one call. The agent keeps using the built-in Grep instead, producing slower and less complete answers. Your tool's description reads "Searches the code index."

Question: What is the most effective fix?

A) Rename the MCP tool to grep_advanced so it is chosen by name similarity with the built-in.

B) Add a system prompt rule stating that the MCP tool must always be used before Grep.

C) Expand the tool's description to state what it returns and when to prefer it.

D) Remove Grep from the agent's allowedTools so it has no alternative.


Answer: Strengthen the MCP tool's description

Why: agents tend to prefer built-in tools, and the documented counter-measure is enhancing MCP tool descriptions so they explain capabilities, outputs and when to prefer them over a plain text search, in enough detail to be chosen on merit (Task 2.4). A one-line description cannot compete with a well-understood built-in.

Why the others are wrong:

  • Name similarity is not the selection mechanism; the description is.
  • An "always" rule in the system prompt is both probabilistic and wrong in the cases where Grep is the better choice.
  • Removes a genuinely useful tool to work around a description problem — plain text search is still the right tool sometimes.
Q67 Discovering what data is available

Scenario: Your internal MCP server exposes 40 database schemas. Agents currently discover them by calling describe_schema repeatedly with guessed names, burning several turns before finding the right one.

Question: What should the server expose?

A) A prompt in the server that instructs the agent to guess names more efficiently.

B) A single list_and_describe_all_schemas tool returning all 40 schemas in full, so one call answers everything.

C) Documentation of the schema names in the project CLAUDE.md, where every session already reads it.

D) MCP resources presenting a catalogue of the available schemas, so the agent sees what exists.


Answer: Expose the catalogue as MCP resources

Why: resources exist to expose content catalogues — issue summaries, documentation hierarchies, database schemas — giving agents visibility into available data without exploratory tool calls (Task 2.4). Tools are for actions; resources are for what exists.

Why the others are wrong:

  • MCP prompts are not part of this exam's scope, and better guessing is still guessing.
  • One call, and 40 full schemas land in context to answer "which one do I need?".
  • Every session reads a copy the server never updates — duplicated server-owned data drifts out of date.
Q68 Build or adopt an MCP server

Scenario: You need two integrations: Jira, used in the standard way, and an in-house deployment system with a bespoke API and team-specific workflows. A well-maintained community MCP server exists for Jira.

Question: What is the right split?

A) Community server for Jira; custom server for deployment.

B) Adopt community servers for both, extending whichever one comes closest to the deployment workflow.

C) Wrap both behind a single custom server that proxies to Jira and to deployment.

D) Build custom servers for both, to keep the stack under your control.


Answer: Community server for the standard integration, custom for the bespoke one

Why: the guidance is to choose existing community MCP servers for standard integrations and reserve custom servers for team-specific workflows (Task 2.4). Jira is the standard case; the in-house system has no equivalent.

Why the others are wrong:

  • There is nothing close enough to extend: the in-house system has a bespoke API no community server was written against.
  • A proxy layer adds a component with no functional benefit over configuring two servers.
  • Rebuilds a solved integration and inherits its maintenance forever.
Q69 Eighteen tools, unreliable choices

Scenario: Your boilerplate-generation agent has 18 tools: five for file operations, four for template rendering, three for dependency lookups, three for git, and three for issue tracking. It routinely picks a plausible but wrong tool, and adding clearer descriptions has produced only modest improvement.

Question: What is the underlying problem?

A) Too many tools degrade selection reliability; restrict each agent to the tools its role actually needs.

B) The tools should be merged into one generic tool with a mode parameter.

C) The agent needs tool_choice: "any" so it commits to a tool instead of hesitating.

D) The tool descriptions must be rewritten again, at greater length and with worked examples.


Answer: Scope the tool set to the agent's role

Why: giving an agent access to too many tools degrades selection reliability by increasing decision complexity, and agents with tools outside their specialisation tend to misuse them. The fix is scoped access — the four or five tools the role actually needs, plus narrowly scoped cross-role tools for specific high-frequency needs (Task 2.3).

Why the others are wrong:

  • One choice on the surface only: the same 18-way decision moves inside a parameter, where the description can no longer differentiate it.
  • "any" guarantees that a tool is called, not that the right one is.
  • Descriptions have already been improved with modest effect — the scenario tells you the bottleneck is elsewhere.
Q70 A tool that can fetch anything

Scenario: Your documentation agent has a generic fetch_url tool. It is being used to pull arbitrary pages — search result listings, unrelated blogs, an internal admin endpoint — when the intended use was retrieving known documentation URLs.

Question: What is the appropriate change?

A) Replace fetch_url with a constrained load_document tool that validates the URL is a documentation source before fetching.

B) Keep fetch_url and add a system prompt rule listing the domains that are allowed, so the agent knows which sources are in scope before it calls.

C) Keep fetch_url and post-filter the responses, discarding content from unexpected domains.

D) Remove the tool and have the agent ask a human to fetch pages for it.


Answer: Replace the generic tool with a constrained alternative

Why: replacing generic tools with constrained alternatives that validate their inputs is the documented pattern — the constraint lives in the tool, where it holds, rather than in an instruction (Task 2.3).

Why the others are wrong:

  • Stating the scope in advance still leaves the capability in place, and depends on compliance every single time.
  • The fetch has already happened by then, including against the internal endpoint.
  • Removes a legitimate capability and puts a human in a loop that does not need one.
Q71 Resuming after changing the code

Scenario: Yesterday you ran a long investigation session named refund-flow that mapped the refund dependencies. Overnight, two of the analysed files were refactored by a teammate. You want to continue the investigation today.

Question: What is the right approach?

A) Fork the session so the old and new versions of the two refactored files can be compared side by side.

B) Resume with --resume refund-flow and list the files that changed.

C) Start a fresh session and re-explore the whole refund flow from scratch, so nothing stale survives.

D) Resume with --resume refund-flow and carry on.


Answer: Resume the named session and state which files changed

Why: named session resumption keeps the prior analysis, which is still mostly valid, and informing the agent about specific file changes lets it re-analyse those targets rather than redo everything (Task 1.7).

Why the others are wrong:

  • Forking explores divergent approaches from a shared baseline; both branches inherit the same stale reading, so nothing gets compared with what is on disk today.
  • Nothing stale survives because nothing survives — it discards a full day of valid analysis over two changed files.
  • The session's picture of those two files is now wrong, and the agent has no way to know it.
Q72 When resuming is the wrong choice

Scenario: A session from three weeks ago holds an architecture analysis. Since then the module was restructured, three files were deleted and the data layer was replaced. Nearly every tool result in that session describes code that no longer exists.

Question: Resume or start fresh?

A) Resume the session and list the changes since then, as in any ordinary resumption.

B) Start fresh, injecting a summary of the conclusions that still hold.

C) Fork the old session so the stale analysis is preserved for reference alongside the new work.

D) Resume the session and run /compact first to discard the stale details before working.


Answer: Fresh session with an injected summary

Why: the choice turns on how much of the prior context is still valid. Resumption is right when it mostly is; when tool results are broadly stale, starting fresh with a structured summary of the durable conclusions is the more reliable option (Task 1.7).

Why the others are wrong:

  • Listing changes works for a couple of edited files, not for a context where almost every observation is obsolete.
  • Preserves the obsolete analysis without producing a usable working context — the fork carries the same dead tool results.
  • /compact summarises the stale content rather than removing it — it cannot tell what is out of date.

7. Scénario 5 — Claude Code for Continuous Integration (Q73–Q86)

Contexte du scénario. Claude Code intégré au pipeline CI/CD : revues de code automatisées, génération de cas de test, feedback sur les pull requests. Il faut concevoir des prompts qui produisent un feedback actionnable et minimisent les faux positifs. Domaines dominants : D3, D4.

Q73 From review output to inline PR comments

Scenario: Your CI job runs a review and receives prose findings. A script tries to parse file paths and line numbers out of the text with regular expressions to post inline PR comments; roughly a fifth of the findings are dropped or mis-anchored.

Question: What is the correct approach?

A) Run with --output-format json and --json-schema so the findings come back machine-parseable and schema-conformant.

B) Post the whole review as a single top-level PR comment.

C) Improve the regular expressions and add fallbacks for the formats that fail.

D) Instruct the prompt to emit a strict Markdown table and parse that.


Answer: --output-format json with --json-schema

Why: these flags exist to enforce structured output in CI contexts, producing machine-parseable findings that can be posted directly as inline comments (Task 3.6). Parsing prose is the problem being solved, not a step to optimise.

Why the others are wrong:

  • Gives up the inline placement that makes review feedback actionable.
  • Iterating on regular expressions treats the symptom of an unstructured contract.
  • A Markdown table is still generated text with no structural guarantee — one malformed row and the parse breaks again.
Q74 Generated tests nobody wants

Scenario: Your CI test-generation step produces tests that assert getters return what was set, mock everything including the unit under test, and ignore the project's fixture helpers entirely. Reviewers delete most of them.

Question: What is the most effective change?

A) Cap the number of generated tests to five per pull request, keeping the review burden small.

B) Document the testing standards, what makes a test valuable, and the available fixtures in CLAUDE.md.

C) Add "write only high-value tests that reviewers will actually keep" to the CI prompt.

D) Have a second Claude instance delete the low-value tests after generation, before reviewers see them.


Answer: Put testing standards and fixtures in CLAUDE.md

Why: CLAUDE.md is the mechanism for giving CI-invoked Claude Code its project context — testing standards, fixture conventions, review criteria. Documenting what a valuable test looks like and which fixtures exist is what lifts generation quality and cuts low-value output (Task 3.6).

Why the others are wrong:

  • Caps the volume of low-value tests without improving any of them.
  • "high-value" is a vague instruction of exactly the kind the guide says fails to change behaviour — the model does not know which tests your reviewers keep.
  • Spends a second review pass compensating for missing context that could simply be supplied.
Q75 Reviewing your own work

Scenario: Your pipeline generates an implementation and then, in the same session, asks the model to review what it just wrote. The reviews are consistently approving and miss defects that a human reviewer finds minutes later.

Question: What is the fix?

A) Ask for the review three times and take the union of the findings.

B) Add an instruction telling the model to be critical of its own work and to assume the code it has just written contains bugs.

C) Enable extended thinking on the review step so it reasons longer.

D) Run the review in a separate, independent instance that does not carry the generator's reasoning context.


Answer: An independent review instance

Why: a model that has just generated code retains the reasoning that produced it, which makes it less likely to question those decisions. An independent instance without that context catches subtle issues that self-review instructions do not (Task 4.6).

Why the others are wrong:

  • Three passes in the same session share the same blind spot.
  • Being told to expect bugs does not remove the reasoning context that is causing the blind spot — the model still believes the decisions it defended a moment ago.
  • More reasoning within the same context reinforces the original conclusions rather than challenging them.
Q76 The same comment, five times

Scenario: The review runs on every push. A pull request updated six times now carries the same three unresolved findings posted six times each, plus repeated comments on issues the author already fixed.

Question: How should re-runs be handled?

A) Deduplicate comments in the posting script by comparing their text, so the same wording is never posted twice on one pull request.

B) Run the review only on the first push of each pull request, so a finding can be raised at most once.

C) Review only the diff of the latest push.

D) Include the prior findings and ask for new or still-unaddressed issues only.


Answer: Feed prior findings back in and ask for new or unresolved issues only

Why: supplying the previous findings lets the model distinguish what is new, what has been fixed and what still stands — which is what prevents duplicate comments without losing coverage (Task 3.6).

Why the others are wrong:

  • Text matching stops literal repeats, misses re-phrased findings, and cannot tell a resolved issue from a duplicated one.
  • Nothing repeats because nothing is reviewed — every subsequent commit goes unchecked, which is where regressions get introduced.
  • A narrow diff view misses issues where the new code interacts with code changed two pushes ago.
Q77 Tests that already exist

Scenario: Test generation proposes a null-input case and an empty-list case for a function whose existing suite already covers both, in the same file. Reviewers spend their time discarding duplicates.

Question: What is the fix?

A) Include the existing test files in context.

B) Generate tests only for files that currently have no test file at all.

C) Compare generated test names against existing ones and drop the collisions.

D) Instruct the prompt to generate only exotic edge cases.


Answer: Put the existing tests in context

Why: the model cannot avoid duplicating what it has not seen. Supplying the existing test files is the documented way to make generation additive (Task 3.6).

Why the others are wrong:

  • Abandons coverage improvement on every file that has partial tests — which is most of them.
  • Name matching misses semantic duplicates written under a different name.
  • Pushes generation towards improbable cases and away from genuine gaps.
Q78 "Be conservative" does not work

Scenario: Your comment-accuracy check flags any comment that is not a literal restatement of the code beneath it, producing dozens of false positives per pull request. You have already added "be conservative" and "only report high-confidence findings" to the prompt, with no measurable improvement.

Question: What actually improves precision?

A) Replace the general instruction with a specific criterion: flag a comment only when it contradicts the code.

B) Add "and please be very careful about false positives" to the instruction.

C) Have the model attach a confidence score and suppress findings below 0.8.

D) Reduce the number of files reviewed per run so each one gets more careful attention.


Answer: A specific categorical criterion

Why: general exhortations like "be conservative" or "only high-confidence findings" do not improve precision; explicit criteria defining what counts as a reportable issue do. Requiring that the behaviour the comment claims contradict what the code actually does converts a matter of taste into a testable condition (Task 4.1).

Why the others are wrong:

  • Reinforces the very instruction that has already been shown not to work.
  • Cuts the volume without changing what the model reports as a defect. Task 4.1 names the remedy the other way around: write specific review criteria "rather than relying on confidence-based filtering".
  • Attention is not the constraint; the definition of a defect is.
Q79 One bad category poisons the rest

Scenario: Your review reports five categories. Four are accurate and valued. The fifth, "performance concerns", is wrong roughly 70% of the time. Developers have started dismissing all review comments without reading them, including the accurate ones.

Question: What should you do?

A) Remove the performance category permanently, as beyond recovery.

B) Keep all five and let developers filter the category they distrust.

C) Temporarily disable the performance category to restore trust, and keep improving its prompt offline.

D) Keep all five and add a disclaimer marking performance findings as experimental, so readers know to weigh them differently from the rest.


Answer: Temporarily disable it while you improve the prompt

Why: a high false-positive category undermines confidence in the accurate ones. Disabling it restores the signal-to-noise ratio of the whole review immediately and returns trust to the four categories that work, while the fix for the fifth continues in parallel (Task 4.1).

Why the others are wrong:

  • Over-corrects: the category is fixable, and "permanently" gives up a genuine class of defect.
  • Puts the filtering burden on the people whose trust you are trying to regain.
  • Readers weigh a finding only after reading it — the disclaimer leaves the noise sitting alongside the good findings, which is what taught them to skip the lot.
Q80 Severity that shifts between runs

Scenario: The same class of issue is labelled "critical" in one pull request and "minor" in the next. Your prompt says to assign severity "based on impact". Downstream automation blocks merges on critical findings, so the inconsistency is expensive.

Question: How do you get consistent classification?

A) Define explicit severity criteria with a concrete code example for each level.

B) Ask the model to justify each severity assignment in prose.

C) Reduce the scale from five levels to two, critical and non-critical.

D) Have a second pass re-rate the severities assigned by the first.


Answer: Explicit criteria with a concrete example per level

Why: "based on impact" is a judgement the model re-derives each run. Defining the levels explicitly, and anchoring each with a code example, is what produces consistent classification (Task 4.1).

Why the others are wrong:

  • A justification explains a decision made on unstated criteria; it does not stabilise it.
  • Fewer levels means fewer ways to disagree, but the boundary between the two remains undefined.
  • A second pass with the same vague definition is inconsistent in the same way.
Q81 Findings in five different shapes

Scenario: Your review prompt lists in detail what each finding must contain — location, issue, severity, suggested fix. The output still varies: some findings are a single sentence, others three paragraphs, some omit the suggested fix, some bury the location in prose.

Question: What is the most effective technique?

A) Add two to four few-shot examples showing findings in exactly the desired shape.

B) Split the review into four passes, one per required element.

C) Rewrite the instructions again with stronger wording about the required format.

D) Post-process the output to reformat findings into the target shape.


Answer: Few-shot examples of the desired output

Why: when detailed instructions alone produce inconsistent output, few-shot examples are the most effective technique for achieving consistent, actionable formatting — a demonstration pins down what a description leaves open (Task 4.2).

Why the others are wrong:

  • Multiplies cost and fragments findings that belong together.
  • Is the approach that has already failed; the instructions are described as detailed.
  • Cannot reconstruct a suggested fix that was never produced.
Q82 Idiomatic code flagged as a defect

Scenario: The reviewer flags your codebase's deliberate patterns as problems: the repository's standard error-wrapping helper, its intentional use of a mutable builder in hot paths, and its convention of returning early from validators. You cannot enumerate every accepted pattern in advance — new ones appear regularly.

Question: What is the most effective fix?

A) Maintain an explicit allow-list of accepted code patterns, updated by whoever first notices the reviewer flagging one of them wrongly.

B) Provide few-shot examples contrasting the codebase's accepted patterns with genuine defects, so the distinction generalises.

C) Exclude from review the files where these patterns occur, so the reviewer stops flagging them.

D) Lower the review's sensitivity so fewer findings of any kind are reported.


Answer: Few-shot examples that teach the distinction

Why: few-shot examples reduce false positives while enabling generalisation — the model learns the principle separating an intentional pattern from a defect, rather than matching a fixed list. The scenario's constraint, that new patterns keep appearing, is what makes generalisation the requirement (Task 4.2).

Why the others are wrong:

  • An allow-list only ever covers what someone has already added to it, and the scenario states new patterns appear faster than that.
  • Stops the false positives by stopping review of those files, so the genuine defects living alongside the deliberate patterns go unreported too.
  • Suppresses real defects at the same rate as false ones.
Q83 Understanding which findings get dismissed

Scenario: Developers dismiss about 30% of findings. You want to know systematically what kind of code triggers the dismissed ones, so you can fix the prompt rather than guess.

Question: What should you add to the structured output?

A) A free-text explanation field for developers to fill in whenever they dismiss a finding, capturing their reasoning.

B) A model-generated confidence score on each finding, on the assumption that dismissals track low confidence.

C) A detected_pattern field on each finding, recording the code construct that triggered it.

D) A monthly manual review of a sample of dismissed findings.


Answer: A detected_pattern field

Why: recording the construct is what lets dismissals be grouped, which makes the pattern analysable — you can see that, say, 80% of dismissals come from one construct and target the prompt at it (Task 4.4).

Why the others are wrong:

  • Depends on developers writing explanations at the moment they are trying to move on; response rates on such fields are poor.
  • Confidence tracks the model's certainty, not the construct that misled it, so the correlation never names the pattern you would rewrite the prompt against.
  • Manual sampling is slow, covers a handful of cases a month, and never scales to the systematic answer the scenario asks for.
Q84 Routing review attention

Scenario: Your review produces around 40 findings per large pull request. Reviewers have time for roughly ten. You want the ten they read to be the ten most likely to be real.

Question: What should the review pass produce?

A) Only the ten findings the model considers most important, with the rest discarded unread.

B) A self-reported confidence on each finding, used to route human attention.

C) Findings sorted alphabetically by file, so reviewers work through them in a predictable order.

D) A single aggregate quality score for the pull request, so reviewers can gauge overall risk.


Answer: Per-finding confidence used for calibrated routing

Why: running verification passes where the model self-reports confidence alongside each finding is what enables calibrated review routing (Task 4.6). The thresholds are then calibrated against outcomes over time, so the routing improves as you learn which confidences held up. Note the difference from Q3: confidence is unreliable as an autonomous decision to escalate, and useful as a prioritisation signal for a human who still sees the finding.

Why the others are wrong:

  • Discarding 30 findings makes the model's judgement final and irreversible; routing keeps them available.
  • A predictable order is still an arbitrary one: file names carry no information about which findings are likely to be real.
  • A gauge of overall risk still tells reviewers nothing about which of the forty findings to spend their time on.
Q85 Iterating towards a correct implementation

Scenario: You are having Claude implement a date-range parser with awkward requirements: overlapping ranges, open-ended ranges, and time zones. Each round of "that's not quite right, try again" produces a different set of defects.

Question: What is the most effective way to iterate?

A) Write a test suite first covering expected behaviour, edge cases and performance requirements, then iterate on the failures.

B) Implement it yourself and ask Claude to review your version, since you understand the awkward requirements best.

C) Ask for three independent implementations and pick whichever one shows the fewest defects on inspection.

D) Describe each defect in prose, in more detail each round, until the description leaves the implementer no room to misread what you want.


Answer: Tests first, then iterate on the failures

Why: test-driven iteration replaces a subjective judgement — "not quite right" — with a specific, reproducible signal. Sharing the failures guides progressive improvement instead of restarting from a new description each round (Task 3.5).

Why the others are wrong:

  • Your grasp of the requirements is better spent writing the tests; handing back the implementation discards the automation for a task it is well suited to.
  • Inspection of three unverified implementations picks the least-bad one and leaves you with no mechanism for improving it.
  • However precise the prose, this is the loop that is already failing; a description of defects is exactly what is being read inconsistently each round.
Q86 Batch submissions under a 30-hour SLA

Scenario: A nightly analysis processes documents through the Message Batches API. Your commitment to internal users is that any document submitted is analysed within 30 hours. Batch processing can take up to 24 hours. Once a batch completes, retrieving the results, re-running as a real-time call any document that came back as an error, and publishing the output take up to 2 more hours.

Question: How often must you submit batches to guarantee the SLA, resubmission of a failed document included?

A) Once every 24 hours, matching the maximum processing time.

B) Every 4 hours at most, so that queueing, processing, and the post-batch retry window all fit inside the 30 hours.

C) Every 6 hours at most — a document arriving just after a submission waits up to 6 hours, plus up to 24 hours of processing.

D) The submission frequency is irrelevant, since batches usually complete in under an hour.


Answer: Every 4 hours at most

Why: the worst case is queueing time, plus 24 hours of processing, plus the 2 hours it takes to retrieve results and re-run a document that failed and is identified by its custom_id. That leaves 30 − 24 − 2 = 4 hours of queueing, so a document arriving just after one submission is still delivered inside the commitment. On this same 30-hour / 24-hour pair, the official example retains 4-hour windows (Task 4.5).

Why the others are wrong:

  • 24 + 24 = 48 hours worst case, a breach.
  • The most attractive distractor, and the subtraction is right for queueing plus processing alone — but 6 + 24 already fills the 30 hours, so the failed document has no time left to be re-run and delivered.
  • Typical latency is not a guarantee, and an SLA is a statement about the worst case.

8. Scénario 6 — Structured Data Extraction (Q87–Q100)

Contexte du scénario. Système d'extraction de données structurées : extraction depuis des documents non structurés, validation par JSON Schema, maintien d'une haute exactitude. Il doit traiter les cas limites proprement et s'intégrer aux systèmes en aval. Domaines dominants : D4, D5.

Q87 What tool use guarantees

Scenario: You moved from asking for JSON in prose to defining an extraction tool with a JSON schema and reading the tool_use response. Parse failures dropped to zero. However, a downstream reconciliation job still rejects about 3% of extractions: line items that do not sum to the stated total, and a shipping address occasionally placed in the billing field.

Question: How should you interpret this?

A) The schema is not strict enough; tightening the field constraints will remove the remaining reconciliation rejections.

B) Expected: the schema eliminates syntax errors, not semantic ones; those need programmatic validation.

C) The extraction tool is misconfigured: conformant output should be semantically correct.

D) The 3% represents documents too complex for the current model, and they should be routed to a larger one instead.


Answer: Syntax is guaranteed; semantics is not

Why: a schema constrains structure and types. It cannot know that line items must sum to the total, or that a given address is the shipping one. Those are semantic properties, checked by validation code after extraction (Task 4.3).

Why the others are wrong:

  • No JSON Schema construct expresses "these numbers must add up to that number", so no amount of tightening reaches this class of error.
  • Rests on a misconception — conformance to a schema says nothing about the truth of the values.
  • A different model produces the same class of error; the missing layer is validation, not capacity.
Q88 A required field with nothing to fill it

Scenario: Your schema marks company_name as required. About one document in five is a personal receipt with no company on it. Reviewers find plausible-looking company names in those extractions that appear nowhere in the source.

Question: What is the fix?

A) Add an instruction telling the model never to invent a company name.

B) Add an enum of the known company names so the model must choose from a list.

C) Make the field nullable — "type": ["string", "null"] — and remove it from required.

D) Keep the field required and validate afterwards that the value appears in the source text.


Answer: Make the field nullable and optional

Why: a required field puts the model between violating the schema and inventing a value, and it invents. Designing fields as nullable when the source may not contain the information is what prevents fabrication: null becomes a legal answer, so the model can say the receipt named no company (Task 4.3).

Why the others are wrong:

  • Leaves the contradiction in place: the schema still demands a value.
  • Forces a wrong choice from the list when the true answer is "none".
  • Detects the fabrication after the fact and produces a failure with no correct value to substitute.
Q89 Categories that do not fit

Scenario: Your category enum is ["billing", "technical", "account"]. Around 8% of documents concern something else — partnership enquiries, legal notices, press requests — and are currently forced into "account", where they are invisible to downstream routing.

Question: How should the schema change?

A) Replace the enum with a free-text string field, so nothing that arrives can fall outside it.

B) Add an "other" enum value plus a nullable category_detail field.

C) Add the three new categories you have observed to the enum, restoring routing for those documents.

D) Keep the enum and add a boolean is_uncategorised flag, so misfiled documents can at least be counted.


Answer: "other" plus a detail field

Why: the "other" + detail-string pattern keeps the enum useful for the common cases while the free-text detail records what the document actually was, preserving the information that falls outside the closed set — which is what makes the category set extensible (Task 4.3).

Why the others are wrong:

  • Nothing falls outside a free-text field because it constrains nothing, and routing needs a closed set of values to switch on.
  • Restores routing for the three kinds you have already seen, and fails again on the fourth kind, next month.
  • Counting the misfiled documents still leaves you without the one thing routing needs: what they actually were.
Q90 The document is genuinely ambiguous

Scenario: A sentiment field with enum ["positive", "negative", "neutral"] is applied to customer letters. Reviewers find that letters mixing praise and complaint are labelled inconsistently — sometimes positive, sometimes negative — and the label is later treated as fact.

Question: What is the appropriate schema change?

A) Instruct the model to pick whichever sentiment dominates the letter, which does yield one consistent rule.

B) Route every mixed letter to a human before extraction runs, so a person decides.

C) Add an "unclear" enum value so the model can record genuine uncertainty instead of guessing.

D) Add a numeric sentiment score from −1 to 1 instead of an enum.


Answer: An "unclear" enum value

Why: an enum with no honest exit forces an arbitrary choice, and downstream systems then treat the arbitrary choice as data. An honest "unclear" is more useful than a confident mislabel (Task 4.3).

Why the others are wrong:

  • A consistent rule for choosing between two genuine sentiments is still the arbitrary choice that is causing the inconsistency.
  • Requires knowing a letter is mixed before extraction — which is what extraction was meant to determine.
  • A score collapses "mixed" and "neutral" onto the same value near zero, hiding the distinction rather than exposing it.
Q91 A step that must run first

Scenario: Your pipeline has an extract_metadata tool that identifies the document type and issue date, and several enrichment tools whose behaviour depends on the document type. The model sometimes calls an enrichment tool first, which then runs with the wrong assumptions.

Question: How do you guarantee the ordering?

A) Describe the required order in each tool description, so the model reads the dependency before choosing.

B) Force the first call with tool_choice: {"type": "tool", "name": "extract_metadata"}, then handle the enrichment steps in follow-up turns.

C) Set tool_choice: {"type": "any"} so the model always calls a tool rather than answering from memory.

D) List extract_metadata first in the tools array, so it is the first candidate the model sees.


Answer: Force the specific tool for the first call

Why: forced tool selection guarantees that a named tool runs, which is how you ensure extraction precedes enrichment. Subsequent steps are then processed in follow-up turns, where the metadata is already available (Tasks 2.3 and 4.3).

Why the others are wrong:

  • A dependency stated in a description informs the choice probabilistically; the scenario asks for a guarantee.
  • "any" guarantees that a tool is called, not which one; answering from memory was never the failure, and the enrichment-first call remains possible.
  • Being listed first does not make a tool the first candidate — array order is not a selection mechanism.
Q92 Dates in four formats

Scenario: Source documents write dates as "15/01/2025", "Jan 15, 2025", "2025-01-15" and "the fifteenth of January". Your schema types the field as a string, and the output preserves whichever form appeared in the document. Downstream systems expect ISO 8601.

Question: What is the right combination?

A) Change the schema to accept any of the four formats and normalise downstream.

B) Keep the strict output schema and add explicit format normalisation rules to the prompt.

C) Add a regular-expression pattern to the schema field so only ISO 8601 validates.

D) Add a second tool call that converts the extracted dates afterwards, keeping the extraction prompt untouched.


Answer: Strict schema plus normalisation rules in the prompt

Why: the schema says what shape the output takes; the prompt says how to map inconsistent source formatting into it — every date emitted as ISO 8601, whichever form the document used. Including format normalisation rules alongside a strict output schema is the documented pairing (Task 4.3).

Why the others are wrong:

  • Pushes an ambiguity — is 01/02 January 2nd or February 1st? — to a consumer with less context than the extractor had.
  • A pattern rejects a non-conforming value without telling the model how to produce a conforming one.
  • Leaving the extraction prompt untouched costs a second call to do what the first could have done in one pass.
Q93 A validation failure worth retrying

Scenario: Pydantic validation rejects an extraction: invoice_date is "15/01/2025" where the model expects an ISO date, and line_items[2].quantity is the string "two" rather than an integer. Your pipeline currently retries with the original prompt unchanged.

Question: What should the retry contain?

A) The original document, the failed extraction, and the specific validation errors, so the model can correct what failed.

B) Only the failed extraction and the errors, without the document, to save tokens on the retry.

C) The original document and a firmer instruction to be more careful with field types this time.

D) The original document only, with a higher token budget so the model has all the room it needs to be thorough.


Answer: Document, failed extraction, and the specific errors

Why: retry-with-error-feedback works by telling the model precisely what was wrong. Without the errors it repeats the mistake; without the document it cannot re-derive the correct value (Task 4.4).

Why the others are wrong:

  • Saves tokens on the retry by removing the only source of the correct values.
  • A firmer instruction still does not identify which two fields failed, or why their values were rejected.
  • Room to be thorough changes nothing: an identical request with a larger budget reproduces the same output.
Q94 A retry that can never succeed

Scenario: Validation flags a missing purchase_order_number. Investigation shows the purchase order number is not in the invoice at all — it lives in a separate procurement record that is not supplied to the extractor. Your pipeline retries these three times before failing.

Question: How should this be handled?

A) Have the model infer the purchase order number from the vendor.

B) Recognise that a retry cannot create absent information, and route the case onward.

C) Increase the retry count, since some extractions genuinely do succeed on a later attempt than the third one.

D) Retry with a stronger instruction to search the whole document and its attachments, since the number must be somewhere.


Answer: Retries do not create information that is not there

Why: the distinction that matters is between format or structural errors — which retries fix — and information genuinely absent from the source, which they cannot (Task 4.4). Recognising the second class is what stops the wasted attempts and routes the case to the process that can supply the missing procurement record.

Why the others are wrong:

  • The vendor does not encode the number, so this manufactures a value for a field used in financial reconciliation.
  • Later attempts help with flakiness; this is a deterministic absence, and every further attempt reads the same document.
  • The number is not somewhere in the invoice waiting to be found — it lives in a record the extractor never receives, so a harder search returns the same absence.
Q95 Numbers that do not add up

Scenario: Some invoices contain arithmetic errors in the source itself: the printed total does not match the sum of the line items. Today the extractor silently reports the printed total, and the discrepancy surfaces weeks later in reconciliation.

Question: How should the schema be designed?

A) Have the model correct the line items so they match the printed total, keeping the record internally consistent.

B) Reject any invoice whose totals disagree and send it back to the supplier for correction.

C) Extract stated_total and calculated_total separately, with a conflict_detected flag.

D) Have the model correct the total so it matches the line items, since the arithmetic is the part that failed.


Answer: Both totals plus an explicit conflict flag

Why: extracting calculated_total alongside stated_total and flagging discrepancies is the self-correction validation design the guide describes. The extractor's job is to report what the document says and surface the inconsistency at extraction time, not to resolve it weeks later (Task 4.4).

Why the others are wrong:

  • Buys internal consistency by rewriting the line items, which assumes the printed total is the trustworthy side — a guess the extractor has no basis to make, and the stored record no longer matches the document.
  • Sending it back stalls a legitimate business record — a supplier's arithmetic error is still an invoice you must process, and the discrepancy is worth recording either way.
  • Assumes the arithmetic rather than the line items is what failed, which is the same unfounded guess in the opposite direction. Source data is altered either way, and reconciliation still has nothing telling it that the invoice disagreed with itself.
Q96 Twelve failures out of a hundred

Scenario: You submitted 100 documents to the Message Batches API. Eighty-eight succeeded; twelve failed because the documents exceeded the context limit.

Question: What is the correct recovery?

A) Resubmit all 100 with chunking applied uniformly, so no document can exceed the limit again.

B) Identify the twelve by custom_id and resubmit only those, chunked so each request fits within the limit.

C) Process the twelve through the synchronous API instead, which has no context limit.

D) Resubmit the twelve unchanged, since batch capacity varies between runs and a quieter run may accept them.


Answer: Resubmit only the failures, chunked, identified by custom_id

Why: custom_id correlates each request with its response, which is what lets you isolate the failures. Resubmitting them with the modification that addresses the cause — chunking oversized documents — is the documented handling (Task 4.5).

Why the others are wrong:

  • Prevents a recurrence by paying for 88 successful extractions a second time.
  • The synchronous API has the same context window — the limit is the model's, not the batch mechanism's.
  • Capacity varies, but the context limit does not: an unchanged oversized document fails identically on the quietest run.
Q97 Before submitting 50,000 documents

Scenario: You are about to run a 50,000-document backfill through the Batches API. Your prompt and schema have been tested on three documents you happened to have open.

Question: What should you do first?

A) Submit all 50,000 twice and use the two runs to see where they do not agree.

B) Submit all 50,000 and iterate on whatever fails, since the failures are the only real evidence.

C) Submit in ten batches of 5,000, adjusting the prompt between batches as problems reveal themselves.

D) Refine the prompt against a representative sample first, before committing the full volume.


Answer: Refine on a sample first

Why: prompt refinement on a sample before batch-processing large volumes is what maximises the first-pass success rate and avoids the cost of iterative resubmission at scale (Task 4.5). Three convenient documents are not a sample.

Why the others are wrong:

  • Agreement measures consistency, not correctness — two identically wrong runs agree perfectly, and the bill is doubled.
  • The failures are evidence bought at full price: a systematic prompt defect is far cheaper to find on a sample.
  • Better than B, but each 5,000-document batch still costs a full day of latency to learn something a small sample answers in minutes.
Q98 Documents that are laid out differently

Scenario: Your research-paper extractor handles papers with a bibliography reliably, but returns null for references on papers using inline citations, and misses methodology details when they are embedded in the results section instead of a dedicated one. The schema and instructions are correct.

Question: What is the most effective fix?

A) Instruct the model to read the whole document twice before extracting, so nothing is missed.

B) Add few-shot examples demonstrating correct extraction from each structural variety — inline citations as well as bibliographies.

C) Add a preliminary classification step that detects the document layout, then dispatches to whichever specialised extractor handles that form.

D) Mark the affected fields as nullable so the null results validate cleanly and stop failing the pipeline.


Answer: Few-shot examples covering the structural varieties

Why: few-shot examples showing correct extraction from varied document structures directly address empty or null extraction of fields that are present but laid out unexpectedly (Task 4.2). Cover every variety you see — embedded methodology as well as dedicated sections. The information is in the document; the model does not recognise it in that form.

Why the others are wrong:

  • A second reading misses the same material: the difficulty is recognising an unfamiliar form, not seeing it.
  • A dispatcher plus one specialised extractor per layout is a lot of machinery for a recognition problem that examples fix inside a single prompt, and every new layout adds another branch to maintain.
  • Stops the pipeline failing by making the wrong answer valid — the references exist, and null is incorrect.
Q99 Findings that vanish from the middle

Scenario: You aggregate the analyses of 30 contracts into a single request for a comparative summary. The summary reliably reflects the first four contracts and the last three; clauses from contracts 10 to 22 are frequently omitted, even though they are present in the input.

Question: How should the input be organised?

A) Ask for the summary as 30 separate per-contract responses.

B) Place a summary of the key findings first, then the detailed results under explicit headers.

C) Sort the contracts so the most important ones sit at the beginning and the end, where attention is reliable.

D) Move to a model with a much larger context window, so the whole comparison fits with room to spare.


Answer: Key findings first, detailed results under explicit headers

Why: models process the beginning and end of long inputs reliably and may omit the middle. Front-loading a findings summary and structuring the body with explicit headers are the documented mitigations (Task 5.1).

Why the others are wrong:

  • Thirty separate answers is not a comparative summary; the comparison across contracts is precisely what nobody produces.
  • Requires knowing which contracts matter before the comparison that determines it, and it still buries whatever lands in the middle.
  • Room to spare does not change where attention concentrates within the window; the middle is still the middle.
Q100 97% accuracy, one week before automating

Scenario: Your pipeline reports 97% aggregate accuracy across all documents and fields. Management wants to remove human review from the 97% and keep reviewers only for the rest.

Question: What must you verify first?

A) Confirm the 3% of failures are randomly distributed by checking that they come from many different customers.

B) Nothing — 97% is above the threshold the business already agreed to, so the decision is made.

C) Re-measure the aggregate on a larger sample to tighten the confidence interval.

D) Break the accuracy down by document type and by field, using stratified random sampling across every segment.


Answer: Break it down by document type and field

Why: an aggregate figure can conceal a segment performing far worse — scanned invoices at 60%, VAT numbers at 40% — while the average stays reassuring. Validating accuracy by document type and field segment is the precondition for automating high-confidence extractions (Task 5.5).

Why the others are wrong:

  • Customer spread says nothing about concentration by document type or by field.
  • An agreed threshold read off an aggregate figure is the exact reasoning the scenario is constructed to challenge.
  • A more precise aggregate is still an aggregate; precision does not reveal structure.

9. Questions transversales et multi-réponses (Q101–Q112)

Ce que couvre cette section. Les concepts qui traversent plusieurs scénarios — stratégies de décomposition, découpage d'outils, reprise après crash, rendu de synthèse — puis quatre items multi-réponses. À l'examen, chaque item indique combien de réponses sélectionner ; ici c'est écrit dans l'intitulé.

Q101 Adding tests to a codebase you do not know

Scenario: The task is "add comprehensive tests to this legacy codebase". You do not know how many modules exist, which are untested, or how they depend on one another. Each thing you discover changes what should be done next.

Question: Which decomposition strategy fits?

A) A fixed sequential pipeline: enumerate the modules, write tests for each in declaration order, then run the whole suite.

B) Dynamic adaptive decomposition: map the structure, then adapt the plan as dependencies surface.

C) A two-pass review: a per-file analysis of every module, followed by a cross-file integration pass over the codebase.

D) One exhaustive prompt describing every module that should be tested, in what order, and to which level of coverage.


Answer: Dynamic adaptive decomposition

Why: open-ended investigation, where the subtasks depend on what earlier steps reveal, calls for adaptive decomposition rather than a fixed plan. Mapping structure, identifying high-impact areas and adapting priorities as dependencies emerge is the guide's own worked example of this task (Task 1.6).

Why the others are wrong:

  • A fixed pipeline requires knowing the module list and the right order up front — precisely what is unknown.
  • Multi-pass review is the pattern for a bounded, known set of files, such as a pull request.
  • An exhaustive prompt presupposes the knowledge the task is meant to produce.
Q102 A review with known aspects

Scenario: Every release, you review the same eight service modules against the same four concerns: security, error handling, logging, and API contract compliance. The scope is stable and known in advance. A single combined prompt produces uneven depth.

Question: Which decomposition fits?

A) One prompt per release, with firmer instructions to cover all four concerns to the same depth.

B) Delegate the whole review to a single subagent with a large context window, so every module and concern is weighed together.

C) Prompt chaining: a fixed sequence of focused passes, one per aspect or per module, then an integration pass.

D) Dynamic adaptive decomposition, letting each pass decide what the next should examine as issues surface.


Answer: Prompt chaining

Why: prompt chaining suits predictable multi-aspect reviews where the steps are known in advance; dynamic decomposition suits open-ended investigation. The distinguishing feature here is that the scope is stable release after release (Task 1.6).

Why the others are wrong:

  • Firmer wording does not change the approach that is already producing the uneven depth you are trying to fix.
  • Seeing every module together is what the combined prompt already does; moving it into a subagent leaves the pass undifferentiated.
  • Adaptivity buys nothing when the aspects and modules are fixed, and it makes results less comparable between releases.
Q103 One tool that does four things

Scenario: Your analyze_document tool is described as "Analyzes a document and returns insights." Callers use it to pull specific data points, to produce summaries, and to check whether a claim is supported by a source. Its output shape varies by call, and downstream code cannot rely on it.

Question: What is the appropriate redesign?

A) Split it into purpose-specific tools — extract_data_points, summarize_content, verify_claim_against_source — each with a defined input/output contract.

B) Keep one tool and have downstream code branch on the shape of the response.

C) Keep one tool and add a mode parameter with three allowed values.

D) Keep one tool and expand its description to explain the three uses.


Answer: Split into purpose-specific tools

Why: splitting a generic tool into purpose-specific tools with defined input/output contracts is the documented remedy. Each resulting tool has one job, one output shape, and a description that can actually differentiate it (Task 2.1).

Why the others are wrong:

  • Makes every consumer responsible for guessing what came back.
  • A mode parameter hides the three-way decision inside an argument, where the tool description can no longer guide it and the output shape still varies.
  • A longer description does not give downstream code a stable contract.
Q104 Two tools nobody can tell apart

Scenario: You have analyze_content ("Analyzes content and extracts key information") and analyze_document ("Analyzes documents and extracts information"). In practice the first is used for web search results and the second for PDFs, but the agent picks between them essentially at random.

Question: What is the most effective fix?

A) Add a system prompt rule stating which tool suits which kind of payload, and trust the agent to obey it.

B) Merge the two into a single tool that accepts both kinds of payload, removing the selection decision altogether.

C) Remove analyze_content, sending web-search results through analyze_document instead.

D) Rename analyze_content and recast its description for web results.


Answer: Rename and rewrite the description to eliminate the overlap

Why: near-identical descriptions cause misrouting, and the fix is to remove the functional overlap — renaming the tool to something like extract_web_results and rewriting its description around the web-specific inputs and outputs it actually handles. The guide uses this exact pair as its example (Task 2.1).

Why the others are wrong:

  • A prompt rule patches the selection and rests on the agent obeying it, while leaving two indistinguishable descriptions in place for it to choose between.
  • Removing the choice is a real option, but merging discards a genuine distinction between two kinds of payload that need different handling.
  • One tool is unambiguous only because every web-search result is now forced through a tool designed for documents.
Q105 Recovering a multi-agent run after a crash

Scenario: A six-hour multi-agent analysis crashes at hour five. Two agents had completed, one was mid-run, three had not started. Restarting from scratch means six more hours.

Question: What should the system have been designed to do?

A) Have the coordinator keep the full run in its conversation history so it can be resumed from the transcript.

B) Run every agent twice in parallel so a crash in one copy leaves the other still running to completion.

C) Checkpoint the whole process to disk every fifteen minutes and restore the last checkpoint after a crash.

D) Have each agent export its state to a known location and have the coordinator maintain a manifest it loads on resume.


Answer: Per-agent state exports plus a coordinator manifest

Why: structured state persistence for crash recovery means each agent exports its state to a known location and the coordinator loads a manifest on resume, injecting the recovered state where it is needed (Task 5.4). Completed work is then skipped rather than repeated.

Why the others are wrong:

  • The coordinator's own context is lost in the crash, and it is the wrong place to hold six hours of subagent output anyway.
  • Doubles the cost of every run to insure against an event that structured state handles for free.
  • Process-level checkpointing does not capture agent-level semantics — which questions were answered and which remain.
Q106 The context fills with discovery output

Scenario: Two hours into an exploration session, the context is dominated by verbose file listings and search results from the discovery phase. The conclusions you care about are a handful of paragraphs. You need to continue in the same session.

Question: What is the appropriate tool?

A) /compact, having first captured the conclusions you must not lose in a scratchpad file.

B) Ask the model to forget the discovery output it no longer needs, freeing the space.

C) Restart the session with only the conclusions, since compaction always loses too much.

D) /compact alone, since it preserves the important information automatically and the session continues.


Answer: /compact, with the critical conclusions persisted first

Why: /compact is the documented way to reduce context usage during extended exploration sessions filled with verbose discovery output. Because compaction is lossy, the durable findings belong in a scratchpad first (Task 5.4).

Why the others are wrong:

  • Asking the model to forget removes nothing: the tokens are already in the request. Taking something out of the context takes a real mechanism — compaction, API-side context editing, or a fresh session — never an instruction.
  • Restarting throws away a session you were told to continue; compaction plus a scratchpad keeps it alive and the findings intact.
  • Compaction summarises; it does not guarantee that any particular detail survives.
Q107 A downstream agent with no room

Scenario: Your synthesis agent receives, from three upstream agents, their full reasoning chains and the complete text of every source they consulted. It exhausts its context before finishing and its output degrades in the second half.

Question: What should change upstream?

A) Have the upstream agents return structured data — key facts, citations, relevance scores — instead of raw content.

B) Give the synthesis agent a model with a larger context window, so everything upstream sends still fits.

C) Have the synthesis agent summarise each upstream contribution before incorporating it, condensing the material where it lands.

D) Run the synthesis agent three times, once per upstream agent, and concatenate the three partial write-ups into one report.


Answer: Upstream agents return structured data

Why: when downstream agents have limited context budgets, the fix belongs upstream: return key facts, citations and relevance scores rather than reasoning chains and raw content (Task 5.1). The reasoning was useful to the agent that produced it; it is not what synthesis consumes.

Why the others are wrong:

  • Making everything fit buys headroom and leaves the ratio of signal to noise exactly where it was.
  • Makes the synthesis agent pay to read everything before condensing it — the cost you are trying to avoid.
  • Three partial syntheses concatenated is not a synthesis; cross-source connections are exactly what is lost.
Q108 An accurate report nobody can use

Scenario: Your synthesis output is factually correct and properly sourced, but readers complain it is unusable. Twelve quarters of revenue figures are described in flowing sentences, a sequence of regulatory events is presented as a bulleted list of fragments, and technical findings are buried in narrative paragraphs.

Question: Where is the defect and what is the fix?

A) In the sourcing: readers are struggling because they cannot trace a claim back to its source quickly enough.

B) In the rendering: each content type should be rendered in its natural form.

C) In the upstream agents, which are returning content in a form the synthesis agent cannot present well.

D) In the volume: the report covers too much and should be split into three shorter reports.


Answer: A rendering defect at the synthesis step

Why: when the content is correct and well sourced but hard to use, the fault is uniform rendering on output — usually everything flattened into prose, because that is the default. Different content types have different natural forms — tables for figures, prose for events and context, structured lists for technical findings — and the synthesis agent should be instructed accordingly (Task 5.6).

Why the others are wrong:

  • Tracing claims is not the complaint: the sourcing is described as sound, and readers are stuck on the presentation.
  • The scenario states the content is correct and sourced; the upstream agents did their job.
  • Splitting an unreadable report produces three unreadable reports.
Q109 Reducing human review safely (Select the 2 correct answers.)

Scenario: Your extraction system currently sends every document to human review. You want to reduce that load without letting error rates rise unnoticed.

Question: Which two measures belong in the design?

A) Set the review threshold at the aggregate accuracy figure already reported by the pipeline.

B) Have the model output field-level confidence scores and calibrate the review threshold against a labelled validation set.

C) Review only the document types that historically produced the most complaints.

D) Maintain ongoing stratified random sampling of high-confidence extractions to measure real error rates and detect novel error patterns.

E) Have the model self-rate each extraction from 1 to 10 and review anything below 8.


Answers: calibrated field-level confidence, plus ongoing stratified sampling

Why: the two halves of Task 5.5. Field-level confidence calibrated on labelled data tells you where to send the reviewers you have; stratified sampling of the extractions you stopped reviewing tells you whether the automation is still safe, and surfaces error patterns nobody anticipated. One routes, the other monitors — dropping either leaves the design incomplete.

Why the others are wrong:

  • An aggregate figure is exactly what conceals a badly performing segment; it is not a threshold.
  • Complaint history is a lagging and biased signal; the errors nobody complained about are the ones you most need to find.
  • Uncalibrated self-rating on an arbitrary scale — the model is confident where it is wrong.
Q110 Reliable escalation triggers (Select the 2 correct answers.)

Scenario: You are writing the escalation section of a support agent's system prompt and must choose which signals genuinely warrant handing off to a human.

Question: Which two are reliable escalation triggers?

A) The customer's request falls into a gap where policy is silent or ambiguous.

B) The conversation has exceeded fifty turns and the context window is filling up.

C) The customer's message registers strong negative sentiment.

D) The agent has made no meaningful progress after a reasonable number of attempts.

E) The agent's self-reported confidence for the case falls below 7 out of 10.


Answers: the policy gap, and the absence of meaningful progress

Why: the reliable triggers are an explicit request for a human, a policy exception or gap, and an inability to make meaningful progress. The policy gap and the stalled case are two of those three (Task 5.2).

Why the others are wrong:

  • A filling context window is a technical constraint with technical remedies (case facts, compaction), not a business reason to involve a human.
  • Sentiment does not correlate with case complexity — a furious customer may have a trivial problem.
  • Self-reported confidence is poorly calibrated, and is high precisely on the hard cases.
Q111 Passing context to a subagent (Select the 2 correct answers.)

Scenario: Your coordinator delegates to a synthesis subagent that must combine the outputs of a web search agent and a document analysis agent, while preserving source attribution.

Question: Which two practices are correct?

A) Give the subagent read access to the other subagents' sessions so it can consult them as needed.

B) Pass the coordinator's entire transcript, so nothing the subagent might need can be missing.

C) Include the complete findings from the prior agents directly in the subagent's prompt.

D) Rely on the subagent inheriting the coordinator's conversation history automatically, as a spawned process would.

E) Use a structured format that separates content from metadata — source URLs, document names, page numbers.


Answers: the prior findings passed explicitly in the prompt, in a content/metadata structure

Why: subagents have isolated context, so everything they need is provided explicitly in the prompt; and separating content from metadata is what preserves attribution through the handoff (Task 1.3).

Why the others are wrong:

  • Cross-session access breaks both the isolation of subagents and the hub-and-spoke routing through the coordinator.
  • The full transcript floods the subagent's context with material irrelevant to its task; "complete findings" is not "everything that happened".
  • A subagent is not a spawned process that inherits its parent's state: there is no automatic inheritance, and this is the single most common misconception about subagents.
Q112 Guarantees versus best effort (Select the 2 correct answers.)

Scenario: Two requirements must hold every single time: identity must be verified before any refund is processed, and no refund above $500 may be issued by the agent.

Question: Which two mechanisms provide a deterministic guarantee?

A) Few-shot examples showing the agent verifying identity first and escalating a $750 refund.

B) A programmatic prerequisite that blocks process_refund until get_customer has returned a verified customer ID.

C) A hook that intercepts the outgoing refund call, blocks amounts above $500, and redirects to escalation.

D) A system prompt stating that verification is mandatory and that refunds above $500 must be escalated.

E) A nightly audit that flags any refund issued without a prior verification call.


Answers: the prerequisite gate, and the tool-call interception hook

Why: both are enforcement in code, evaluated on the actual call and its actual arguments. Prerequisite gates and tool-call interception hooks are the two mechanisms the guide names for deterministic compliance (Tasks 1.4 and 1.5).

Why the others are wrong:

  • Demonstrations shape behaviour probabilistically: an agent that has been shown the right sequence can still depart from it, and no example is evaluated against the arguments of the call actually being made.
  • An instruction shapes behaviour probabilistically too. It lowers the failure rate without making it zero, and "every single time" is the stated requirement.
  • Detects violations after the money has moved. Detection is not prevention.

10. Couverture — les 30 task statements

Chaque task statement du guide est couvert par au moins un item. Utilisez ce tableau pour cibler une révision : si un domaine ressort faible sur votre rapport de score, reprenez d'abord ses questions.

Task Sujet Questions
1.1 Boucle agentique, stop_reason, historique Q13, Q14
1.2 Coordinateur-subagents, hub-and-spoke, décomposition, raffinement itératif Q7, Q48, Q49, Q50, Q51
1.3 Outil Task, passage de contexte explicite, parallélisme Q43, Q44, Q45, Q46, Q47, Q111
1.4 Prérequis programmatiques, handoff structuré, requêtes multi-sujets Q1, Q25, Q26, Q112
1.5 Hooks : normalisation PostToolUse, interception d'appels Q15, Q16, Q112
1.6 Prompt chaining contre décomposition dynamique Q12, Q101, Q102
1.7 Sessions : --resume, fork_session, reprise ou session neuve Q52, Q71, Q72
2.1 Descriptions d'outils, découpage, renommage, system prompt keyword-sensitive Q2, Q27, Q103, Q104
2.2 Erreurs structurées MCP, catégories, isRetryable Q17, Q18, Q53
2.3 Distribution d'outils, tool_choice, outils contraints Q9, Q28, Q69, Q70, Q91
2.4 Serveurs MCP : portées, ${VAR}, resources, community contre custom Q64, Q65, Q66, Q67, Q68
2.5 Outils built-in : Grep, Glob, Read/Write/Edit Q59, Q60, Q61, Q62, Q63
3.1 Hiérarchie CLAUDE.md, @import, /memory, .claude/rules/ Q29, Q30, Q31, Q32
3.2 Slash commands et skills : portées, context: fork, allowed-tools, argument-hint Q4, Q33, Q34, Q35, Q36, Q37
3.3 Règles par chemin, patterns glob Q6, Q32
3.4 Plan mode contre exécution directe, subagent Explore Q5, Q38, Q39
3.5 Raffinement itératif : exemples, TDD, pattern interview, groupement des correctifs Q40, Q41, Q42, Q85
3.6 CI/CD : -p, --output-format json, contexte CLAUDE.md, isolation de session Q10, Q73, Q74, Q75, Q76, Q77
4.1 Critères explicites, faux positifs, sévérité Q78, Q79, Q80
4.2 Few-shot : format, cas ambigus, généralisation, variété structurelle Q81, Q82, Q98
4.3 tool_use + JSON Schema, nullable, enums, normalisation de format Q87, Q88, Q89, Q90, Q91, Q92
4.4 Validation, retry avec feedback, limites du retry, detected_pattern Q83, Q93, Q94, Q95
4.5 Message Batches API : pertinence, custom_id, calcul de SLA Q11, Q86, Q96, Q97
4.6 Revue indépendante, multi-passes, routage par confiance Q12, Q75, Q84
5.1 Case facts, trimming des tool results, lost-in-the-middle, données structurées en amont Q19, Q20, Q99, Q107
5.2 Escalation : déclencheurs fiables et non fiables, correspondances multiples Q3, Q21, Q22, Q23, Q24, Q110
5.3 Propagation d'erreurs, échec d'accès contre résultat vide, coverage annotations Q8, Q54, Q55
5.4 Grandes codebases : scratchpad, /compact, manifeste de reprise Q105, Q106
5.5 Revue humaine, échantillonnage stratifié, confiance par champ calibrée Q100, Q109
5.6 Provenance, conflits de sources, dates, rendu par type de contenu Q56, Q57, Q58, Q108

CCA Révision — Practice Questions · 112 items · Source : guide officiel CCAR-F v1.0 · 2026

↑