FR

D3 — Claude Code configuration and workflows (20 %)

Twenty per cent of the score, six task statements, and a subject that looks deceptively soft: files in a repository. One section per task statement, in the guide's own order and under the guide's own wording, so that a question spotted on exam day maps to a section here without translation.

Table of contents
  1. 3.1 — Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization
  2. 3.2 — Create and configure custom slash commands and skills
  3. 3.3 — Apply path-specific rules for conditional convention loading
  4. 3.4 — Determine when to use plan mode vs direct execution
  5. 3.5 — Apply iterative refinement techniques for progressive improvement
  6. 3.6 — Integrate Claude Code into CI/CD pipelines
What this course owns, what it defers elsewhere

D0 and D1 are assumed. D1 taught the machinery inside an agent: the loop, subagents, hooks as SDK callbacks, sessions. D3 is the layer above it: the files and flags that configure a running Claude Code, and the working habits that decide whether a session goes well.

The guide places neighbouring notions elsewhere, and this course honours that. The hook mechanism — events, permissionDecision, updatedInput — is tested in TS 1.5; subagent context isolation in TS 1.2 and 1.3; sessions, --resume and fork_session in TS 1.7; the context window in D0, and its management in D5 — /compact included, which the guide lists among the skills of TS 5.4, so the pointer is D5.4. Each of those gets a sentence and a pointer here, never a second development. What D1 explicitly deferred to this course — subagent definitions under .claude/agents/, and hooks as a place to put a rule rather than as an API — lands in 3.2.

1. 3.1 — Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization

Open Claude Code in a repository it has never seen and it starts from nothing: it must rediscover the build command, the test runner, where handlers live, what was tried last month. CLAUDE.md is the file that stops the rediscovery. It is Markdown, Claude Code reads it at the start of every session, and its contents are appended to your prompt.

Which file, though. The guide tests three levels, and they differ on exactly one axis: who receives them.

Level Location Shared with
User ~/.claude/CLAUDE.md Just you, across all your projects — never travels through version control
Project ./CLAUDE.md or ./.claude/CLAUDE.md The whole team, through the repository
Directory CLAUDE.md in a subdirectory The whole team; loaded when Claude works in that subdirectory

That third column is the entire task statement. A user-level file lives in a home directory. Committing the repository does not carry it; cloning the repository does not produce it. It reaches a second machine only if something copies it there.

The signature of a hierarchy fault is "it works for everyone except the newest person", or except the machine rebuilt from an image, or except CI. An instruction the veterans apply and the last arrival never received is sitting at user level. The fix is to move it to the project level and commit it.

The distractors name real things and explain nothing. A different Claude Code version would not remove one targeted instruction. A CLAUDE.md that has grown too long degrades adherence for the whole team, not for one workstation. Having everyone copy the file by hand reproduces manually what committing already does, and diverges again at the first edit.

Keeping the file modular

Two mechanisms, and the guide names both.

@import. A path preceded by @, anywhere in the file, pulls another file in. Relative paths resolve against the file containing the import, not the working directory, and imports may nest to a maximum depth of four hops.

# packages/api/CLAUDE.md

This package owns the public REST surface.

@../../standards/http-errors.md
@../../standards/database-migrations.md

This is the guide's own skill: a monorepo maintainer imports into each package's CLAUDE.md the standards that package actually needs, instead of loading every standard everywhere.

.claude/rules/. Instead of one growing file, one file per topic:

.claude/
├── CLAUDE.md            # what is true of the whole project
└── rules/
    ├── testing.md
    ├── api-conventions.md
    └── deployment.md

Rules without path scoping load at launch, with the same standing as .claude/CLAUDE.md. Adding a paths field to a rule makes it conditional. That is TS 3.3, below, and it is the only thing that actually shrinks what gets loaded.

Reaching for @import to reduce context. Imported files are expanded inline at launch, right where you referenced them. Imports buy you organization, not a smaller context: every byte still loads. A question that asks how to stop irrelevant conventions from consuming the window is asking for path-scoped rules (3.3), and @import is the distractor placed next to them.

Why the file stops working

CLAUDE.md is guidance, not enforced configuration. Every line competes with every other line for attention, so the longer the file grows the less reliably any single rule is followed. That is not a defect to work around; it is the property that dictates how you write.

Three habits follow from it, and they are worth knowing because they are also the shape of the wrong answers:

And the rule that does not belong in the file at all: anything that must never happen. "Never push to main" written in CLAUDE.md is a hope. A hook is code that runs at a fixed point and can refuse the action. The mechanism is D1's TS 1.5; where a given rule belongs is 3.2, below.

/memory is the diagnostic command of this task statement: it lists your memory file locations across user and project scope and lets you open them, which is how you check whether the instruction you are looking for is where you think it is.

The guide is the law here. Its skill reads "using the /memory command to verify which memory files are loaded and diagnose inconsistent behavior across sessions", and the answer is /memory. Field nuance, so you are not surprised at a terminal: the current documentation splits the job, with /memory listing and opening the files and /context reporting which ones actually loaded into the running session.

Test yourself on this section

Q1 A rule the cloud workspaces never receive

Scenario: Half the team codes in ephemeral cloud workspaces rebuilt from a base image at every checkout. Payment calls written there go straight to the gateway, while laptops keep routing through PaymentsClient as agreed. The instruction ships in the onboarding dotfile bundle, which every new hire unpacks into ~/.claude/CLAUDE.md on their laptop; the base image is built from another recipe and never runs that step.

Question: What is the diagnosis and the fix?

A) The workspaces open their session below the repository root, so the file one level up never loads; start at the root.

B) The wording reads as a preference rather than an obligation, and a fresh environment has no local habit to fall back on; make it imperative.

C) It sits at user scope, outside version control; commit it at project scope instead.

D) The repository never declares it; move the restriction into settings.json.


Answer: the instruction lives at user scope instead of project scope

Why: a user-scope file belongs to a home directory and is never distributed with the repository, so it reaches a machine only through whatever copies it there and an environment rebuilt from an image starts without it. Anything the whole team must follow belongs in a committed project file.

Why the others are wrong:

  • Nothing is waiting further up the tree: the file is not in the image at all.
  • Wording cannot decide the behavior of a session that never read the text.
  • That file governs permissions, hooks and tool access; it is not where standing instructions to the model live.

All bank questions on 3.1

2. 3.2 — Create and configure custom slash commands and skills

When you have typed the same multi-step instruction twice, it should stop being something you type. The guide keeps two mechanisms for that, and treats them as distinct.

Slash command Skill
Lives in .claude/commands/review.md .claude/skills/review/SKILL.md
Invoked Explicitly, /review Explicitly, or automatically when Claude matches your request to its description
Carries A prompt A folder: instructions, reference files, scripts
Configurable — Frontmatter: context, allowed-tools, argument-hint, and more

Both come in two scopes, and the split is the same one as 3.1:

Scope Commands Skills Shared with
Project .claude/commands/ .claude/skills/ The team, through version control
Personal ~/.claude/commands/ ~/.claude/skills/ Just you

As soon as the scenario says a command must be available to everyone after a clone, the guide's answer is .claude/commands/. Three wrong answers are available to be offered: ~/.claude/commands/, which produces a working command that no colleague ever sees; CLAUDE.md, which carries standing instructions and not invocable definitions; and a commands array declared in some .claude/config.json, which sounds like tool configuration and names nothing that exists.

The frontmatter the guide tests

A skill is a directory holding a SKILL.md: YAML frontmatter, then the procedure.

---
name: dependency-audit
description: Audits third-party dependencies for unused, outdated and duplicated
  packages. Use when reviewing dependencies or before a release.
context: fork
allowed-tools: Read, Grep, Glob, Bash
argument-hint: "[package-manager] — npm, pnpm or cargo"
---

1. Read the lockfile and list every direct dependency.
2. Grep the source for each one and flag those with no import.
3. Report unused, outdated and duplicated packages in three tables.

The description is the field everything hinges on: Claude loads only names and descriptions at startup and matches your request against them semantically. A skill that never fires almost always has a description that does not overlap the way you actually phrase requests.

Three optional fields carry the task statement.

context: fork runs the skill in an isolated subagent context. The skill's own content becomes the prompt driving that subagent; it does not inherit your conversation, and what comes back to the main session is the result rather than the whole trace. That is the answer whenever a skill's output volume is the problem: codebase-wide analysis that prints every path it inspected, brainstorming that explores five options you did not keep.

allowed-tools restricts tool access during skill execution, for example limiting a skill to file-write operations so it cannot do anything destructive. That is the guide's wording and the one to answer with. Field nuance, for your own machine: the current documentation describes the field as pre-approving the listed tools for the invoking turn, so Claude may use them without asking, rather than removing the others; the field that actually removes a tool from the pool is disallowed-tools.

argument-hint is the hint shown at autocomplete time, so a developer who invokes the skill without arguments is told which parameter it expects.

A skill runs, produces a wall of output, and the session that follows has visibly lost the thread. The answer is context: fork. Two neighbouring options are plausible and neither moves the output. allowed-tools changes what the skill may do, not where what it writes lands. And a "be concise" line added to the skill body asks the model to write less rather than to write elsewhere. Reduced or not, the output still arrives in the main session.

Personal variants, and why the name matters

Skills with the same name are resolved by source, and a personal skill wins over a project skill. Clone a repository that ships .claude/skills/review/, keep your own ~/.claude/skills/review/, and /review runs yours, silently, on your machine only.

The guide's instruction is to give a personal variant a different name: ~/.claude/skills/review-strict/ rather than a second review. Reusing the team's name does not modify their configuration, but it does mean you and your colleagues run different procedures under one command and nothing in either setup says so. Renaming is also the usual fix when a skill of yours is being shadowed and you cannot work out why it never fires.

Which surface owns which rule

The guide's last skill on this task statement is choosing between a skill and CLAUDE.md. The full picture is worth holding, because the wrong answers on this topic are always another real mechanism:

Surface Loads Fires on Use it for
CLAUDE.md Every session, always Nothing — it is simply present Universal standards: naming, layout, "always do X"
Skill On demand Explicit /name, or a semantic match on its description Task-specific procedures and reference material
Slash command On invocation Explicit /name A repeatable prompt
Hook — A lifecycle event A rule Claude must not be able to skip
Subagent On delegation An explicit hand-off Work you want done in an isolated context

The line between the first two is the guide's: always-loaded universal standards versus on-demand task-specific workflows. Detailed procedures placed in CLAUDE.md pay their token cost in every conversation, including the ones they have nothing to do with.

The line to the hook is different in kind. CLAUDE.md and skills are instructions Claude follows; a hook is code that runs. If skipping the rule is unacceptable, it does not belong to instruction-following at all.

The events, and PreToolUse as the enforcement primitive, are D1's TS 1.5 and are not re-taught here. What differs on this side of the line is how a hook says no. A Claude Code hook is a command declared in settings, so it answers with an exit code: 2 blocks, and what it wrote on standard error is fed back to Claude as the reason. An Agent SDK hook is a callback in your process and answers with a JSON permissionDecision. Same event, same intent, two conventions. And exit 1 is the trap on the Claude Code side: it looks like failure and blocks nothing.

# .claude/hooks/protect-main.sh — refuse a push to main, every time.
# A PreToolUse hook receives the tool name and its input as JSON on stdin.
if grep -qE 'git +push.*\bmain\b'; then
  echo "Pushing to main is not allowed; open a pull request instead." >&2
  exit 2          # 2 blocks and explains. 0 allows. 1 does neither.
fi

Subagents are the other neighbour. They are defined as Markdown files with YAML frontmatter under .claude/agents/, and /agents creates one interactively. Two facts belong to this task statement rather than to D1: a subagent does not inherit your skills, and a custom subagent loads only the skills its skills frontmatter field names.

---
name: accessibility-reviewer
description: Reviews frontend changes for accessibility regressions.
tools: Read, Grep, Glob
skills: accessibility-audit, contrast-check
---

Review the diff for WCAG violations. Report findings; never edit files.

The rest — why a subagent's isolated context is the point, and what must travel in the prompt because of it — is D1's TS 1.2 and 1.3.

Putting a long procedure in CLAUDE.md so it is "always available". It is always loaded, which is not the same benefit and is a permanent cost. A PR review checklist has no business in context while you are debugging.

Reaching for a hook to package a workflow. A hook reacts to a tool event; it is not something a developer chooses to run. If the scenario says a person invokes the thing, a hook is not the answer.

Test yourself on this section

Q3 A skill that floods the window it runs in

Scenario: Your /orphan-assets skill walks the media bucket and the codebase to list images nothing references. It prints every path it checks on the way, and the work that follows in the same session has visibly lost the thread.

Question: What is the right change?

A) Restrict the skill's allowed-tools to Read and Glob, so the run produces less to read.

B) Have the skill append each inspected path to reports/orphans.log and print only the total it found.

C) Run it in a second terminal and paste the verdict back.

D) Declare context: fork in the skill's frontmatter, so only the verdict returns.


Answer: run the skill in a forked context

Why: the path-by-path trace is worth producing, it simply does not belong in the conversation. A forked context runs the walk inside an isolated subagent and hands back the conclusion alone.

Why the others are wrong:

  • Tool permissions decide what the skill may do, not how much it narrates.
  • The listing still crosses the session on its way to disk.
  • Isolation by hand, unrepeatable, and nothing carries it to the next person.
Q6 Packaging a workflow the whole team reuses (Select the 2 correct answers.)

Scenario: Your team runs the same accessibility audit over and over. It has to reach everyone, keep its verbose trace out of the working session, and run both at a developer's terminal and from the pipeline.

Question: Which two moves deliver that?

A) Commit it as .claude/skills/a11y-audit/SKILL.md with context: fork in its frontmatter.

B) Keep it as a personal ~/.claude/skills/a11y-audit/SKILL.md so you can iterate without disturbing anyone.

C) Invoke it by name at the terminal, and headless from the pipeline through claude -p.

D) Paste the audit steps into ~/.claude/CLAUDE.md so every session carries them.

E) Fire it from a PostToolUse hook after each write, so nobody has to remember it.


Answers: commit the forked skill with the repository, and invoke it by name locally or headless from the pipeline

Why: a committed skill meets all three requirements at once — versioned, so everyone gets it at clone time; forked, so its trace stays out of the session; and reachable from both sides, by name at the terminal and headless in the pipeline, where --output-format json hands the next step something it can act on.

Why the others are wrong:

  • User scope is private and outside version control, which is the one requirement it cannot meet.
  • User scope again, and that file carries standing instructions rather than an invocable workflow.
  • A hook reacts to a tool event; it does not package a workflow somebody chooses to run.

All bank questions on 3.2

3. 3.3 — Apply path-specific rules for conditional convention loading

Splitting CLAUDE.md into .claude/rules/ files makes the instructions maintainable, and loads exactly as many tokens as before: that is the problem 3.1 left open. Adding a paths field to a rule is what makes the loading conditional.

---
paths:
  - "**/*.test.ts"
  - "**/*.test.tsx"
---

# Test conventions

- One `describe` per exported symbol, `it` sentences in the present tense.
- Build fixtures with the factories in `test/factories`; never inline literals.
- Never mock the database — use the test database from `test/db.ts`.

paths takes a list of glob patterns. The rule's body enters context only when Claude is working with a file that matches one of them; the rest of the time it costs nothing. A rule file with no paths field loads unconditionally, like the project CLAUDE.md.

The globs are ordinary:

Pattern Matches
**/*.ts Every TypeScript file, any directory
src/**/* Everything under src/
*.md Markdown files at the project root only
terraform/**/* Everything under terraform/

Why not a CLAUDE.md in the subdirectory

Both mechanisms scope instructions. They scope them by different things, and the guide tests the difference.

A subdirectory CLAUDE.md is attached to a place. It is the right answer when the convention really is local: everything under terraform/ obeys these tagging rules, and nothing outside does.

A path-scoped rule is attached to a pattern. It is the right answer when the files sharing a convention do not share a directory. The canonical case is test files, colocated with the code they cover and therefore scattered through the whole tree. Writing a CLAUDE.md in each of forty directories to say the same three things about tests is unmaintainable, and a single root CLAUDE.md saying them costs every session that never opens a test.

Read the scenario for one word: where do the files live?

"Test files colocated throughout the codebase", "Terraform files under infra/<product>/ in a dozen folders", "migration scripts spread across services" → path-scoped rule with a glob. "Everything in this one package" → a directory CLAUDE.md is defensible.

Three wrong answers are available here. A CLAUDE.md per subdirectory, seductive because it sounds more local, but bound to the folder and unable to follow a file type. Consolidating into the root CLAUDE.md under section headings, which replaces an explicit match with the model's inference and loads everything anyway. And a skill per file type, which fails on the word automatically in the scenario: a skill is invoked or matched against a request, not against the file being edited.

The guide is the law here. It writes that path-scoped rules load "when editing matching files". If an option turns on the word editing, that is the one. Field nuance, for your own machine: the current documentation says path-scoped rules trigger when Claude reads a file matching the pattern, so in practice the rule is already in context before your first edit. Same mechanism, same paths, same globs; only the trigger differs.

Using path scoping to enforce something. A rule that loads is still a rule Claude reads, with the adherence properties of any instruction. "Every Terraform resource must carry a cost_center tag" as a path-scoped rule shapes the writing; it does not guarantee the outcome. If the requirement is a gate, it is a hook or a CI check. The two are complements, not alternatives: the rule gets it right most of the time, the gate catches the rest.

Test yourself on this section

Q2 Standards for files that sit in every product folder

Scenario: Every product team keeps its own infrastructure under infra/<product>/, so the Terraform files matching infra/**/*.tf are scattered across a dozen folders. Each resource must pin its provider version and carry a cost_center tag, and you want that applied whenever one of those files is edited, without anyone invoking anything.

Question: Which setup delivers that?

A) A PostToolUse hook that fires after each write and appends the tagging standards to the session.

B) A .claude/rules/terraform.md whose frontmatter paths glob covers those files.

C) A tflint policy in CI that fails the build whenever a resource is untagged or its provider is unpinned.

D) The standards written up in docs/infrastructure-standards.md, the place engineers already look.


Answer: a path-scoped rule file whose glob follows the file type across the tree

Why: the glob matches on the file being edited, wherever it sits, and the body loads only while such a file is in play — which is what a convention attached to a file type rather than to a directory needs.

Why the others are wrong:

  • A hook runs a command on a tool event, once the write has already happened; it is not a channel for standards meant to shape the writing.
  • Rejection after the fact, not guidance during: the loop becomes a red build somebody has to go back and fix.
  • A document no session opens shapes nothing that gets written.

All bank questions on 3.3

4. 3.4 — Determine when to use plan mode vs direct execution

Plan mode is a permission mode in which Claude reads, searches and runs read-only commands, asks its clarifying questions, and returns a plan, without editing anything. Direct execution is the ordinary mode: you describe the change and Claude makes it.

# Enter it for a whole session…
claude --permission-mode plan "Restructure the billing monolith into services"

# …or press Shift+Tab to cycle into it mid-session, /plan for a single prompt.

The choice is not about caution, it is about where the decisions are.

The scenario says Mode Because
"Dozens of files", "45+ files" Plan The blast radius has to be mapped before it is created
"Several valid approaches", "different infrastructure requirements" Plan Something must be chosen, and choosing after writing means rewriting
"Service boundaries", "restructuring" Plan Architectural decisions propagate
"A clear stack trace", "one function" Direct The scope is already established
"Add a date validation check" Direct Understood change, nothing to arbitrate

Plan mode earns its keep for one reason: iterating on a plan is far cheaper than letting a run happen and cleaning up after it. The place to course-correct is before a line is written. And when the plan comes back, read it rather than skim it: an unread plan is an approval you did not actually give.

The Explore subagent

A subagent runs in its own context window and returns a summary; the reads, the searches and the dead ends stay with it. That isolation is D1's subject (TS 1.2 and 1.3). What belongs here is what the guide asks of it in this task statement.

Explore is the built-in read-only subagent for discovery. Claude delegates to it to search or understand a codebase without changing it, and only its summary reaches the main conversation. That is the answer whenever a multi-phase task is running out of window during its discovery phase: the exploration cost is paid in a context you are going to throw away, and the main window stays available for planning and then implementing.

It is not tied to plan mode. You can ask for exploration through the subagent outside plan mode when all you want is a summary of how something works.

Plan mode decides, Explore keeps the window clear, and the guide combines them. They answer two different shortages: plan mode when something has to be settled before any file changes, Explore when a discovery phase fills the window with reads you will never need again. Neither implies the other, and neither excludes the other: a multi-phase task that has to be scoped and is running out of room during discovery takes both, which is the combination the guide names. The distractors answer the context shortage alone, with compaction or a restart, and leave the decision unmade.

The combination the guide asks about

Plan mode and direct execution are not a permanent allegiance. The common shape is both, in order: plan the library migration, approve the plan, then execute it directly. The approved plan becomes the scope of the writing phase. The arbitration happened before any file changed.

The complexity is stated in the scenario; it is never something you infer. Count the signals: multiple files, competing approaches, boundaries to draw → plan mode. A stack trace, one file, a change everybody already understands → direct execution.

The subtlest distractor proposes starting in direct execution and switching to plan mode if it turns out to be complicated. It sounds prudent and it ignores that the scenario has already established the complexity: you would be discovering, at cost, something you were told. Two blunter ones: plan mode everywhere on principle, which spends deliberation on a one-line fix; and giving exhaustive instructions up front in direct execution, which assumes you know the right structure before anyone has read the code.

Treating plan mode as a safety feature and direct execution as the risky one. They answer a question about the task, not about your appetite for risk. When a scenario carries two tickets, answering plan mode for both costs you the settled one: a failing test that states the expected value has already written down the answer that planning would go looking for.

Test yourself on this section

Q7 A stale scale factor and an unsettled ingest path

Scenario: Two tickets open the same morning on a smart-metering platform. The first: a failing regression test pins toKilowattHours(), which still applies the scale the meters used before the current firmware shipped; the test states the value it expects, and the scale is a single constant, METER_SCALE. The second: retire the nightly polling job and have meters publish readings as they take them, which means settling where backpressure belongs, what becomes of readings buffered through a network outage, and how both paths run side by side while the fleet upgrades over months.

Question: Which mode fits which ticket?

A) Plan mode for both, since a change that reaches customer meters should be designed before it is written.

B) Plan mode for the scale factor so the firmware assumption gets recorded, and direct execution for the ingest change, revising the design as problems surface.

C) Direct execution for the scale factor, plan mode for the ingest change.

D) Direct execution for both, one meter model at a time.


Answer: run the scale factor directly, plan the ingest change

Why: deliberation pays where approaches compete, boundaries are unsettled, and one early decision propagates across many files. The ingest change carries three open questions before a line is written; the scale factor carries a failing test that already states the answer.

Why the others are wrong:

  • Spends deliberation on a ticket whose expected value is already written down in the test.
  • Reverses both readings at once, deliberating over the settled ticket and improvising the unsettled one.
  • Incremental delivery still commits to a backpressure design nobody has chosen, and a fleet halfway through an upgrade is the worst place to discover that.

All bank questions on 3.4

5. 3.5 — Apply iterative refinement techniques for progressive improvement

The four techniques in this task statement come from the guide's own list. They answer four different failures, and the exam question is always which failure is this.

Concrete input/output examples

The failure: the same prose instruction produces a different result each run. The remedy the guide calls most effective is not more prose. It is two or three concrete pairs showing the transformation.

Normalise these customer records.

  "Dupont,  Marie-Claire"  →  { first: "Marie-Claire", last: "Dupont" }
  "jean pierre MARTIN"     →  { first: "Jean Pierre", last: "Martin" }
  "O'Brien, Sean (Jr.)"    →  { first: "Sean", last: "O'Brien", suffix: "Jr." }

Three pairs settle what a paragraph of definition left open: where the comma goes, what happens to casing, that a suffix gets its own field. An example shows the result instead of describing it, and the boundary cases you choose to include are the ones you were worried about.

When the scenario says a written instruction was rewritten once or twice and the output still varies, the answer is worked examples. The distractors are all plausible next moves and all repeat the failure. Rewriting the prose more precisely is the approach that has already failed twice; more words on an ambiguity do not resolve it. Asking the model to explain its interpretation before each run turns every invocation into a negotiation instead of fixing the specification once. Splitting the transformation into one call per field multiplies invocations without ever defining the ambiguous term. And a validator that rejects bad output catches one shape of error and re-runs the same unconstrained request.

Test-driven iteration

The failure: you cannot tell whether the code is right, so neither can Claude. Write the test suite first — expected behaviour, edge cases, performance requirements — then have the implementation written against it, then iterate by handing back the failures.

$ pytest tests/test_migration.py
FAILED test_null_effective_date - AssertionError: expected None, got '1970-01-01'
FAILED test_duplicate_account_ids - KeyError: 'legacy_id'

Two conditions make this work rather than merely feel rigorous. The tests must be a reliable source of truth: a suite that passes on broken code turns iteration into false confidence. And "done" has to be the gates having run and been observed, not a summary that says they passed.

The guide's sharpest version of this is for edge cases. Rather than describing the problem, supply the specific case: this input, this expected output, for instance the null value your migration script mishandles.

The interview pattern

The failure: the specification is incomplete and nobody knows it yet. Ask Claude to interview you before implementing.

Before implementing the cache for the pricing API, ask me whatever you need to
settle the design. Don't write code until I've answered.

  1. Invalidation by TTL, or on write events?
  2. Is stale data acceptable while the cache is unreachable?
  3. Per-tenant or global?
  4. What read volume should this hold at?

Its value is question 2. You had not thought about the cache being unreachable, and you would have found out in production. Reach for it in unfamiliar domains, where the implications are not obvious — cache invalidation strategies, failure modes — and where several approaches are valid and the right one depends on context you have not stated.

One message, or one at a time

The failure: several problems at once, and the fixes fight each other.

The problems are Give feedback Because
Interacting — fixing one changes what the others need All in one detailed message The model has to see the interaction to reconcile it; fixing them one at a time means re-doing the first fix
Independent — each stands alone Sequentially You validate each step, and nothing you approve gets undone by the next round

Error handling that is missing and input validation that is wrong are one message: a unified answer has to place them together. Renaming a variable and changing a return type are two conversations.

Four techniques, four different failures. Output that varies between runs is an under-specified transformation, and the answer is worked pairs. Code nobody can judge is a missing gate, and the answer is tests written first. A design you cannot state is an incomplete specification, and the answer is letting Claude interview you. Fixes that undo each other are a feedback question, and it turns on interdependence, never on how many problems there are.

Batching independent problems to save round-trips. It is more token-efficient and it gives up the per-step validation that made sequential feedback worth choosing: you find out at the end which of the five changes was wrong.

Splitting interacting problems for control. Each fix is made without sight of the others, so the second one dismantles the first. The criterion is interdependence, never the number of problems.

Test yourself on this section

Q8 Redaction that lands differently every time

Scenario: You ask for support transcripts to be "stripped of anything personal" before they reach the analytics store. One pass masks the order number and the next leaves it; one pass swaps the customer's first name for a token and the next deletes the whole sentence. Two rewrites of the written instruction have not converged.

Question: What is the most effective next move?

A) Point the prompt at the company's data-classification policy and tell it to apply the categories defined there.

B) Add a check that rejects any output still containing a long digit sequence, and re-run the ones it rejects.

C) Have a second instance grade each redaction.

D) Supply two or three excerpts paired with their redacted form.


Answer: worked excerpt pairs, before and after

Why: an example settles what prose left open by showing the result rather than describing it again, and a handful of pairs cover far more of the boundary than another paragraph of definition.

Why the others are wrong:

  • A classification policy names categories; it never says what to do with a sentence that happens to contain one.
  • Catches one shape of leak and says nothing about the rest, and the re-run repeats the same ambiguous request.
  • The grader inherits the same undefined standard, so it flags by the same guesswork.

All bank questions on 3.5

6. 3.6 — Integrate Claude Code into CI/CD pipelines

Claude Code is interactive by default. Invoked in a pipeline with no one to answer a prompt, it waits, and the runner eventually kills it at the step timeout. Everything in this task statement follows from that.

-p, and the shape of the output

-p, spelled out --print, runs Claude Code as a one-shot command: it processes the prompt, writes the result to standard output, and exits. It reads stdin and writes stdout, so it pipes like any other shell tool.

That gets the job to finish. Getting something the next step can act on takes two more flags: --output-format json makes the reply a JSON envelope, and --json-schema constrains the payload to a schema you supply. The conforming object lands in the envelope's structured_output field, at a fixed place, so nothing downstream has to parse prose.

- name: Review the diff
  env:
    ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
  run: |
    claude -p "Review $(git diff origin/main...HEAD) for security issues and bugs." \
      --output-format json \
      --json-schema '{
        "type": "object",
        "properties": {
          "issues": {"type": "array", "items": {
            "type": "object",
            "properties": {
              "severity": {"enum": ["critical", "major", "minor"]},
              "file": {"type": "string"},
              "line": {"type": "integer"},
              "message": {"type": "string"}
            },
            "required": ["severity", "file", "message"]
          }}
        }
      }' \
      | jq '.structured_output.issues' > findings.json

findings.json is now a list a later step can post as inline pull-request comments through the GitHub API. The API key comes from the repository secrets, never from the workflow file.

What the CI instance knows

A CI-invoked Claude Code has no session history and nobody to ask. The guide's answer to that is CLAUDE.md: the committed project file is the mechanism that carries testing standards, fixture conventions and review criteria to the instance running in the pipeline. Document which fixtures exist and what makes a test worth writing, and test generation stops producing low-value tests that re-cover ground the suite already holds, which is also why you put the existing test files in context when you ask for new ones.

claude -p loads the same project context an interactive session would, CLAUDE.md included. Bare mode — the --bare flag — is the opt-out: it skips auto-discovery of hooks, skills, plugins, MCP servers and CLAUDE.md to start faster and run deterministically.

Be deliberate here, because a common misreading of this pair is wrong: it attributes that skipping to -p itself. The documentation attributes it to --bare and states that without it, claude -p loads the same context an interactive session would. The guide's own TS 3.6 depends on CLAUDE.md reaching the CI instance, which it could not do if -p dropped it. The pipeline gets your CLAUDE.md unless you asked for bare mode; and if you did ask for it, you also gave up the standards the guide is telling you to rely on.

Reviewing is a second instance's job

The session that wrote the code is the worst reviewer of it. It carries the reasoning that produced the code and questions its own decisions least, the same bias as an author proof-reading their own text. Review runs in an independent instance: a separate session, or a subagent with no memory of how the code was built. Fresh eyes catch what the original run talked itself past.

Two operational rules complete the picture, and both are guide skills:

Four associations settle most CI questions. A job hanging on input → -p. A pipeline that has to parse findings → --output-format json with --json-schema. A CI instance ignoring the project's testing conventions → the committed CLAUDE.md carries them. Reviewing code Claude just wrote → an independent instance.

The distractors come in two families. Invented mechanisms: a CLAUDE_HEADLESS environment variable, a --batch flag, a --no-tty option. They are plausible because headless and batch really are the vocabulary of this domain, and none of the three exists. An unrecognised option does not change the mode. Workarounds: redirecting stdin from /dev/null, which treats the symptom without turning on the documented mode; wrapping the call in timeout, which turns a long hang into a short one and still produces nothing; and regex-parsing the prose, which moves the fragility into your own code instead of removing it.

Re-running the review from scratch on every push. Claude has no memory between runs, so it re-reports everything it found last time. The pull request accumulates duplicates until nobody reads the bot. The fix is not to review less often; it is to include the prior findings and ask for the delta.

Letting the job that generated the code also review it, to save a run. One job costs less than two, and it buys the review the guide names as the least effective one: the reasoning that wrote the code is still in the window. Why an independent instance catches what a self-review misses, and how to split a large review into per-file and cross-file passes, is developed in D4.6.

Test yourself on this section

Q4 A nightly job the runner has to kill

Scenario: A nightly job runs claude "Summarize the changes merged today" and the runner kills it at the step timeout. The log shows Claude Code sitting at an interactive prompt.

Question: What makes the command usable in an unattended job?

A) Export CLAUDE_CI=1 in the job environment.

B) Add --no-tty to the invocation.

C) Add -p (spelled out as --print) to the invocation.

D) Wrap the call in timeout 600 so the runner is never the one to kill it.


Answer: the print flag, which is the documented non-interactive mode

Why: it processes the prompt, writes the result to stdout and exits without waiting on input — which is what an unattended job needs.

Why the others are wrong:

  • No such environment variable exists, so the command runs unchanged and still waits.
  • No such flag exists either, and an unrecognized option does not switch the mode.
  • Turns a long hang into a shorter one; the job still produces nothing.
Q5 The flag-retirement job cannot read its own input

Scenario: A weekly job runs claude -p to list the feature flags that are safe to remove, and retire_flags.mjs expects records carrying flag, owner and last_seen. Some runs come back as a bulleted list, some as a paragraph, and the field names move between runs, so the job fails more often than it succeeds.

Question: What is the correct fix?

A) Append "reply with a JSON array and nothing else" to the prompt, and retry the step whenever the parse fails.

B) Feed the first run's prose into a second claude -p call whose only job is to turn it into the record shape.

C) Run with --output-format json and --json-schema, then read structured_output.

D) Keep a run that parsed cleanly and paste it into the prompt as a worked example of the layout.


Answer: constrain the shape with the two structured-output flags

Why: the flags exist for exactly this — one makes the reply machine-readable, the other imposes the fields, and the conforming payload arrives in a fixed place. Nothing downstream has to guess.

Why the others are wrong:

  • Rests on the model complying, which is what already fails; a retry repeats the same unconstrained request.
  • The second call is as unconstrained as the first, and now two shapes can drift.
  • An example nudges the layout without guaranteeing it, and the field names are exactly what moves.

All bank questions on 3.6


CCA Revision — D3 — Claude Code configuration and workflows (20 %) · 2026

↑