These twenty questions come from the free CCAR-F Form 1 on this site, picked so that every exam domain appears. The real exam draws four scenarios from a published bank of six, so these questions are written inside those same settings.
Answer each one before you open it. Then read every explanation, not just the one for the right answer: knowing why the tempting option loses is what carries over to the questions you have not seen.
The resolution loop's dispatcher branches on response.stop_reason. One branch executes the tool Claude asked for and sends the result back for another iteration; the other stops and returns the reply to the customer. Which pair of literal values does the dispatcher branch on?
- A"function_call" to run the requested tool, "end_turn" to finish the turn
- B"tool_call" to run the requested tool, "complete" to finish the turn
- C"tool_result" to run the requested tool, "stop" to finish the turn
- D"tool_use" to run the requested tool, "end_turn" to finish the turn
Show the answer and why each option wins or loses
A
Pairs the right terminal value with the wrong iteration value. "function_call" is what other tool-calling APIs name this event, so the refund branch never fires and every tool request falls through to the reply.
B
Reads naturally and matches the vocabulary of several other vendors' loops, but neither string is emitted here, so the dispatcher matches nothing and the loop stalls on its first tool request.
C
"tool_result" is a real name in this API, for the block carrying a tool's output back in the next user message. It is not a stop_reason, and "stop" is not one either.
DCorrect
Correct. "tool_use" is the value Claude returns when it wants a tool executed and "end_turn" the value it returns when the turn is finished, which is exactly the two-branch split the dispatcher needs.
The principle
The agentic loop continues while stop_reason is "tool_use" and terminates when it is "end_turn". Those two literals are what the control flow switches on.
67% of 390 answers on this site got it right on the first try.
Tier-2 staff say escalated disputes arrive with only a one-line reason, so they re-interview the customer from scratch. escalate_to_human currently accepts a single free-text reason string. Which change most directly removes the rework?
- AGrant Tier-2 direct read access to the session transcript so they can reconstruct what was already established
- BExtend escalate_to_human to carry the verified customer ID, root-cause analysis, and recommended action
- CAttach the full tool-call and tool-result history of the session to every escalation the agent files
- DRequire the agent to draft a customer-facing apology and resolution summary before it escalates
Show the answer and why each option wins or loses
A
Removes the information gap, but shifts the reading cost onto the human: a transcript still has to be mined for the customer ID, the finding, and the proposed remedy. It restores raw data where a conclusion is needed.
BCorrect
Correct. A structured handoff carries exactly what the receiving human needs to act — who the customer is, what was found, and what the agent recommends — without them replaying the conversation.
C
Preserves everything the agent saw, and is useful for later debugging, but raw tool output is even denser than a transcript and buries the analysis Tier-2 needs at the top.
D
Improves the customer's experience of the handoff, and is worth doing, but the apology text is addressed to the customer and carries none of the case facts Tier-2 is missing.
The principle
Escalation is an interface between an agent and a human who lacks the agent's context. The handoff payload should be a structured conclusion — identity, root cause, recommended action — not the raw material the conclusion came from.
78% of 285 answers on this site got it right on the first try.
lookup_order returns orderDate as a Unix epoch integer while the billing MCP server returns ISO 8601 strings. The agent has twice quoted a 1970 delivery date to a customer on a call. Which change reconciles the two representations in code?
- AA PreToolUse hook whose updatedInput adds a canonical date-format argument to each of the two lookup calls
- BA PostToolUse hook that rewrites every timestamp field in a returned result into one format
- CAn orderDate property typed as an ISO 8601 date string in lookup_order's own JSON input schema
- DA loop check that rejects any result whose orderDate is not ISO 8601 and re-issues that call
Show the answer and why each option wins or loses
A
PreToolUse can rewrite a call before it executes, which is the right instrument when the defect is in the arguments being sent. The epoch integer is chosen by each backend as it answers, so nothing in the outgoing call decides it.
BCorrect
Correct. PostToolUse runs on the result after the call returns and before the model reads it, which is the point at which a tool's own output can be reshaped in code rather than left to the model's judgment.
C
A tool's input schema is a genuine constraint and does reject malformed arguments before a call goes out. It governs the arguments Claude sends rather than the payload the tool hands back, so it never sees the value at issue.
D
Rejecting a non-conforming result is deterministic and would keep the 1970 date away from the customer, but the backend returns the same representation every time, so the call is re-issued until the loop abandons a lookup that succeeded.
The principle
Hook phase decides what a hook can touch: PreToolUse sees the outgoing call, PostToolUse sees the returned result. Reconciling formats that two backends produce is work on the result, and a tool's input schema has no view of it.
73% of 282 answers on this site got it right on the first try.
On a disputed charge the coordinator verifies the caller with get_customer and pulls the order with lookup_order, then spawns the billing-dispute subagent, which calls both tools again on the same account. Tool calls per dispute have roughly doubled. Which two changes address the cause?
Choose 2
- AGive the subagent one named slice of the case, the disputed charge and its payment history, and nothing wider
- BAdd a result cache in front of get_customer so the subagent's repeat lookup is served without a second call
- CSplit the record types between them, so the coordinator reads identity and the subagent reads orders, each once
- DTell the subagent in its prompt to skip any lookup the coordinator has most likely already run for this case
- EHave both agents work the whole ticket and then reconcile their two findings, so that duplicated conclusions merge
Show the answer and why each option wins or loses
ACorrect
Correct. A subagent handed a defined slice stops re-deriving what the coordinator already holds; scope, not memory, is what removes the duplicate calls.
B
Cuts the cost of the second lookup and is worth having on a hot path, but both agents still spend a turn deciding to make it and the duplicated reasoning remains.
CCorrect
Correct. Partitioning by record type is the other axis of the same decomposition: each source is read by exactly one agent, so coverage stays complete without overlap.
D
Sets a sensible expectation and costs nothing to try, yet the subagent cannot see which lookups the coordinator ran, so it is guessing about state it has no access to.
E
Produces one consistent answer and catches contradictions between the two agents, but it accepts the duplicated work as the price rather than removing it.
The principle
Decomposition means giving each agent a distinct slice, a subtopic or a source type, so the work is covered once. Reconciling afterwards, or hoping one agent guesses what another already did, leaves the duplication in place.
38% of 111 answers on this site got it right on the first try.
Since the system prompt gained the line "always look up the customer's history before deciding," the agent calls get_customer twice per turn — once to verify identity and again before each order action. Both tool descriptions are detailed and accurate. What is the most likely cause?
- AThe prompt's "look up the customer's history" wording binds that instruction to get_customer on each turn
- BThe get_customer description overlaps lookup_order enough that the agent hedges by calling both tools
- CThe agent cannot see its own earlier tool results within a turn, so it re-queries identity before acting
- Dtool_choice is set to "any", which forces at least one tool call before the agent may answer in text
Show the answer and why each option wins or loses
ACorrect
Correct. Keyword-sensitive system prompt language creates unintended tool associations: "look up" and "customer" read as a standing directive to invoke get_customer, overriding otherwise sound descriptions.
B
Would be the usual first suspect and is worth ruling out, but the stem states both descriptions are detailed and accurate, and the symptom is repetition of one tool rather than misrouting between two.
C
Would explain re-querying, and context loss does cause duplicate calls in other systems, but tool results are appended to the conversation and remain visible for the rest of the turn.
D
Would guarantee a tool call rather than prose, which is a real effect of that setting, but it does not explain why the same tool is chosen a second time in a single turn.
The principle
System prompt wording competes with tool descriptions for control of tool selection. When behavior changes right after a prompt edit, read the new wording for keywords that map onto a tool name before suspecting the tool definitions.
59% of 150 answers on this site got it right on the first try.
When the payments backend times out, process_refund returns isError: true with the text "Operation failed." Policy rejections return the same payload. The agent now retries policy rejections indefinitely and gives up on real timeouts. What should the tool return instead?
- AA distinct human-readable error string per failure mode so the agent can match on the wording it gets
- BA retry loop inside the tool itself, surfacing only the final outcome after three attempts have failed
- CThe backend's own exception class name in a code field, so the agent can branch on the code it gets
- DAn errorCategory value and an isRetryable boolean alongside a readable description of what failed
Show the answer and why each option wins or loses
A
Gives the agent more to go on than one generic string, but it puts the retry decision back on string matching, which breaks the moment someone rewords a message.
B
Handles transient failures close to the source, which is genuinely good practice, but it does nothing for the policy rejections that should never be retried at all.
C
A stable code is a real improvement on prose and is something the agent can match exactly rather than interpret. It exports the backend's internal taxonomy, so the agent must carry its own map of which classes are worth retrying.
DCorrect
Correct. Structured error metadata lets the agent branch on isRetryable rather than infer intent from prose, and errorCategory separates transient failures from business-rule violations.
The principle
Uniform failure payloads strip an agent of the information it needs to recover. Errors should carry a category and an explicit retryability flag so the retry decision is data, not interpretation.
87% of 128 answers on this site got it right on the first try.
run_query lets the billing subagent compose arbitrary queries against the billing database. Twice last quarter it returned invoice rows belonging to a different account whose contact name matched the caller's. Which two changes most directly address this?
Choose 2
- AReplace run_query with a lookup_invoice tool that takes an invoice ID and rejects invoices outside the account in dispute
- BRewrite run_query's description so it warns that the tool is a last resort, for questions no specific billing tool covers
- CRemove run_query from the billing subagent's tool set once lookup_invoice covers the lookups the role actually needs
- DAdd a PostToolUse hook that drops any returned row whose account id is not the account currently in dispute
- EUse forced tool selection so lookup_order runs first on each billing turn and establishes the account in context
Show the answer and why each option wins or loses
ACorrect
Correct. A constrained replacement makes the account boundary part of the tool contract, so a cross-account row cannot come back however the model phrases its request.
B
Sharpens selection and would reduce casual reaching for it, but the capability is unchanged: a tool that can read any account still will when the description is outweighed.
CCorrect
Correct. The narrow tool only helps if the broad one is gone; leaving both in the set keeps the unconstrained path one selection away.
D
Filtering the result set is deterministic and does keep the foreign rows out of the model's context, which is worth having as a backstop. The query still ran unconstrained, so the read happened and the tool can be asked for anything.
E
Establishes the account early and is a sound use of forced selection, but nothing then stops a later run_query from reaching past that account anyway.
The principle
Guardrails belong in the tool interface, not in the instructions wrapped around it. Narrowing a generic tool means both building the constrained replacement and retiring the general tool it replaces.
33% of 60 answers on this site got it right on the first try.
Refund work must be scoped to a verified account, and get_customer is the tool that verifies one. On tool_choice "auto" the agent opens roughly one dispute in eight with lookup_order instead, then works whichever account that order belongs to. What settles the opening call?
- ASet tool_choice to "any" for the opening turn, so the agent must call some tool before it may reply in text.
- BUse forced tool selection to name get_customer as the opening call, then let the dispute continue on auto.
- CState in the system prompt that get_customer is the required first call on every dispute the agent picks up.
- DWithhold lookup_order until get_customer has returned, then supply it for the turns that follow in the dispute.
Show the answer and why each option wins or loses
A
Guaranteeing a tool call rather than conversational text is the right instrument when an agent answers from memory, and it would fix that failure. It does not say which tool, so the eighth dispute still opens on lookup_order.
BCorrect
Correct. Forced tool selection names the one tool that must be called, which is how a specific call is guaranteed to run first; the rest of the dispute then proceeds on auto in the follow-up turns.
C
Prompt instructions steer selection, cost nothing to add, and would carry most disputes. A stated preference competes with the model's own reading of the tools rather than deciding the call, so a fraction still goes the other way.
D
Varying the tool set by turn does make the wrong opening impossible and is a real technique. It asks the caller to track dispute state and rebuild the tool set each turn, where a single tool_choice value does the same job.
The principle
tool_choice "any" guarantees that some tool is called; forced selection guarantees which one. When a specific call has to come first, name it and let the following turns run on auto.
71% of 122 answers on this site got it right on the first try.
A developer who joined the atlas monorepo last week keeps producing commit messages in the wrong format, while everyone else's sessions get it right. Her clone is current and the convention is written down. What is the most likely cause?
- AThe convention lives in each teammate's ~/.claude/CLAUDE.md, which version control never carries to a new clone.
- BHer project-level .claude/CLAUDE.md is being shadowed by a directory-level CLAUDE.md that omits the commit rules.
- CThe convention sits in a .claude/rules/ file whose paths frontmatter excludes the files she has actually been editing.
- DHer sessions run long enough that the root CLAUDE.md is loaded and then displaced as earlier context is dropped.
Show the answer and why each option wins or loses
ACorrect
Correct. User-level configuration applies only to the user who wrote it; it is never shared through the repository, so a teammate who was never told to create it simply does not have the rule.
B
Directory-level CLAUDE.md files do add scoped instructions, and diagnosing an omission there is a real exercise — but that would affect everyone editing those directories, not one new clone.
C
A paths mismatch is a genuine cause of rules that never fire, and worth checking. It would not explain a convention that works for every other developer in the same files.
D
Context pressure does degrade adherence in very long sessions, and shortening them helps. The pattern here is per-person rather than per-session, which points at configuration scope.
The principle
When behaviour differs by person rather than by file or by session, suspect the configuration hierarchy: user-level instructions are invisible to teammates.
59% of 116 answers on this site got it right on the first try.
The tech lead has written /atlas-scaffold, which generates a handler, a test and docs from one route definition. Every maintainer across the four packages must have it the moment they clone or pull atlas, with no per-machine setup, and it must not sit in context on the sessions that never run it. Where does the command file belong?
- AIn .claude/commands/ inside the atlas repository, which version control carries to every clone and pull
- BIn ~/.claude/commands/ on each maintainer's machine, copied there from a reference file kept in the repo
- CIn a .claude/rules/ file with a paths glob, so that it activates whenever a route definition is edited
- DIn the root CLAUDE.md, under a heading naming the scaffold procedure that the four packages all share
Show the answer and why each option wins or loses
ACorrect
Correct. A command file under the repository's .claude/commands/ is version-controlled, so it arrives with the code and needs no install, and its body is read when someone invokes it rather than loaded into every session.
B
This is a real commands directory and the command runs perfectly for whoever placed a copy there, but the user-level path sits outside version control, so each maintainer installs it by hand and the copies drift apart.
C
Path-scoped rules do load conditionally and would keep scaffold guidance out of unrelated sessions, which is half of what is being asked. A rule file is context that activates on a matching edit, not an artefact anyone invokes by name.
D
The root file travels with the repository and reaches every maintainer, so the distribution half is satisfied. It is always-loaded context rather than an invocable command, so it is carried by every session and nothing answers to /atlas-scaffold.
The principle
Two properties decide the placement: whether version control carries the file, and whether its contents load on invocation or in every session. Of the surfaces available, .claude/commands/ is the only one that gives both.
69% of 104 answers on this site got it right on the first try.
The team's dep-audit skill walks every package manifest in atlas and prints a few thousand lines of findings. Developers report that running it mid-task leaves Claude vague about the refactor they were in the middle of. Which frontmatter change addresses that?
- ASet allowed-tools to Read and Grep so the skill cannot modify any files while it walks the package manifests.
- BSet context: fork so the skill runs in an isolated sub-agent and only its summary returns to the session.
- CSet an argument-hint so developers scope the audit to a single package rather than to the whole monorepo.
- DMove the skill into ~/.claude/skills/ so each developer keeps a personal copy tuned to the packages they own.
Show the answer and why each option wins or loses
A
Restricting allowed-tools is a sound safety measure for a skill that only needs to read, and it prevents destructive edits. It does nothing about the volume of output landing in the conversation.
BCorrect
Correct. context: fork runs the skill in an isolated sub-agent context, so its verbose output never pollutes the main conversation and the refactor in progress stays intact.
C
An argument-hint advertises the parameter a skill takes, and an audit scoped to one package would indeed print less. It relies on the developer narrowing scope rather than isolating the output structurally.
D
A personal variant in the user-level skills directory avoids affecting teammates and can be tuned per developer. The output still lands in whichever session invokes it.
The principle
Skills that produce verbose or exploratory output belong in a forked context; isolation preserves the main conversation regardless of how the developer invokes them.
64% of 108 answers on this site got it right on the first try.
A read-through cache in front of the pricing service is next up in packages/api. Nobody has settled how entries are invalidated, what a miss during a deploy should do, or whether a stale read is ever acceptable. A developer opens a session to build it. How should the session start?
- AAsk for three candidate designs up front and compare them once written, keeping whichever handles invalidation best
- BHave Claude question the developer about invalidation, deploys and staleness until the requirements settle
- CWrite the test suite from the acceptance criteria first, then iterate by feeding the failing test output back
- DBuild the cache with conventional defaults and refine it against the review comments the first pull request draws
Show the answer and why each option wins or loses
A
Three designs make the tradeoffs concrete and can be compared quickly. It builds the thing three times to elicit requirements a conversation would settle, and the comparison has no agreed criteria to judge against.
BCorrect
Correct. The interview pattern surfaces the decisions nobody has made — invalidation, deploy-time misses, tolerable staleness — before any code exists that quietly assumes an answer to each of them.
C
Test-first is the right driver when the expected behaviour is known and can be written down, as it is for a normalizer with fixed output shapes. Here the tests would encode guesses, and passing them would prove nothing.
D
Shipping something conventional is fast and puts a real artefact in front of reviewers. Review comments arrive after the design is chosen, which is the expensive moment to learn the team wanted different staleness rules.
The principle
When the requirements themselves are unsettled, have the model interview you before it writes anything. Test-first and iterate-on-review both assume the expected behaviour has already been agreed.
59% of 108 answers on this site got it right on the first try.
The extraction prompt tells the model to flag any invoice field that looks questionable. Roughly 30% of invoices come back flagged, most for fields the AP reviewers find are fine, and the reviewers have started clearing the flag queue without reading it. Which change most directly improves flag precision?
- ATell the model to be conservative and flag only high-confidence problems, leaving borderline fields unmarked.
- BHave the model attach a 1-10 certainty rating to each flag so reviewers can work the strongest ones first.
- CReplace it with named conditions - total mismatch, missing tax ID, date outside the billing period - with examples.
- DRun a second extraction pass on every invoice and flag only the fields where the two passes disagree.
Show the answer and why each option wins or loses
A
This does dial down flag volume, and lower volume is part of what reviewers want. But "be conservative" and "high-confidence" are the same vague instruction in a quieter register: the model still has no definition of what counts as questionable, so the flags that survive are an arbitrary subset rather than a more accurate one.
B
Certainty ratings give reviewers an ordering, which is a real improvement to queue triage. They do not improve the underlying judgment, though - a model that cannot tell a genuine problem from a normal value will rate its bad flags confidently too, so the top of the sorted queue stays polluted.
CCorrect
Correct. Explicit categorical criteria with concrete examples define what a flag means, so the model is classifying against a rule rather than guessing at "questionable." That is what moves precision, and it is what restores reviewer trust in the flags that remain.
D
A second pass and a disagreement filter would catch genuine instability in the extraction, and it is a reasonable consistency check. It costs a full extra call per invoice and still measures agreement rather than correctness: two passes reading the same vague instruction can agree on the same wrong flag.
The principle
Vague quality instructions cannot be fixed by softening them or by filtering their output. Precision comes from specific categorical criteria that define what qualifies, illustrated with examples.
79% of 100 answers on this site got it right on the first try.
A field-level audit shows payment_terms flags are wrong on about 40% of the invoices they fire on, while vendor_name and stated_total flags are almost always right. Reviewers have begun skimming past all three. Rewriting the payment_terms criteria will take a full sprint. What should you do in the meantime?
- AKeep all three categories running but mark payment_terms as lower-confidence in the reviewer UI so reviewers weight it.
- BRaise the flagging threshold across all three categories so fewer invoices land in the reviewer queue.
- CTurn off payment_terms flagging while its criteria are rewritten, leaving the two accurate categories running.
- DSuspend all flagging until payment_terms is rewritten and the full rule set has been re-validated.
Show the answer and why each option wins or loses
A
Labelling the weak category is honest and keeps some payment_terms coverage alive, which has value. It also leaves the noise in the queue and asks reviewers to do the discounting themselves - the behaviour they have already shown they will not do, since they are skimming past all three today.
B
A uniform threshold rise would reduce total queue volume, which addresses the reviewers' workload complaint. It suppresses the accurate categories at the same rate as the inaccurate one, so you lose real findings from vendor_name and stated_total to fix a problem that lives in only one category.
CCorrect
Correct. The damage from a high false-positive category is not confined to that category - it teaches reviewers that flags in general are noise. Disabling the one bad category preserves the trust the two accurate ones have earned, and costs only the payment_terms coverage you were not getting value from anyway.
D
Stopping everything guarantees no further false positives reach reviewers, which is a defensible way to protect trust. It also discards two categories that are working, leaves invoices unchecked for a sprint, and treats a localised precision problem as though the whole rule set were suspect.
The principle
False positives are contagious across categories: one noisy category undermines confidence in accurate ones. Disable the offending category specifically rather than degrading or suspending the whole system.
70% of 104 answers on this site got it right on the first try.
A newly onboarded carrier starts sending shipping manifests on Monday, about 200 a night, on a layout the pipeline has never processed. The AP team wants them in the same nightly Message Batches run as the other 1,200 documents from the first night. What should happen before that batch goes out?
- ASubmit them with the rest, then resubmit whatever failed by custom_id once the morning's failure list is in.
- BRun a sample of the carrier's manifests, refine the extraction prompt on the results, then batch all 200.
- CSubmit them with the rest under a tightened extract_manifest schema, so a misread layout is rejected on arrival.
- DExtract this carrier's manifests synchronously for the first week, then fold them into the nightly batch run.
Show the answer and why each option wins or loses
A
Resubmitting failed custom_ids is the documented way to handle batch failures and is worth planning for anyway. A layout the prompt has never met fails as a class rather than at random, so most of the 200 come back for a second paid run.
BCorrect
Correct. Refining the prompt on a sample before batch-processing a large volume is what raises the first-pass success rate, and a dozen sample manifests cost a fraction of the resubmission a bad first night would force.
C
A tighter schema does keep malformed extractions out of the AP ledger and is good hygiene on any new document type. Rejection is not extraction: the batch is paid for in full and every rejected manifest still has to be worked afterwards.
D
Synchronous calls surface each failure immediately and would certainly expose the new layout fast. It forgoes the batch saving for a week on volume that has no latency requirement, to learn what a sample of a dozen answers in an hour.
The principle
A batch is committed and paid for all at once, so first-pass quality has to be established before submission. Prompt refinement on a sample is what buys that, and it is cheapest on a document type nothing has been tuned against.
63% of 87 answers on this site got it right on the first try.
Shipping manifests sometimes record quantities informally, such as "2 pallets (48 cases)". The extraction records 2 on some manifests and 48 on others with no discernible pattern. You are adding three few-shot examples of informal quantity phrasing. What makes them most likely to generalise to phrasings you have not seen yet?
- AEach example covers one more of the specific phrasings seen in the last month of incoming manifests.
- BThe examples sit at the very end of the prompt, immediately before the document being processed.
- CEach example shows the complete schema output, so every field is demonstrated alongside quantity.
- DEach example states why the chosen unit was picked over the plausible alternative in that document.
Show the answer and why each option wins or loses
A
Covering observed phrasings guarantees those exact cases are handled, and for a fixed vocabulary that would be enough. Manifest phrasing is open-ended, so pure pattern coverage produces a lookup table: the model matches the demonstrated strings and is back to guessing on the fourth phrasing that shows up next month.
B
Placement near the document is a sound prompt-construction habit and keeps the examples salient. It changes where the model reads the examples, not what they teach - three examples that only show inputs and outputs teach the same thing wherever they sit in the prompt.
C
Showing the full schema output does reinforce the overall shape of a correct extraction, which is useful when several fields are unstable. Here only quantity is unstable, and padding each example with the other fields dilutes the signal about unit choice rather than sharpening it.
DCorrect
Correct. Examples that show the reasoning - why cases rather than pallets was the right unit here - teach the judgment rather than the string. That is what lets the model generalise to novel phrasings instead of matching pre-specified cases.
The principle
Few-shot examples for ambiguous cases should show why one reading was chosen over a plausible alternative. Reasoning generalises to novel inputs; input-output pairs alone only cover what they demonstrate.
47% of 107 answers on this site got it right on the first try.
For a three-issue dispute the coordinator concatenates the identity check, the order history and the billing findings into one long prompt, with billing in the middle. Identity and order details reach the customer reply; the billing findings are routinely missing from it. Which two changes address this positional effect?
Choose 2
- APlace a key-findings summary at the beginning of the aggregated prompt
- BRe-order the three blocks by recency so the newest findings sit first
- CDrop the billing block and re-fetch those findings on demand instead
- DOrganise the detail beneath an explicit section header for each block
- ERaise the reply's max token limit so all three issues fit in one message
Show the answer and why each option wins or loses
ACorrect
Correct. Long inputs are processed reliably at the beginning and the end, so a summary of all three blocks placed first puts the billing findings where they will be read.
B
Keeps the freshest work near the front and is a sensible ordering rule in a long session, but it only moves which block occupies the middle; whichever one lands there is omitted next.
C
Cuts the prompt down and would help if the failure were volumetric. It is positional, so re-fetching returns the same findings to the same weak position, at the cost of an extra call.
DCorrect
Correct. Explicit section headers organising the detailed results are the guide's second mitigation for position effects, giving the middle material structure the model can locate.
E
Prevents a long resolution message being cut off mid-sentence, a real failure elsewhere. Here the reply completes; it simply never contains the billing material in the first place.
The principle
Findings in the middle of a long input can be omitted while the beginning and end come through. The remedy is placement and structure — summary first, explicit section headers — not trimming.
85% of 46 answers on this site got it right on the first try.
get_customer returns three accounts matching "J. Alvarez" in the same city. The agent currently selects the one with the most recent order and continues into the refund flow. What should it do instead?
- AEscalate to a human at once, since ambiguous identity is a policy gap the agent cannot resolve alone
- BRead back the most recent order's details and let the customer correct the record if it looks wrong
- CRank the candidate accounts by profile completeness and proceed with the highest-scoring match found
- DAsk the customer for one more identifier, such as an order number or the email on the account
Show the answer and why each option wins or loses
A
Is the right instinct when policy is silent on a request, and it does prevent a wrong-account refund, but handing off a question the customer can answer in one turn spends human capacity needlessly.
B
Involves the customer rather than guessing silently, which is better than the current behavior, but it anchors on a guess and invites a distracted customer to confirm someone else's order.
C
Replaces one heuristic with a more defensible one, and completeness scoring is useful for ranking search results, but any heuristic still resolves identity without evidence from the customer.
DCorrect
Correct. Multiple matches call for clarification, not selection. One additional identifier resolves the ambiguity from authoritative input and keeps the case in first-contact resolution.
The principle
When tool results return multiple matches, request an additional identifier rather than selecting heuristically. Escalation is for explicit customer requests, policy gaps, and genuine lack of progress — not for questions the customer can answer.
84% of 98 answers on this site got it right on the first try.
The billing subagent's invoice query times out after two internal retries. It returns "billing data unavailable" to the coordinator, which then abandons the entire resolution — including the damaged-item claim it had already substantiated. What should the subagent return?
- AAn empty result marked successful, so the coordinator proceeds with the concerns it can still resolve
- BThe same status plus a directive telling the coordinator to escalate this ticket to a human reviewer
- CThe failure type, the query it attempted, any partial results, and the alternatives still available
- DA retry request asking the coordinator to re-delegate the identical query after a short backoff period
Show the answer and why each option wins or loses
A
Keeps the workflow alive, which is the improvement being sought, but it disguises a failure as a valid empty result, so the coordinator reports "no billing issues found" for an issue it never saw.
B
Ensures a human eventually sees the case and is appropriate once options are exhausted, but it escalates the whole ticket over one unavailable subquery instead of resolving what is already resolvable.
CCorrect
Correct. Structured error context lets the coordinator decide intelligently: keep the substantiated claim, annotate the billing gap, and choose whether an alternative path is worth attempting.
D
Delegates recovery upward, and backoff is the right response to a transient failure, but the subagent already retried twice, and the coordinator has no information to decide differently.
The principle
Generic failure statuses hide the context a coordinator needs, and both silent suppression and whole-workflow termination are anti-patterns. Propagate failure type, what was attempted, partial results, and alternatives.
71% of 85 answers on this site got it right on the first try.
Three days into the HTTP client migration, the session has filled with verbose discovery output and answers about atlas's wrapper classes have started drifting toward typical adapter patterns. Select the TWO moves that address this.
Choose 2
- ARun /memory to release the context that the discovery output has been consuming.
- BRun /compact to reduce context usage now that the window is full of discovery output.
- CRe-enter the work with --resume so the session reloads with a smaller footprint.
- DKeep a scratchpad file of key findings and consult it when answering later questions.
- ESet context: fork on the session so the verbose discovery output lands elsewhere.
Show the answer and why each option wins or loses
A
/memory is a real command, and its name makes it sound like a context-pressure tool. What it actually does is report which memory files are loaded, for diagnosing inconsistent behaviour across sessions.
BCorrect
Correct. /compact is named for exactly this state: reducing context usage during an extended exploration session once the window has filled with verbose discovery output.
C
--resume does continue a named investigation across working sessions, which is useful here on other days. Resumption restores prior context rather than shrinking it, so it does not relieve a full window.
DCorrect
Correct. A scratchpad file of key findings, referenced when answering later questions, is the named counter to context degradation, because the findings live outside the window.
E
context: fork is real, but it is SKILL.md frontmatter that runs a skill in an isolated sub-agent context. It is set on a skill, not on a running session, so it cannot be applied to this one mid-migration.
The principle
Two named mechanisms cover this state: /compact for reducing context usage when an extended session fills with verbose discovery output, and scratchpad files that record key findings outside the window. /memory, --resume and context: fork are each named for a different job.
49% of 39 answers on this site got it right on the first try.
Right-rates are first attempts on this site's free CCAR-F form, as of September 15, 2026.
Take all 60 free CCAR-F questionsWhere to go next
- The full free form. All 60 questions of Form 1, in study mode or timed exam mode: CCAR-F Form 1.
- What the exam covers. The domains and their weights: CCAR-F domains.
- Cost, length and passing score. CCAR-F exam facts.