AI Certificates

Free practice questions

20 free CCAR-P (CCA-P) practice questions with explanations

Matthew Hartman · CCAR-P certified, scored 965/1000 · 1 min read

Twenty original Claude Certified Architect — Professional practice questions, spread across all 7 exam domains. Try each one, then open the answer: every option is explained, the wrong ones too, and each question shows how many people on this site got it right on the first try. The exam is also written CCA-P; it is the same exam.

These twenty questions come from the free CCAR-P Form 1 on this site, picked so that every exam domain appears. The questions are long scenarios where more than one option sounds reasonable, which is why the explanation on each wrong option matters as much as the key.

Answer each one before you open it. Then read every explanation, not just the one for the right answer: knowing why the tempting option loses is what carries over to the questions you have not seen.

Question 1D1 Solution Design & Architecture

Vantera Compliance must map a 1,400-page banking regulation to its internal control library of 900 controls. Single-call attempts, even with the full document in context, miss obligations in later sections and produce inconsistent mappings between runs. Which decomposition approach best addresses the failure?

  1. ASplit the regulation into equal 50-page chunks and run the full obligation-to-control mapping independently on each chunk, then concatenate the per-chunk results into one combined mapping.
  2. BHave Claude produce a condensed summary of the regulation first, then perform the obligation-to-control mapping against the summary so the whole task fits comfortably in a single context.
  3. CKeep the single-call design but move to the most capable available model, place the mapping instructions at both the start and end of the prompt, and lower the temperature for consistency.
  4. DDecompose by task stage: extract obligations section-by-section into a normalized schema, map each extracted obligation to candidate controls in a second stage, then aggregate and deduplicate the results.
Show the answer and why each option wins or loses
  1. A

    Equal page-count chunks cut across the document's semantic structure, splitting obligations mid-clause and severing cross-references that mappings depend on.

  2. B

    Summarization is lossy compression; regulatory mapping demands obligation-level completeness, and dropped detail becomes silently unmapped requirements.

  3. C

    A stronger model and instruction repetition may nudge quality but keeps one call responsible for two distinct cognitive tasks over an overloaded context, so run-to-run inconsistency persists.

  4. DCorrect

    Separating extraction from mapping gives each call one focused job over a right-sized input, the intermediate schema makes coverage auditable, and aggregation handles overlaps explicitly.

The principle

Effective decomposition splits along the task's natural stages and the document's semantic structure, producing verifiable intermediate artifacts between stages. Arbitrary size-based chunking and lossy summarization both sacrifice exactly the completeness the task exists to guarantee.

61% of 975 answers on this site got it right on the first try.

Question 2D1 Solution Design & Architecture

Halvorsen Freight, a third-party logistics provider, asks for 'GenAI in customer service.' Ticket analysis shows 40% of tickets are shipment-status lookups already answerable from the transportation management system (TMS) API, 25% are complex exceptions (damaged or misrouted freight), and customer satisfaction score (CSAT) correlates almost entirely with exception-resolution speed. Budget covers one initiative this year. Which Claude-based translation of the business need delivers the most value?

  1. ADeflect the 40% status tickets with a Claude chatbot grounded in help-center articles and shipment FAQs, since status lookups are the single largest ticket category by volume.
  2. BGenerate empathetic, brand-consistent reply templates for each ticket type so agents respond in a consistent tone and satisfaction improves across the service operation.
  3. CSummarize closed tickets into a weekly digest of themes and root causes so management gains visibility into emerging service trends and staffing needs.
  4. DBuild a tool-using Claude assistant that triages exception tickets, pulls live TMS and carrier data, and drafts a resolution plan for the human agent.
Show the answer and why each option wins or loses
  1. A

    Status lookups need a deterministic API integration, not retrieval over articles, and deflecting them barely touches CSAT because satisfaction is driven by exceptions.

  2. B

    Tone templates optimize a secondary axis; the data says speed of exception resolution, not phrasing, drives satisfaction.

  3. C

    Trend summaries are useful analytics but are an observability layer, not an intervention on the metric the business actually cares about.

  4. DCorrect

    It targets the 25% of tickets that drive the outcome metric, uses tool calls for live operational data, and augments agents on the work where model reasoning adds real leverage.

The principle

When budget forces one bet, translate the request by tracing which work actually moves the business metric, not which category has the highest volume. High-volume deterministic lookups belong in conventional automation; Claude's leverage is in the complex, judgment-heavy slice that drives the outcome.

64% of 723 answers on this site got it right on the first try.

Question 3D1 Solution Design & Architecture

Corvex Legal runs a Claude pipeline for NDAs: intake, OCR, Claude extraction to JSON, then load into the contract-lifecycle system. Paralegals quietly correct about 12% of extracted fields inside the contract lifecycle management (CLM) system after load. Extraction accuracy has not improved in six months despite two prompt rewrites. What is the PRIMARY architectural gap?

  1. AThe extraction prompt lacks enough few-shot examples demonstrating edge-case clause language, so unusual NDAs fall outside the demonstrated patterns.
  2. BThere is no feedback loop: the paralegals' corrections never flow back into an evaluation set that drives prompt and pipeline iteration.
  3. CThe pipeline uses a mid-tier model and should be upgraded to the most capable model available to raise extraction accuracy on difficult clauses.
  4. DThere is no pre-processing stage that scores OCR quality and rejects low-confidence scans before extraction, so garbled inputs reach the model.
Show the answer and why each option wins or loses
  1. A

    More few-shot examples might help, but without knowing which fields fail there is no principled way to choose examples, which is why the two blind prompt rewrites changed nothing.

  2. BCorrect

    The corrections are ground-truth labels being generated daily and discarded; capturing them closes the input-processing-output-feedback loop and turns improvement from guesswork into measurement.

  3. C

    A model upgrade is a blind bet; without measurement you cannot know whether the 12% failures are model-limited, prompt-limited, or input-limited.

  4. D

    OCR gating may address some failures, but you cannot know what fraction stem from scan quality until the correction data is captured and analyzed.

The principle

An end-to-end architecture is not complete at output delivery; it needs a feedback path that captures downstream human corrections as labeled evaluation data. Systems that discard this signal plateau, because every improvement attempt is unmeasured guesswork.

64% of 655 answers on this site got it right on the first try.

Question 4D2 Claude Models, Prompting & Context Engineering

A national insurance carrier processes 1.8 million inbound claims emails per day, routing each into one of 12 claim categories. The routing decision is a short classification with a p95 latency budget of 500 ms, and finance has capped the AI line item at $4,000/day. A pilot with the most capable frontier model achieved 97% routing accuracy but would cost roughly 20x the budget at full volume. Which approach best fits the workload?

  1. AKeep the frontier model but batch requests overnight in bulk processing windows, confirming the amortized cost fits within the $4,000/day cap
  2. BUse a mid-tier model with extended thinking enabled so deliberate reasoning holds routing accuracy near the 97% pilot level at lower per-token cost
  3. CUse the fastest, lowest-cost model with a few-shot prompt containing labeled examples of each category, and measure whether accuracy stays acceptable
  4. DFine-tune the frontier model on historical routing data so each routing request needs far fewer prompt tokens and the marginal cost per classification drops
Show the answer and why each option wins or loses
  1. A

    Overnight batching breaks the 500 ms real-time routing requirement; it treats the cost symptom while destroying the latency contract.

  2. B

    Extended thinking increases both latency and output-token cost, which moves the workload in the wrong direction for a simple, high-volume classification task.

  3. CCorrect

    A bounded 12-way classification is exactly the workload where a small, fast, cheap model plus well-chosen few-shot examples typically closes most of the accuracy gap, and an accuracy measurement validates the trade before committing.

  4. D

    Fine-tuning adds program cost and operational complexity before the cheaper lever, prompting a smaller model, has even been evaluated.

The principle

Model selection should start from the task's intrinsic difficulty, not from the most capable model that works. Simple, high-volume, latency-bound classification is the canonical fit for the smallest model in the family, upgraded with prompt engineering and validated with an eval before scaling.

62% of 397 answers on this site got it right on the first try.

Question 5D2 Claude Models, Prompting & Context Engineering

A legal-tech firm analyzes merger agreements for cross-document inconsistencies: about 200 requests per day, each carrying roughly 150K tokens of contract text, with attorneys reviewing every output. Errors that reach attorneys are expensive to catch, users tolerate multi-minute turnaround, and daily model spend at current volume is under $900 even on the most capable model. The team is debating whether to cut costs with a smaller model. What should they do?

  1. AStay on the most capable model, since at 200 requests/day the absolute cost is small relative to the cost of missed inconsistencies
  2. BSwitch to the smallest model and compensate with a far more detailed system prompt describing the legal reasoning steps to follow
  3. CSplit each 150K-token contract into 10K-token chunks and run them in parallel through a mid-tier model to cut per-request cost and turnaround time
  4. DRoute straightforward merger agreements to a small model and complex ones to the most capable model, using contract length as the routing heuristic the team maintains
Show the answer and why each option wins or loses
  1. ACorrect

    Low volume, high stakes, and generous latency tolerance mean capability dominates the trade-off, and the total spend is already immaterial next to attorney review costs.

  2. B

    Prompt detail does not substitute for reasoning capability on genuinely hard cross-document analysis; this optimizes a cost axis that is not the constraint.

  3. C

    Chunking destroys exactly the cross-document context the task depends on, trading correctness for a saving that was never needed.

  4. D

    Tiered routing is sound engineering, but document length is a poor proxy for analytical difficulty and adds routing complexity to save a few hundred dollars a day.

The principle

Cost optimization only matters when cost is actually the binding constraint. When volume is low, error costs are high, and latency is slack, the correct trade is the most capable model, and premature downgrading or chunking optimizes an axis nobody is constrained on.

40% of 304 answers on this site got it right on the first try.

Question 6D2 Claude Models, Prompting & Context Engineering

A retail bank's support chatbot must never give personalized investment advice. The current implementation prepends a 900-word policy paragraph, including this prohibition, to the start of every user message; behavioral rules, tone guidance, and the ban are interleaved in one block. In testing, the bot still occasionally recommends specific funds when users push back twice. What is the most effective structural fix?

  1. AAdd a second reminder of the prohibition at the end of each user message, closest to the model's response, so the ban is the very last instruction the model reads before it answers
  2. BMove the role definition and prohibitions into the system prompt as explicit, separately stated rules, including what the bot should say instead when asked for investment advice
  3. CLower the temperature to 0 and cap top-p so the model samples its most probable tokens and adheres more consistently to the instructions it was given
  4. DPost-process each response with a keyword filter that blocks replies mentioning fund names, tickers, or share classes before they reach the customer
Show the answer and why each option wins or loses
  1. A

    Repeating instructions inside user turns still leaves policy competing with conversational content and remains vulnerable to user pushback framing.

  2. BCorrect

    The system prompt is the designated channel for role and constraints, and pairing each prohibition with a specified fallback behavior gives the model a compliant action instead of just a ban to erode.

  3. C

    Sampling settings control output variability, not instruction adherence; a model that yields to pushback at temperature 0.7 can do so identically at 0.

  4. D

    A keyword filter treats the symptom, blocks legitimate factual mentions of funds, and misses advice phrased without any ticker or fund name.

The principle

Guardrails belong in the system prompt as distinct, explicit rules, and each 'never do X' is far more robust when paired with 'do Y instead.' Burying policy inside user-turn preambles leaves it structurally equal to the content it must override.

48% of 287 answers on this site got it right on the first try.

Question 7D3 Integration

A freight dispatch assistant has accumulated 42 tools, including warehouse admin functions and invoice voiding it has never legitimately needed. It now picks the wrong tool on about 8% of routine status updates, and a security review found that a prompt injection in a delivery note could trigger invoice voiding. What is the most effective remediation?

  1. AAdd detailed system-prompt guidance describing exactly when each of the 42 tools should and should not be used, with worked examples for each
  2. BReduce the agent's tool list to the tools its workflows actually require, and move admin-grade actions to a separately gated integration
  3. CUpgrade the assistant to a larger, more capable model tier so tool-selection accuracy improves across the full 42-tool list without redesign
  4. DRewrite each tool description to be longer and more distinctive so the model can discriminate more reliably among the 42 tools
Show the answer and why each option wins or loses
  1. A

    More routing prose treats the symptom: the context stays bloated, selection remains probabilistic across 42 options, and the dangerous tools are still reachable by injection.

  2. BCorrect

    Trimming to workflow-required tools shrinks the selection space (fixing accuracy) and removing admin actions from the agent's reach eliminates the injection blast radius rather than mitigating it.

  3. C

    A stronger model may raise selection accuracy but leaves the real problem — an over-broad, high-privilege capability surface — largely intact.

  4. D

    Longer descriptions consume more context per turn and still leave 42 tools, including destructive ones, exposed to a compromised instruction stream.

The principle

Capability bloat is both an accuracy problem and a security problem: every attached tool enlarges the selection space and the blast radius of a hijacked agent. The fix is to scope the tool list to the job, not to explain the bloated list better or throw model capability at it.

85% of 275 answers on this site got it right on the first try.

Question 8D3 Integration

A hospital assistant queries patient records through an electronic health record (EHR) integration that authenticates with a single service account holding organization-wide read access, chosen because provisioning per-user credentials was deemed too slow. During a pilot, a cardiology nurse received data on an oncology patient she has no treatment relationship with — a reportable access violation. Which change addresses the root cause?

  1. AAdd a system-prompt rule, reinforced with few-shot examples, instructing the model to return records only for patients within the requesting user's own department
  2. BPost-filter the assistant's responses with a second model pass that detects and redacts patient identifiers and out-of-department data before display
  3. CPropagate the requesting clinician's identity to the EHR on each call so the EHR's own authorization rules scope each query to that user's permitted records
  4. DDowngrade the shared service account from read-write to read-only and enable comprehensive audit logging with periodic access reviews on its queries
Show the answer and why each option wins or loses
  1. A

    Prompt instructions are not a security boundary; a model can be manipulated or simply err, and the data has already crossed the trust boundary into its context.

  2. B

    Redaction operates after unauthorized data has been retrieved into the model's context, and it filters output formatting rather than enforcing access rights.

  3. CCorrect

    Acting-as-user (on-behalf-of) delegation makes the system of record enforce its existing authorization model per request, so unauthorized rows never reach the assistant at all.

  4. D

    The violation was a read, so read-only changes nothing, and audit logging detects violations after they happen instead of preventing them.

The principle

When an agent fronts a permissioned system with a broad service account, every user inherits the service account's access — the model becomes a confused deputy. Authorization must be enforced by the system of record using the end user's identity, not approximated by prompts or output filters.

77% of 271 answers on this site got it right on the first try.

Question 9D3 Integration

A fintech is rolling out an internal MCP server that exposes payment-operations tools — refund issuance, chargeback lookups, and account-hold toggles — to Claude clients used by three different teams. Today any authenticated employee client can invoke any tool, and the server trusts whatever tool name and arguments the client sends. Compliance has flagged the design before launch. Which TWO changes most directly close the authorization gaps?

Choose 2

  1. ARequire an OAuth flow so every tool invocation carries a short-lived token scoped to the requesting employee's role
  2. BHave the server verify per invocation that the calling user is entitled to the specific tool, rather than trusting the client's presented tool list
  3. CMove the MCP server behind the corporate VPN so only managed devices on the internal network can reach its payment-operations endpoints
  4. DLog each tool call with its arguments and results, retained for quarterly compliance review of refund and account-hold activity
  5. ELower the client model's temperature so tool-invocation behavior is more deterministic and predictable across the three teams
Show the answer and why each option wins or loses
  1. ACorrect

    Short-lived, role-scoped tokens bind each call to a real user identity and permission set instead of a shared blanket credential.

  2. BCorrect

    Server-side enforcement means a modified, misconfigured, or prompt-injected client cannot invoke tools its user is not entitled to — the server, not the client, is the policy point.

  3. C

    Network placement controls who can reach the server, not what an authenticated employee is allowed to do once connected — the flagged gap is authorization, not reachability.

  4. D

    Audit logs are a detective control; they document a rogue refund after the fact rather than preventing an unauthorized one.

  5. E

    Temperature shapes sampling behavior, not permissions — a well-behaved model with excessive rights is still an authorization failure.

The principle

MCP integrations need the same authorization rigor as any API: identity bound to each request and policy enforced server-side. Network perimeter, logging, and model behavior tuning are complements, but none of them decides whether a specific user may perform a specific action.

64% of 121 answers on this site got it right on the first try.

Question 10D4 Evaluation, Testing & Optimization

A support-triage assistant scores 92% classification accuracy on a 500-ticket curated eval set, and the team declares it production-ready. Two weeks after launch, agents complain that responses take 11 seconds at p95 during peak hours, and one response echoed a customer's full credit card number from the ticket body. Leadership asks why evaluation didn't catch either problem. What was the core evaluation design flaw?

  1. AThe 500-ticket eval set was too small a sample to produce a statistically reliable estimate of classification accuracy
  2. BThe team used a hand-curated eval set instead of sampling tickets directly from live production traffic during representative peak hours
  3. CThe evaluation measured only one dimension, omitting latency and safety/security metrics that production success actually depends on
  4. DThe team should have used expert human graders rather than automated scoring so nuanced failure modes could be caught before launch
Show the answer and why each option wins or loses
  1. A

    500 items is a reasonable size for a classification eval; a larger set would sharpen the accuracy estimate but still measure only accuracy.

  2. B

    Curated sets are legitimate and often necessary; sampling live traffic changes the distribution but would not add latency or PII-handling measurement by itself.

  3. CCorrect

    The failures were on axes the evaluation never measured — production readiness requires defining metrics across accuracy, latency, cost, and safety/security, not accuracy alone.

  4. D

    Human graders improve judgment quality on the measured dimension but cannot surface p95 latency or data-leak risks that were never in scope.

The principle

An evaluation only protects you on the dimensions it measures. Define a metric suite spanning accuracy, latency, cost, and safety/security up front, because a system can pass a single-axis eval while failing in production on every axis you didn't measure.

64% of 257 answers on this site got it right on the first try.

Question 11D4 Evaluation, Testing & Optimization

To build an eval set for a document Q&A system, a team samples 300 real user queries from production logs, filtering to conversations that received a thumbs-up so that reference answers are trustworthy. The system scores 96% on this set, but a manual audit of live traffic estimates only 78% of production answers are acceptable. What is the most likely explanation for the gap?

  1. AFiltering to thumbs-up conversations selected for queries the system already handles well, excluding the failure modes behind the gap
  2. BThe eval set of 300 queries is too small, so the 96% figure carries a wide confidence interval that overlaps the production estimate
  3. CThe manual audit graded against stricter acceptability criteria than the automated eval did, so the two percentages are not directly comparable
  4. DProduction traffic drifted in topic mix after the eval set was collected, making the sampled queries unrepresentative of current usage patterns
Show the answer and why each option wins or loses
  1. ACorrect

    Sampling only from successful interactions is survivorship bias — the eval measures performance on the happy path and cannot detect the failures that make up the 18-point gap.

  2. B

    A 300-item sample gives roughly a ±2-3 point margin at these rates, which cannot explain an 18-point discrepancy.

  3. C

    Grader-criteria mismatch can shift numbers a few points, but it is speculative here, while the sampling bias is structural and certain given how the set was built.

  4. D

    Drift is possible over long periods, but the set was drawn from recent logs; the filtering step is the documented, immediate source of bias.

The principle

An eval dataset must represent the full failure distribution, including hard cases, edge cases, and adversarial inputs — not just interactions that already succeeded. Selection criteria that correlate with success turn an evaluation into a happy-path demo.

69% of 242 answers on this site got it right on the first try.

Question 12D4 Evaluation, Testing & Optimization

A contract-analysis system handles 2,000 queries per day. The eval framework runs deterministic checks (JSON schema, required citation fields) on every response and an LLM-judge for answer quality on a 10% sample. The team can afford expert human review of only 50 responses per week. To make the mixed-methodology framework most trustworthy, where should the scarce human review be focused?

  1. AOn a random sample drawn from all weekly production responses, so the human scores form an unbiased estimate of overall answer quality
  2. BOn the highest-traffic query categories, weighted by volume, since errors in those categories affect the largest number of users
  3. COn responses that already failed the deterministic schema and citation checks, since those are confirmed defects worth documenting in depth
  4. DOn responses where the LLM-judge verdict is borderline or conflicts with other signals, to calibrate the judge and expose blind spots
Show the answer and why each option wins or loses
  1. A

    A random sample of 50 out of 14,000 weekly responses is too sparse to estimate quality precisely and duplicates what the judge already covers at scale.

  2. B

    Volume-weighted review optimizes for exposure, not for validating the automated layers the framework depends on; the judge already scores those categories.

  3. C

    Deterministic failures are already reliably detected by machine; spending expert hours re-confirming known defects adds no new information.

  4. DCorrect

    In a tiered eval stack, human judgment is most valuable auditing the automated judge where it is least certain — disagreements and borderline cases — because judge calibration determines whether the other 13,950 automated verdicts can be trusted.

The principle

Mixed-methodology frameworks are hierarchies of trust: deterministic checks validate structure cheaply, LLM judges scale quality assessment, and humans calibrate the judges. Scarce expert review should target where automated evaluators are most likely wrong, because that is what makes the automated majority of the framework trustworthy.

61% of 238 answers on this site got it right on the first try.

Question 13D5 Governance, Safety & Risk Management

A regional bank deploys a Claude-powered assistant that answers customer questions about their own accounts. The team's only safety control is a system-prompt instruction: 'Never reveal information about accounts other than the authenticated user's.' During a red-team exercise, a tester uses a multi-turn injection to get the model to summarize another customer's transactions, because the retrieval layer fetches records by whatever account number appears in the conversation. Which change most directly fixes the vulnerability?

  1. AStrengthen the system prompt with explicit refusal examples and restate that only the authenticated user's account number may drive record retrieval
  2. BAdd an output classifier that scans responses for account numbers and identifiers not belonging to the authenticated user and blocks matches
  3. CEnforce the verified caller's identity as a mandatory filter on data lookups so records outside that scope are excluded before generation
  4. DLog conversations with their retrieval traces and route flagged sessions to the bank's fraud team for structured next-business-day review
Show the answer and why each option wins or loses
  1. A

    Prompt instructions are probabilistic and remain bypassable by injection; they cannot serve as the sole access-control boundary.

  2. B

    An output filter treats the symptom after the sensitive data has already entered the model's context, and pattern matching on summaries or paraphrases is unreliable.

  3. CCorrect

    Authorization enforced deterministically at the data layer means unauthorized records never enter the context, so no prompt attack can exfiltrate them.

  4. D

    Next-day review detects breaches after the fact; it is monitoring, not a control that prevents the disclosure.

The principle

Access control must be enforced deterministically in the system layer (retrieval, APIs, permissions), not delegated to the model via instructions. If sensitive data never enters the context window, no amount of prompt injection can leak it.

72% of 233 answers on this site got it right on the first try.

Question 14D5 Governance, Safety & Risk Management

A hospital's clinical documentation assistant summarizes physician dictations into structured notes. During validation, reviewers find that in about 2% of notes the summary includes a plausible-sounding medication dosage that was never dictated. The engineering team insists this is a retrieval bug, but the system performs no retrieval — it works only from the dictation transcript provided in the prompt. What is the most accurate characterization of this failure?

  1. AA context-window overflow in which the model drops portions of the transcript under length pressure and reconstructs them from surrounding context
  2. BA prompt-formatting defect that clearer delimiting of the transcript, separating instructions from dictated content, would resolve
  3. CIntrinsic hallucination: the model generates content unsupported by its provided input, an inherent failure mode requiring output verification
  4. DTraining-data contamination that can be corrected by fine-tuning the model on the hospital's own corpus of historical clinical notes
Show the answer and why each option wins or loses
  1. A

    Overflow truncation is a real failure mode but would present as missing content and would not occur on transcripts that fit comfortably in context.

  2. B

    Better delimiting can reduce confusion between instructions and data, but it cannot guarantee the model stops generating unsupported specifics.

  3. CCorrect

    Fabricating specifics not present in the source is intrinsic hallucination, a probabilistic property of generation that mitigation can reduce but engineering cannot fully eliminate — hence verification against the source is required.

  4. D

    Fine-tuning shifts style and domain fluency; it does not remove the underlying tendency to generate plausible unsupported details, and could make fabrications more convincing.

The principle

Hallucination is an inherent property of generative models, not an implementation bug to be patched away. Architectures handling high-stakes content must assume some fabrication rate and add verification — such as source-grounding checks or human review — for consequential fields like dosages.

72% of 224 answers on this site got it right on the first try.

Question 15D5 Governance, Safety & Risk Management

An insurance company uses an LLM to draft claim-denial letters. To satisfy regulators, every letter is routed to a claims adjuster who must click 'Approve' before it is sent. Adjusters each receive roughly 400 letters per day, approve 99.7% of them, and average 11 seconds per review. A regulator's audit later finds several denials that misstated policy terms and were approved. What is the fundamental flaw in this human-in-the-loop design?

  1. AThe review step is placed after generation instead of before it, so adjusters cannot shape the drafts while policy terms are being applied
  2. BThe review volume and approval pattern indicate rubber-stamping — the checkpoint exists procedurally but provides no meaningful validation
  3. CAdjusters lack the legal training to evaluate policy language and should be replaced with attorney reviewers qualified to parse contract terms
  4. DThe approval interface should require a short written justification for each approval so reviewers are forced to slow down and engage
Show the answer and why each option wins or loses
  1. A

    Post-generation review is the normal and appropriate placement for output validation; the flaw is not the position of the checkpoint.

  2. BCorrect

    11-second reviews at 400 per day with near-total approval is automation bias in action — the human checkpoint exists on paper but is operationally incapable of catching errors.

  3. C

    The adjusters are the right domain experts for claims; escalating credentials does not fix a workload that makes genuine review impossible.

  4. D

    Mandatory justifications add friction and process without addressing volume; reviewers under the same load will produce boilerplate justifications.

The principle

HITL is only a control when the human has the time, information, and incentive to genuinely evaluate each item. Effective designs triage by risk — sampling low-stakes outputs while concentrating full review capacity on high-consequence ones — rather than routing everything through a checkpoint that volume turns into a rubber stamp.

75% of 182 answers on this site got it right on the first try.

Question 16D6 Stakeholder Communication & Lifecycle Management

A VP of Customer Experience opens a kickoff call by saying, "We've decided we need an AI chatbot on our support site — how fast can you build it?" She has budget approved and a launch date already communicated internally. The architect has not yet seen any support metrics or spoken to the support team. What should the architect do FIRST?

  1. AAccept the chatbot direction and begin scoping channels, languages, escalation paths, and integration points so delivery stays aligned with the approved timeline
  2. BSketch a candidate architecture and phased rollout plan for the chatbot so the VP has something concrete to react to in the next meeting
  3. CAsk what business problem prompted the decision and what measurable outcome would define success before validating the chatbot as the solution
  4. DRequest that the announced launch date be revisited, since committing to a solution this early in discovery carries avoidable delivery risk
Show the answer and why each option wins or loses
  1. A

    Scoping the stated solution locks in the VP's assumption before anyone has verified a chatbot actually addresses the underlying problem.

  2. B

    A concrete architecture makes the unvalidated solution feel even more decided and skips the discovery the situation is missing.

  3. CCorrect

    Structured discovery starts by probing the problem and success criteria, because a pre-decided solution may not fit the actual need.

  4. D

    Reopening the timeline before understanding the problem creates conflict without adding information; the date may be fine once the real need is known.

The principle

When a stakeholder arrives with a solution instead of a problem, the architect's first move is to surface the underlying business problem and measurable success criteria. Validating or redirecting the solution comes after discovery, not before — otherwise you inherit someone else's untested assumption as your requirement.

85% of 224 answers on this site got it right on the first try.

Question 17D6 Stakeholder Communication & Lifecycle Management

During discovery for an AI document-processing system, the operations director insists the system must run fully automated end to end, while the compliance officer states that every output touching customer accounts must be human-reviewed. Both were interviewed separately and each believes their requirement is settled. The project sponsor asks the architect to "just write up the requirements." What should the architect do FIRST?

  1. ABring both stakeholders together over the conflict and facilitate agreement on risk tolerance and decision criteria before finalizing requirements
  2. BDocument both requirements exactly as stated and design a configurable architecture flexible enough to support either operating mode later
  3. CAdopt the compliance officer's position as the binding requirement, since regulatory stakeholders take precedence over operational efficiency preferences
  4. DPropose human review for a sampled percentage of account-touching outputs and full automation for the remainder as a balanced starting point
Show the answer and why each option wins or loses
  1. ACorrect

    Unresolved stakeholder conflict is a discovery finding; the architect brings the parties together and drives alignment on criteria rather than hiding the conflict in the document.

  2. B

    Recording contradictory requirements defers the conflict to implementation, where it becomes far more expensive and political to resolve.

  3. C

    Compliance input matters, but unilaterally overriding one stakeholder without a facilitated decision breeds resistance and may over-constrain the design unnecessarily.

  4. D

    Proposing a compromise neither party agreed to substitutes the architect's judgment for a stakeholder decision that hasn't actually been made.

The principle

Conflicting stakeholder requirements are not resolved by documentation tricks, unilateral rulings, or invented compromises — they are resolved by making the conflict visible and facilitating an explicit decision on risk tolerance and priorities. Requirement gathering includes reconciling stakeholders, not just recording them.

74% of 230 answers on this site got it right on the first try.

Question 18D6 Stakeholder Communication & Lifecycle Management

An architect recommends a mid-tier model for a claims-triage workload after testing shows it meets quality targets. In the review meeting, the CTO pushes back: "If we're doing this, I want the most capable model available — I don't want us looking cheap." Which response best communicates the recommendation?

  1. AExplain that the mid-tier model's context window, token throughput, and rate limits are more than sufficient for the prompt sizes and request volumes involved in the claims-triage workload
  2. BShow that both models meet the quality bar on the company's own test cases, that the choice saves projected monthly cost at lower latency, and name when it would be revisited
  3. CAgree to launch on the most capable model to preserve the executive relationship, planning an internal downgrade to the mid-tier once usage data accumulates
  4. DDefer the choice to the engineering team's technical judgment, since model selection is an implementation detail below the level of an executive review
Show the answer and why each option wins or loses
  1. A

    This answers in technical vocabulary the CTO didn't ask about; it explains capacity, not why the choice serves the business.

  2. BCorrect

    It translates the trade-off into measured quality, cost, and latency on the company's own data, and names the conditions under which the decision would change.

  3. C

    Conceding a decision you believe is wrong, then reversing it without reopening the conversation, undermines trust in both directions and hides the trade-off instead of communicating it.

  4. D

    Model selection drives cost, latency, and quality — exactly the outcomes executives own — so deflecting it as an implementation detail dodges the conversation.

The principle

Trade-off communication with executives means translating technical choices into business outcomes — measured quality on their data, cost, and speed — and stating the conditions under which the decision would be revisited. Jargon, quiet concessions, and deflection all fail because they avoid the actual trade-off.

67% of 201 answers on this site got it right on the first try.

Question 19D7 Developer Productivity & Operational Enablement

A platform team rolls out Claude Code to 14 engineers and results diverge: some get output that follows the repo's conventions and test patterns, others get generic code that ignores the monorepo structure. Each engineer wrote their own personal notes file for the assistant, and several have none. The engineering manager proposes a mandatory two-hour prompt-writing training. What is the most effective way to make assistant output consistent across the team?

  1. ARun the proposed two-hour training so each engineer learns to describe the repo's conventions and test patterns in their prompts
  2. BHave the strongest prompter publish their personal notes file and settings in the team wiki with instructions for the other engineers to copy them locally
  3. CCheck a project-level memory file (e.g., CLAUDE.md) and shared tool settings into the repository so every session starts from the same context
  4. DStandardize the whole team on a single pinned model version so engineers receive identical completions for identical prompts
Show the answer and why each option wins or loses
  1. A

    Training improves individual skill but leaves the underlying cause intact: each session still starts with different context, so divergence returns as people forget or interpret the training differently.

  2. B

    A wiki page depends on every engineer manually copying and updating it; it drifts from the repo and is not versioned or reviewed alongside the code it describes.

  3. CCorrect

    Committing project context and tool configuration to the repository gives every engineer's session the same conventions, structure, and commands automatically, and changes flow through normal code review.

  4. D

    Model pinning removes one variable but does not address the real driver of divergence, which is that engineers are supplying different (or no) project context.

The principle

Team-level inconsistency in AI-assisted development is almost always a configuration problem, not a skill problem. Shared, version-controlled context files and tool settings make consistency a property of the repository rather than of individual habits, and keep that context evolving under code review like any other artifact.

82% of 224 answers on this site got it right on the first try.

Question 20D7 Developer Productivity & Operational Enablement

One month after a team adopted an AI coding agent under an 'allow everything' policy, two incidents hit: a session ran a destructive database reset while fixing a failing test, and shared transcripts were found to contain production API keys the agent had read from a local .env file. Which TWO changes most directly prevent both classes of incident from recurring?

Choose 2

  1. ARotate the exposed keys immediately and circulate a detailed incident writeup so engineers monitor their agent sessions more closely going forward
  2. BReplace the allow-everything policy with a checked-in permission configuration that auto-approves safe read and test commands but gates destructive or state-changing operations
  3. CConfigure the agent for read-only access repo-wide so it can analyze code and propose changes but not execute commands or edit files on its own
  4. DMove secrets out of files the agent can read and add deny rules for secret file patterns (e.g., .env, key material) in the shared tool configuration
  5. ERequire a designated senior engineer to review the day's agent session transcripts in full at day's end before any of that work is merged
Show the answer and why each option wins or loses
  1. A

    Rotation and awareness are necessary incident response, but they change nothing structural: the same permissive configuration will produce the same exposures and destructive actions again.

  2. BCorrect

    A tiered, version-controlled permission policy directly addresses the destructive-command incident by gating dangerous operations while keeping low-risk operations frictionless.

  3. C

    Read-only mode prevents the incidents but destroys most of the tool's value; it is an over-correction that teams quickly abandon or work around, recreating the original risk informally.

  4. DCorrect

    Excluding secret files from the agent's readable context, enforced by shared deny rules rather than individual discipline, directly addresses the credential-exposure incident at its source.

  5. E

    End-of-day transcript review is detective, not preventive: the destructive command and the key exposure would both have already happened by the time anyone reads the transcript.

The principle

Safe agent adoption is a least-privilege problem: permissions should be tiered by blast radius and enforced through shared, version-controlled configuration, and secrets should be structurally unreachable rather than protected by human vigilance. Controls that rely on people watching harder, or that remove capability entirely, either fail silently or get bypassed.

63% of 91 answers on this site got it right on the first try.

Right-rates are first attempts on this site's free CCAR-P form, as of September 15, 2026.

Take all 63 free CCAR-P questions

Where to go next

Common questions

Are these real CCAR-P exam questions?
No. They are original questions written from the published CCAR-P exam guide, covering the same domains and objectives. Nothing in them comes from the live exam.
How should I use free CCAR-P practice questions?
Answer first, then read the explanation for every option, the wrong ones included. When you miss one, write down the principle in a sentence; the same idea comes back in other settings. Once you are scoring well here, sit the full free form in timed exam mode to practice pacing.
Where can I practice a full-length CCAR-P exam for free?
Form 1 on this site is a complete 63-question CCAR-P practice exam, free with no account. It runs in study mode, with explanations after each answer, or in timed exam mode.
What does the percentage on each CCAR-P question mean?
It is the share of first-try answers on this site that were right. A low number on a question you got wrong means it is a common miss; a high one means it is worth rereading the explanation.

Practice this

A full-length 63-question CCAR-P form. Free, no signup.

Start form 1 · study mode

Not sure where you stand? 3-minute readiness check →

Sets 2–6 carry the rest of the CCAR-P bank. Full access unlocks those and every set for the other three certifications, in one payment.

Get full access $39