Claude Certified Architect — Foundations Practice Exam
Practice exam for Anthropic's Claude Certified Architect — Foundations certification. Covers prompt design, agentic patterns, code review pipelines, evaluation strategies, structured outputs, and operational best practices, plus tool design and MCP integration, agent guardrails and escalation, Claude Code permissions, hooks and subagents, and context management with prompt caching.
100 questions · 10 free preview
Studying more than one? every exam for $79
Free sample questions
- Sample · question 1 · Narrowing label definitions to stop over-triggering
You run a Claude-powered ticket-classification pipeline that tags incoming customer requests with one of seven labels. Production logs show the model consistently picks the "billing" label for messages that mention any dollar amount, even when the message is clearly about feature requests that happen to reference pricing. The prompt currently says "Tag the ticket with the most relevant label." Which prompt change is most likely to fix the over-triggering without breaking other labels?
- A.Add an explicit "Only use the billing label when the user is asking about an existing charge, refund, or payment failure" instruction.correct
- B.Append a final instruction "Be careful and think before answering."
- C.Switch to a tool call where the model selects from an enum of labels.
- D.Lower the model's temperature to 0.
Why: Over-triggering on a lexical pattern (any dollar amount → billing) is a definition gap in the prompt. Narrowing when billing applies fixes the root cause. A tool-call with an enum still lets the model pick the wrong enum value; temperature changes don't address the model's mistaken concept of what 'billing' covers.
Open this question on its own page → - Sample · question 2 · Severity and confidence metadata for review findings
A code-review automation surfaces findings to engineers; the team complains that noisy findings drown out the genuine bugs. The team is unwilling to lose any genuine bug, so they want to keep recall high but reduce the noise downstream. Which output design best supports this tradeoff while keeping all of Claude's findings available for later analysis?
- A.Ask Claude to emit each finding with `severity` (low/med/high) and `confidence` (0–1) fields, and threshold programmatically downstream.correct
- B.Lower Claude's temperature and tell it "only report bugs you are highly confident about."
- C.Have Claude self-filter and emit only findings it would tag as high severity.
- D.Wrap Claude's review in a second pass that re-reads each finding and keeps only the ones flagged "true positive."
Why: Structured metadata on every finding preserves recall at the model layer (nothing is dropped) while letting downstream code threshold cheaply and adjust over time. Self-filtering and re-pass approaches throw away information you can never recover later for analysis.
Open this question on its own page → - Sample · question 3 · Tool-based verification to prevent hallucinated APIs
Your Claude-based migration tool reads a developer's TypeScript file and proposes a new version that swaps a deprecated library for its replacement. Engineers report it sometimes proposes correct-looking code that references functions which do not exist in the replacement library. Which architectural change is the most effective fix?
- A.Increase the static prompt context to include a full snapshot of the replacement library's public surface, pasted as plain text.
- B.Convert the pipeline to an agentic task where Claude is given a `lookupSymbol(name)` tool and instructed to verify each replacement function exists before proposing it.correct
- C.Lower the temperature to 0 and add the instruction "do not hallucinate."
- D.Run the proposal through a static analyzer in a second pass and have it reject calls to unknown functions.
Why: Function hallucinations come from the model guessing what's plausible. The fix is to make verification cheaper than guessing — give Claude a tool to look up symbols and require their use. Static snapshots get stale; a downstream rejector lets the failure happen and then catches it, wasting the model's effort.
Open this question on its own page → - Sample · question 4 · Precision vs recall trade-offs in extraction
Your team uses Claude to extract entities from medical referral letters. After deploying a stricter prompt, precision rose from 86% to 94% but recall dropped from 91% to 78%. The medical reviewers say missing entities is worse than spurious ones because doctors are trained to ignore extra detail but rarely catch omissions. What is the most appropriate next step?
- A.Ship the stricter prompt and start a project to backfill the missing entities with a manual review queue.
- B.Revert to the more lenient prompt and accept the lower precision — recall is the load-bearing metric for this use case.correct
- C.Run both prompts in parallel and average their outputs.
- D.Train a downstream model to fill in entities the stricter prompt missed.
Why: When stakeholders explicitly state which error mode is more harmful, the deployed configuration should match that preference. Reverting to the lenient prompt is the cheapest correct action; the alternatives either cement the wrong tradeoff or stack complexity on top of a misaligned policy.
Open this question on its own page → - Sample · question 5 · Evidence fields for auditable structured output
Your support automation summarizes customer threads into JSON for downstream routing. Product wants to start measuring how often the summary's "topic" field is wrong. Which output-design change best supports that measurement?
- A.Add a `rationale: string` field explaining why the topic was chosen.
- B.Add an `evidence: { thread_msg_ids: string[] }` field listing which messages drove the topic decision.correct
- C.Lower the temperature so the summary is more deterministic.
- D.Have Claude emit the top three candidate topics, not just one.
Why: Wrongness is only auditable if you can trace the decision back to source. Listing the messages that influenced the topic gives reviewers a concrete artifact to disagree with. Rationales are useful but unfalsifiable; multiple topics multiply the labeling cost without anchoring the disagreement.
Open this question on its own page → - Sample · question 6 · Tool schemas for strict JSON conformance
You need to extract structured information from invoices and pass it to a downstream system that rejects any deviation from a strict JSON schema (every field required, no extras). Which approach gives the most reliable schema conformance?
- A.Define a tool whose input parameters are the schema and have Claude call it with the extracted values.correct
- B.Embed the schema in the prompt as text and instruct Claude to "output JSON that matches the schema exactly."
- C.Pre-fill the assistant turn with `{` and parse the completion as JSON.
- D.Generate freely and run a downstream JSON repair step.
Why: Tool input parameters are validated by Claude's runtime against the declared schema — this is the highest-reliability path to conformance. Prompt-text schemas and pre-fills don't get machine-level validation; downstream repair masks problems instead of preventing them.
Open this question on its own page → - Sample · question 7 · Routing scarce human review capacity
Your contracts-review pipeline flags clauses that may need human escalation. Reviewer capacity is roughly 8% of the daily volume — anything more sits in the queue for days. Which routing rule maximizes the value of that 8% reviewer capacity?
- A.Random sample 8% of all flagged clauses.
- B.Always route the first 8% of the day's clauses; cap at that and skip the rest.
- C.Route the clauses where Claude's confidence is lowest, regardless of severity.
- D.Route the highest-severity clauses where Claude's confidence is below a threshold — high severity AND low confidence.correct
Why: Scarce reviewer time should hit the intersection of "mistake matters most" and "mistake is most likely." That is high severity AND low confidence. Pure low-confidence routing dilutes attention on low-stakes items; random or first-N strategies don't use the model's signal at all.
Open this question on its own page → - Sample · question 8 · Execution-based evals for SQL generation
Your team is adding a new prompt for SQL generation and wants to set up an eval harness before shipping. The team can hand-label about 80 (question, expected-SQL) pairs over the next two weeks. Which eval setup gives the most useful signal?
- A.Score generated SQL with string-equality against the labeled expected SQL.
- B.Execute both the generated and expected SQL against a known dataset and compare result sets.correct
- C.Have Claude score its own generated SQL for correctness on a 1–10 scale.
- D.Track only latency and cost in production; trust the model to be correct.
Why: Two SQL statements can differ in syntax (alias names, join order, whitespace) but produce identical results — and result-correctness is what users care about. String equality is too brittle; self-scoring leaks the eval signal back into the system under test; production-only metrics don't measure correctness.
Open this question on its own page → - Sample · question 9 · Prompt caching for repeated context
A documentation Q&A app sends the same 90 KB knowledge-base prelude with every user query. Cost has grown linearly with query volume. Which is the most appropriate first-pass cost optimization?
- A.Switch to a smaller Claude model and accept some quality regression.
- B.Enable prompt caching on the static prelude so identical preludes don't re-tokenize on every call.correct
- C.Move the knowledge base into a vector store and retrieve only relevant chunks per query.
- D.Compress the knowledge base into a smaller summary that Claude reads each time.
Why: The prelude is stable across queries — prompt caching reduces both input cost and latency on hits while keeping full context. Retrieval is a larger architectural change with new correctness risks; summarization loses information; switching models trades quality you may not need to give up.
Open this question on its own page → - Sample · question 10 · Shadow-mode rollout for prompt changes
You are about to ship a major change to your support-classification prompt. Engineering and product agree the change is right, but no one wants to risk a regression hitting all users. The team has logging on classification accuracy with a roughly one-day lag. Which rollout strategy fits these constraints best?
- A.Replace the prompt all at once, since the team agrees the change is right.
- B.Run the old and new prompts in shadow mode (both run, only the old is acted on) for one day, compare metrics, then cut over.correct
- C.Make the prompt configurable per user and let users opt in.
- D.Run the new prompt on 50% of traffic immediately; if metrics look fine, expand to 100% the next day.
Why: Shadow mode is the lowest-risk option: no user is exposed to the new prompt's outputs until you have confirmed metrics match, which fits the one-day measurement window. Canary or 50% rollouts expose real users to potential regressions before metrics arrive; opt-in selection is biased.
Open this question on its own page →
Like the sample?
Study guides for this exam
Other practice exams
- DatabricksDatabricks Data Engineer Associate100 questions · $19
- SnowflakeSnowPro Core198 questions · $19
- SnowflakeSnowPro Advanced: Data Engineer100 questions · $19
- SnowflakeSnowPro Advanced: Architect100 questions · $19
- AWSAWS Cloud Practitioner (CLF-C02)100 questions · $19
- AWSAWS Solutions Architect Associate (SAA-C03)100 questions · $19
- Microsoft AzureAzure Fundamentals (AZ-900)100 questions · $19
- Microsoft AzureAzure Administrator (AZ-104)100 questions · $19