Claude Certified Architect prep: 7 design patterns that keep coming up
· 3 min read
Scenario questions about building with Claude share a shape: a system misbehaves, four fixes are offered, and the tempting ones treat the symptom. These seven patterns explain most of the "best answer" choices in our practice bank.
CertCram is an independent practice-exam publisher and isn't affiliated with Anthropic.
1. Define criteria instead of asking for "care" or "confidence"
When a model over-flags or mislabels, the usual root cause is an under-specified instruction. "Only report issues you're confident about" or "be careful" gives the model nothing concrete to apply. Explicit criteria do: report bugs and security issues; skip naming preferences and patterns already used elsewhere in the codebase.
Tell-tale distractors: adding "think carefully", lowering temperature, or bolting on a second model to filter the output.
2. Keep recall at the model layer, filter downstream
To cut noise without losing real findings, have the model emit everything it finds, tagged with fields such as severity and confidence, and apply thresholds in code. Self-filtering inside the prompt ("only report high-severity issues") silently discards information you can never recover or analyze later.
3. Make verification cheaper than guessing
Hallucinated function names, wrong file paths and invented APIs come from the model filling gaps with plausible text. The durable fix is to give it a way to check — a tool that looks up a symbol, reads a file or searches the codebase — and to tell it to verify before asserting. Pasting a static snapshot of a whole library into the prompt goes stale and bloats context; catching errors in a later pass lets the failure happen first.
4. Use tool schemas for strict structured output
When a downstream system rejects anything that doesn't match a JSON schema, define a tool whose input schema is that JSON schema and have Claude call it. That's more reliable than describing the schema in prose, pre-filling an opening brace, or repairing malformed JSON afterwards.
5. Evaluate outcomes, not strings
Good evals measure what users care about. For SQL generation, execute the generated and reference queries and compare result sets — two queries can differ textually and still be equivalent. Don't use the model under test as its own primary grader, and don't judge correctness from production latency and cost alone.
6. Spend human review where severity meets uncertainty
Reviewer time is scarce. Route items that are both high-impact and low-confidence — that's where mistakes are most likely and most expensive. Random sampling is useful for measuring quality, but it isn't a routing strategy, and routing on confidence alone wastes reviewers on low-stakes items.
7. Ship prompt changes like code
- Shadow mode: run the new prompt alongside the old one, act only on the old output, compare metrics, then cut over.
- Respect the stakeholder's error preference: if missing an entity is worse than an extra one, a prompt that raises precision but drops recall is a regression, whatever the headline accuracy says.
- Cache stable context: when every request repeats a long, unchanging prefix — a knowledge base, instructions, examples — prompt caching cuts cost and latency without changing behavior.
Using the patterns on exam day
For each option, ask:
- Does it fix the root cause, or mask the symptom?
- Does it preserve information, or throw it away early?
- Is it the simplest change that works before adding infrastructure?
Options that add a second model, a vague instruction or an after-the-fact repair step are usually distractors when a more direct fix exists.
Work through 10 free Claude Certified Architect practice questions — each includes a written explanation of why the best answer wins.