Claude Certified Architect — Foundations · Free practice question 9 of 10
Prompt caching for repeated context
A documentation Q&A app sends the same 90 KB knowledge-base prelude with every user query. Cost has grown linearly with query volume. Which is the most appropriate first-pass cost optimization?
- A.Switch to a smaller Claude model and accept some quality regression.
- B.Enable prompt caching on the static prelude so identical preludes don't re-tokenize on every call.
- C.Move the knowledge base into a vector store and retrieve only relevant chunks per query.
- D.Compress the knowledge base into a smaller summary that Claude reads each time.
Show answer and explanation
Correct answer: B. Enable prompt caching on the static prelude so identical preludes don't re-tokenize on every call.
Why: The prelude is stable across queries — prompt caching reduces both input cost and latency on hits while keeping full context. Retrieval is a larger architectural change with new correctness risks; summarization loses information; switching models trades quality you may not need to give up.
More free Claude Certified Architect — Foundations questions
- Narrowing label definitions to stop over-triggering
- Severity and confidence metadata for review findings
- Tool-based verification to prevent hallucinated APIs
- Precision vs recall trade-offs in extraction
- Evidence fields for auditable structured output
- Tool schemas for strict JSON conformance
- Routing scarce human review capacity
- Execution-based evals for SQL generation
- Shadow-mode rollout for prompt changes