Claude Certification Blog

Which Claude model to use: Haiku, Sonnet or Opus

Deciding which Claude model to use is examined as a method, not as a recall question. The exams grade the evidence behind the choice — which is also why the answer survives the lineup changing underneath it.

12% of CCAO-F4 objectives, 1 is modelsTiers, not versions

7 min read

Classify the task, write down the quality bar it has to clear, pilot the lightest tier that could plausibly clear it, and escalate only when a measured result misses that bar. That you should not reflexively reach for the largest model is the half everyone already knows — it is one of the shortcuts that recur across all four exams. What the exams actually score is the half nobody writes down: the evidence behind the choice.

Which Claude model to use, in one pass

The decision is a constrained optimisation, not a ranking. Fix a minimum acceptable quality in terms specific to the task — extraction accuracy, required fields present, citations that hold, a reviewer accepting the draft — and then minimise cost and latency subject to that bar. Ranking the tiers against each other in the abstract answers a question nobody asked.

The order matters and it runs upward. Start at the lightest tier that could plausibly work and buy capability on measured failure, rather than starting at the top and negotiating down on price.

Models are one objective of four

Worth noticing, because it changes what to revise. The Associate domain is called Product and Model Selection, and its four published objectives are product features, model types, aligning selection to task requirements, and context limits — when to restart, summarise or persist.

Model types is one of four published objectives in the Product and Model Selection domainPRODUCT AND MODEL SELECTIONProduct featuresModel typesAligning selection to the taskContext limits and memoryONLY ONE OF THE FOUR IS ABOUT MODELS
Three of the four are about what surrounds the model, which is where most of the marks are.

So a scenario that opens with a model question is often settled by a feature, a source or a context decision instead. Choosing a project over a one-off chat, or resolving what happens when a conversation outgrows its working set, are separately examinable — and neither is fixed by a heavier tier. The tier names themselves come from that objective, which lists Haiku, Sonnet and Opus as examples of capability tiers rather than as a fixed catalogue.

What actually justifies moving up a tier

Three questions, in order, and a no at any of them means the answer is not a bigger model.

Three checks before escalating a model tier, each with what a no meansTHE OUTPUT MISSED THE BARWas the bar written down first?No → Not an escalation decision yetIs the gap capability, or context?No → Fix the prompt or the sourcesDoes the heavier tier measurably clear it?No → You have bought cost, not qualityTHREE YESES BEFORE YOU PAY MORE
The first gate catches the commonest failure, which is deciding the bar after seeing the result.

That first gate is worth dwelling on. If the threshold gets defined after the cheaper output comes back, it will be defined to fit what came back — the bar quietly moves to meet the result instead of the result having to meet the bar. Written down first, it is a test; written down afterwards, it is a justification.

The second gate matters just as much. A heavier tier buys reasoning capability. It cannot supply a source nobody provided, resolve two documents that contradict each other, or repair an instruction that never said what the output was for.

All three tiers are traps if you pick by reflex

Always reaching for the most capable is a named failure — but so is choosing the lightest because the output is short, and so is assuming the middle is safe without evaluating it. Short output is not the same as a simple task. The balanced-sounding option is not automatically the defensible one.

Cheaper is not cheaper if someone has to fix it

Per-request price is never the unit of comparison. The comparison that counts runs end to end: retries, reviewer time, rework, escalations, and the cost of an error that reaches someone. A lighter tier that halves generation time and doubles the reviewing is more expensive, not less, and the exam expects you to notice.

Latency has the same trap in a different shape. Optimising average response time while ignoring the tail means the deadline that actually matters is the one you miss. Averages are not commitments.

One workflow, more than one tier

The reframe that makes the answer better than the question: the unit of decision is the step, not the application — though what shape those steps form is decided separately. Asked which model a system should use, the defensible answer is frequently "more than one".

ExamWhere the decision is examinedWeight
CCAO-FDomain 3 · Product and Model Selection12%, about 7 of 60 items
CCDV-FDomain 1 · LLM and Claude Platform Foundations17%, about 9 of 53 items
CCAR-FDomain 1 · Product Foundations, and Domain 2 · API and Model Fundamentals12% and 19% of 60 items

One pipeline can route extraction to a light tier, synthesis to a heavier one, and a final check to a qualified human — and the ladder does not stop at the largest model, because a reviewer is the last tier. How those steps actually hand work to each other is a contract in its own right. A scenario offering one model for the whole system is usually offering the wrong shape of answer. Both Claude Certified Associate – Foundations and Claude Certified Architect – Foundations examine this, from different ends.

Why this survives the lineup changing

The published objective names those tiers as examples, and the exam guides say plainly that lineups, surfaces and feature availability change frequently. That is not a caveat bolted onto the exam — it is why the exam is written the way it is. Items turn on relative fit, so they keep working when a new tier ships and an old one is retired.

Which means the preparation that pays is the method, not the spec sheet. Verify current specifics against official Anthropic documentation close to your test date rather than relying on any guide, including this one. How the rest of the blueprint is weighted is set out in what the Claude exam actually tests, and the Claude certification study guide covers how to build practice around a domain like this one. If you are still choosing an exam, the guide to choosing a Claude certification compares the four.

Key takeaways

  • Set the quality bar before you look at a tier. Defined afterwards, it will be defined to fit the result you already have.
  • Escalate on measured failure, not on nerves — and only when the gap is capability rather than a prompt, a source or missing context.
  • Every tier is a trap when chosen by reflex, including the middle one. Short output is not the same as a simple task.
  • Compare end to end. Retries, reviewer time and rework decide which option was actually cheaper.
  • Choose per step, not per system, and remember the last tier in the ladder is a qualified human.

Practise the decision, not the spec sheet

Timed mock exams scored per domain, with the reason each wrong option fails — so you can tell whether your selection reasoning holds up on scenarios you have not seen, which is the only part of this that a changing lineup cannot invalidate.

Start practising free

Questions

Frequently asked

The follow-up questions people search next.

Which Claude model should I use for a task?

Classify the task, write down the quality bar it has to clear, then pilot the lightest tier that could plausibly clear it and measure. Escalate only when the measured result misses the bar you set beforehand. The answer is a configuration, not a favourite.

Do I need to memorise Claude model specs for the exam?

No, and trying to is the wrong preparation. The published objective names model types as capability tiers rather than versions, precisely because lineups change. Items are written so the reasoning decides the answer, not the spec sheet.

What is the difference between Haiku, Sonnet and Opus on the exam?

They function as relative positions — a fast, cost-efficient end, a general-purpose middle, and a high-capability end — rather than as facts to recall. Assuming the middle is automatically safe is a named trap, as is always reaching for the top.

How much of the Claude certification is model selection?

On the Associate exam it is Domain 3 at roughly 12% of the blueprint, about 7 of 60 items, and models are one of that domain’s four published objectives. It also appears on the Developer and Architect Foundations exams under platform fundamentals.

When is it right to move up to a more capable Claude model?

When a measured result misses a threshold you defined in advance, and the gap is genuinely capability rather than an unclear prompt, missing context or an absent source. A heavier tier cannot manufacture evidence that was never supplied.

Keep reading

Related posts

Not affiliated with, or endorsed by, Anthropic or Pearson VUE. Details are summarised from publicly published program information and can change — always confirm against the official exam guide before booking.

We use cookies and privacy-friendly analytics to understand usage and improve Cred Farmer. Essential features work either way. See our Cookie Policy.