Claude Certification Blog

Product and model selection on CCAO-F: the lightest configuration that clears the bar

Product and model selection on CCAO-F is the domain where the tempting answer is the most capable one and the credited answer is the lightest one that clears a bar you wrote before you looked. Four objectives, about seven items, one habit to unlearn.

12% of the paperAbout 7 of 60 itemsFour objectives

8 min read

Product and model selection on CCAO-F carries an approximate 12% weight, roughly seven of the 60 items, across four objectives: choosing among product features such as Projects, research and artifacts, telling the Haiku, Sonnet and Opus model types apart, aligning a model with a task's cost, speed and quality requirements, and managing context limits by knowing when to restart, summarise or persist. One rule sits under all four: pick the lightest configuration that reliably meets a defined quality threshold, and step up only when measured results say it was missed.

12% approximate weight
7 of 60 items at that weight
4 published objectives
3 model tiers named in the blueprint

Product and model selection on CCAO-F: four objectives, seven items

The ids are Cred Farmer's, the subjects paraphrase the blueprint, and the last column is the official documentation behind each objective. The full blueprint is on the CCAO-F certification page.

ObjectivePublished subjectWhat it turns onOfficial reading
O11Product features: Projects, research, chat, artifactsThe lightest surface the shape of the work calls forWhat are Projects
O12Model types: Haiku, Sonnet, OpusTiers of capability, cost and speed, not version numbersModels overview
O13Aligning a model with cost, speed and qualityA quality bar written first, then the cheapest tier that clears itChoosing the right model
O14Context limits and memoryWhen to restart, when to summarise, when to persistUsage and length limits

Source: Section 6 of the official CCAO-F exam guide, Version 1.0, effective July 2026, mirrored in the bank's objectives file. The second of the guide's three illustrative items belongs here: high-volume short customer-reply drafts where speed and cost outweigh deep reasoning, credited to a faster, lower-cost model. The rationale treats defaulting to the top model as wasted cost and latency, and reserves it for complex reasoning.

Three tiers, one direction of travel

The blueprint names Haiku, Sonnet and Opus as types, not versions. Read them as tiers: speed and cost efficiency for straightforward, high-volume work; the general-purpose balance; the highest capability for demanding reasoning at higher cost or latency. Which numbered model currently sits in each tier is dated in which Claude model to use; the exam judgment does not change with it.

Anthropic's own selection guide documents two legitimate starting points, efficiency-first and capability-first, and both end in the same place: a candidate tested against representative cases and a stated requirement. Starting at the top because the task is important is not defensible, and neither is staying light because the bar was never written down.

Write the bar, pilot the lightest plausible tier, measure, then stop or step up one tier and pilot againESCALATE ON EVIDENCE ONLYWrite the quality bar firstaccuracy, completeness, reviewer acceptancePilot the lightest plausible tierrepresentative cases, not the demoMeasure against the barquality, end-to-end cost, latencyClears itstop hereMisses itstep up, pilot again
The right-hand exit loops back to the pilot. The bar itself never moves.

Cost is the whole workflow

The third objective says cost, speed and quality, and the trap is reading cost as the price per request. On this exam the cost of a configuration includes the reviewer minutes it consumes, the rework it causes, the retries it needs and the damage a wrong answer does. A faster tier that saves seconds per item and doubles the corrections a human makes is slower end to end. The arithmetic runs the other way too: the most capable tier on deterministic field extraction adds cost and latency with no measurable quality gain.

Features and context: the two non-model objectives

The feature objective names Projects, research mode, chat and artifacts, and the discriminator is the shape of the work, not the richness of the surface. A one-off question needs a conversation. Current external evidence calls for research; material already uploaded does not. A work product to be revised and handed on wants an artifact; recurring work with shared instructions wants a Project. Which surface suits which job is covered in Projects, artifacts and connectors on CCAO-F.

The context objective names three moves that answer three different situations. Restart when the task has changed or the history now misleads. Summarise when the same work continues but the detail is too much to carry. Persist when material will be needed by a conversation that has not happened yet. The recognisable wrong answer appends one more correction to a conversation already holding three budget figures.

Restart, summarise or persist, and the situation each move answersTHREE MOVES, THREE SITUATIONSRestartnew task, or a history that now misleadsSummarisesame work, too much history to carryPersist in a Projectmaterial a future conversation will need
A summary can drop a caveat or a number, so validate the carried-forward state before relying on it.

The traps, and the tell in each

Distractors in this domain are extremes without evidence: always the top model, always the cheapest, a threshold that quietly moves. Five recur often enough to learn.

TrapThe tellThe correction
The top tier by defaultThe most capable model chosen because the task matters, with no measured gap.Pilot the lightest plausible tier and step up on evidence.
The cheapest tier because the output is shortBrevity of the answer taken as simplicity of the reasoning.Judge complexity by ambiguity and error cost, not by length.
Per-request price as the whole costTwo configurations compared on token price alone.Add review time, rework, retries and the cost of a wrong answer.
The bar lowered after the resultsA cheaper tier misses the threshold, so the threshold moves.Fix the bar before the comparison and keep it there.
A feature for its own sakeResearch for a file already uploaded; a Project for a one-off question.Select a feature only where a requirement calls for it.

The bar has to exist before the comparison does

Neither starting light nor starting capable is the answer. The exam scores whether the choice was measured against a requirement written down in advance. An option with no bar in it is a preference.

A worked Domain 3 item

Written for this article: one defensible answer, and a named reason each of the others fails. Decide before reading the key.

Written for this article · single response

A support operations lead wants Claude to sort three thousand inbound tickets a week into twelve categories. Before piloting, the team set a bar of 95% agreement with human labels. A two-week pilot on the fastest, lowest-cost tier scored 96% overall, but nearly all of its misses fell in one category, formal complaints with legal exposure, which the team had said must never be missed. What is the BEST next step?

  • A. Move the whole workflow to the most capable tier, since misses in a legal category are serious.
  • B. Keep the fast tier for the bulk, route anything showing signals of the complaint category to a more capable tier or a human check, and re-measure recall on that category before scaling.
  • C. Accept the result and scale, since 96% clears the 95% bar the team set.
  • D. Trim the prompt to lower the fast tier's cost further, then rerun the pilot.

Answer: B

The controlling fact is a category that must never be missed, hidden inside an average that looks fine. A is the top tier by default: nothing shows the tier caused the misses, and the bulk already clears the bar. C lets the average stand in for the requirement. D optimises cost before the quality gap is closed. B tiers by risk, spends capability only where the pilot showed it was needed, and measures again before scaling.

How to practise Domain 3

Seven items is a small domain to lose marks in, and the free timed CCAO-F mock exam filtered to Domain 3 shows which of the five traps you fall for, because every option carries a reason it fails. Practice data is not exam data: a Domain 3 percentage counts what you got right on those items, and Anthropic publishes no conversion from a raw percentage to the scaled score with its 720 cut.

Still choosing a track? The guide to choosing a Claude certification compares the four scopes, and the developer paper's version of this material is covered in model selection and optimization on CCDV-F. The Claude certification study guide shows where a domain lesson fits in a plan for the Claude Certified Associate Foundations exam. Checked against the official sources on 17 September 2026.

Key takeaways

  • Twelve percent, four objectives, about seven items. Features, model types, alignment with cost, speed and quality, and context limits.
  • Haiku, Sonnet and Opus are tiers, not versions. Learn the direction of travel.
  • Write the bar, pilot light, escalate on evidence. The bar never moves after the results come in.
  • Cost is the whole workflow. Reviewer minutes, rework and the price of a wrong answer count on both sides.
  • Features follow the work; context moves answer situations. Restart, summarise and persist are three answers, not three strengths of one.

Seven items where the tempting answer is the wrong one

The free CCAO-F mock runs on the 60-item allocation with no account, and its Domain 3 filter puts the selection items in front of you with a reason every option fails. Practice evidence, not a verdict.

Practise Domain 3 free

Questions

Frequently asked

The follow-up questions people search next.

What is CCAO-F Domain 3?

CCAO-F Domain 3 is Product and Model Selection, an approximate 12% of the associate paper. Its four objectives cover choosing product features such as Projects, research and artifacts, telling the Haiku, Sonnet and Opus model types apart, aligning a model with cost, speed and quality, and managing context limits.

Do I need to know the difference between CCAO-F Haiku Sonnet Opus tiers?

At the level the blueprint names: Haiku for fast, low-cost, high-volume work, Sonnet as the general-purpose balance, Opus for the most demanding reasoning at higher cost or latency. Version numbers change faster than exam guides, so learn the direction of travel rather than a catalogue.

What do CCAO-F model selection questions look like?

A task with stated volume, stakes and budget, and four configurations. The guide’s own sample is high-volume short customer replies where speed and cost matter more than deep reasoning, credited to a faster, lower-cost model. The pattern rewards matching the tier to a stated requirement.

Is the most capable model ever the right CCAO-F answer?

Yes, when the scenario shows a lighter configuration missed a defined quality threshold, or when ambiguity and error cost are plainly high. What the exam scores against is choosing it by default. The evidence for stepping up has to be in the stem.

Does a Domain 3 practice score tell me I am ready?

No. It is practice data, and Anthropic publishes no conversion from a raw percentage to the scaled score with its 720 cut. Use the domain filter in the free mock to see which of the four objectives you miss and which trap each miss was.

Keep reading

Related posts

Not affiliated with, or endorsed by, Anthropic or Pearson VUE. Details are summarised from publicly published program information and can change. Always confirm against the official exam guide before booking.

We use cookies and privacy-friendly analytics to understand usage and improve Cred Farmer. Essential features work either way. See our Cookie Policy.