Claude Certification Blog

CCAO-F practice questions: ten worked scenarios across the seven Associate domains

CCAO-F practice questions are worth the time only when every option is explained. Here are ten original worked scenarios across all seven Associate domains, weighted like the blueprint, each with its key and a reason every distractor fails.

10 worked itemsSingle and select-two7 official domains

31 min read

These ten CCAO-F practice questions are original Cred Farmer items written to the CCAO-F Exam Guide v1.0 (July 2026): a named organisation, a measured symptom, a constraint, then four options, or five where the item asks you to select two. They cover all seven official domains in rough proportion to weight, every wrong option carries a reason it fails, and none is drawn from a live exam form or from the question bank behind the login. Answer each one cold, then score yourself on the miss map.

The official CCAO-F exam guide publishes three CCAO-F sample questions of its own in Section 8 and says they are not drawn from the live item bank. The Claude certification sample questions hub works one item per code; this post is its Associate spoke. The blueprint and fee are on the Claude Certified Associate - Foundations page, and the guide to Claude certification practice exams covers what official practice material exists.

The ten are spread by official weight, so output evaluation at 21 percent gets two items and troubleshooting at 10 percent gets one; three ask for two answers. Read the stated constraint before the options, decide what it rules out, then compare what is left.

Seven CCAO-F domains by official weight, from 10 to 21 percent, with one or two items from this set in eachOfficial weight · items hereD1 Prompting and Task Execution14% · 1 itemD2 Output Evaluation and Validation21% · 2 itemsD3 Product and Model Selection12% · 1 itemD4 Workflow Integration and Solution Design16% · 2 itemsD5 Configuration and Knowledge Management12% · 1 itemD6 Governance, Risk, and Responsible Use15% · 2 itemsD7 Troubleshooting and Optimization10% · 1 item
Weights from the CCAO-F Exam Guide v1.0, Section 6; the item counts are this post's ten.

How do you match the prompt to the task in front of you?

Domain 1, Prompting and Task Execution, is 14 percent of the blueprint, and its items rarely ask you to write a prompt. They ask which kind of task a team actually has, then which move fits: decomposition for a multi-part request, iteration for a near miss, a different strategy when the task type was misread. The item below turns on that last case.

Question 1 · Written for this post · Domain 1 · single response

The marketing team at Fenwick & Vale, a homeware retailer, asked Claude for the single best campaign idea for a spring launch, with supporting evidence. The reply gave one polished concept and quoted three customer survey statistics that the team has never collected. No research has been commissioned yet, and the team needs a wide field of options for a workshop on Wednesday. How should the team change the prompt?

  • A. Keep the request for one recommended idea but add an instruction to cite a verifiable source for every statistic it uses, so the team can check each figure before the workshop begins.
  • B. Ask for three fully developed campaign concepts, each with a projected percentage uplift in spring revenue, so the workshop can compare them on expected return in a single pass.
  • C. Reframe the task as brainstorming: ask for fifteen distinct concepts across named angles such as price, sustainability and gifting, with no evidence claims, and evaluate a shortlist later.
  • D. Switch the conversation to research mode and ask Claude to find published homeware market statistics first, so that the single recommended idea rests on real figures rather than invented ones.
Reveal the answer and the reason for every option

Answer: C. Choose the prompting strategy for the task type.

A. Demanding citations for statistics about the retailer's own customers, which nobody has collected, invites fabricated references. The framing still asks for one conclusion when the team needs a wide field.

B. Revenue projections for concepts that have never run are invented numbers with a decimal point. The prompt narrows to three ideas before the workshop has explored the space at all.

C. Correct. Brainstorming across named angles gives the workshop breadth, banning evidence claims removes the pressure that produced the invented survey figures, and evaluation is deferred to a step with real criteria.

D. Research mode finds published sources; it does not create the team's own customer data, and the team's problem is idea generation, not market statistics. It keeps the single-idea framing that narrowed the field.

The objective is adapting prompting strategy to task type. The team's need on Wednesday is breadth, which is a brainstorming task, and the prompt they used asked for a single conclusion with evidence, which is an analysis task. Analysis prompts push the model to justify, and when no data exists the justification gets invented, which is exactly what the survey statistics were. Reframing as brainstorming with named angles produces a wide field, explicitly forbids evidence claims so there is nothing to fabricate, and separates the later evaluation step, where the team's own criteria and any research it commissions can be applied. The distractors keep the analysis framing and try to make the evidence problem go away: by demanding citations for figures that do not exist, by asking for revenue projections that would also be invented, or by researching public statistics before the team has even decided what ideas to test.

Which output can you trust, and what shape should it take?

Output Evaluation and Validation is the heaviest domain at 21 percent, so two of the ten sit here. The first tests what an internal contradiction tells you about the process behind a document, and what must happen before anyone edits it. The second is a select-two on output format: two audiences with two next actions need two containers. In both, the tempting wrong answers make the output tidier, not truer.

Question 2 · Written for this post · Domain 2 · single response

Kestrel Components asked Claude to summarise supplier risk for a board paper. The table in the output states that 7 of 22 suppliers are single-sourced, the narrative below it says nearly half are, and one supplier appears under two different countries. The paper goes to the board on Monday and the operations director has asked for a tidy final version tonight. What should the analyst do first?

  • A. Treat the mismatches as a signal that the underlying figures may be wrong, and check the single-source count and each supplier's country against the supplier master data before editing anything.
  • B. Ask Claude to reconcile the table with the narrative and to state clearly which version is correct, then adopt whichever figure it settles on for the final version of the paper.
  • C. Edit the narrative so it matches the table, because tabulated figures are usually computed more carefully than prose, and then correct the duplicated supplier's country entry by hand before formatting.
  • D. Add a note stating that the figures are indicative and subject to confirmation, keep both the table and the narrative exactly as they are, and send the paper to meet tonight's deadline.
Reveal the answer and the reason for every option

Answer: A. Inconsistency is evidence about the source, not the format.

A. Correct. The disagreement between table and narrative shows the figures cannot be trusted as produced, so the check runs against the supplier master data first and editing follows the verified numbers.

B. The model that produced two conflicting figures has no independent access to the truth. Asking it to choose returns a confident answer with the same evidential basis as the original, which is none.

C. Tables are not more reliable than prose when the same unverified process generated both. Reconciling the text to the table tidies the paper while the count itself remains unchecked against any source.

D. A caveat does not stop the board acting on a wrong count, and it leaves a visible contradiction in a governance document. Validation before the deadline is possible; polish before validation is the trap.

Identifying inconsistencies in a response is the objective, and the point of identifying them is what they tell you. When a table and its narrative disagree and a supplier is placed in two countries, the output is showing that at least one figure is wrong and that the process producing the figures is unreliable, so neither the table nor the prose can be trusted by default. The first step is therefore to go back to the authoritative record, the supplier master data, and establish the true count and countries; only then is editing meaningful. The distractors each pick a version without evidence: the model's own adjudication, a preference for tables over prose, or a caveat that lets both versions go to the board. A tidy paper with the wrong single-source count is worse than a late one, and the director's deadline does not change which figures are true.

Question 3 · Written for this post · Domain 2 · select two

An operations analyst at Harborline Bank has Claude produce a list of 40 reconciliation exceptions from the overnight run. The payments team needs to load the list into its case-tracking system, which imports rows with fixed field names, and the chief operating officer wants something she can read in two minutes before an 8 a.m. call. The analyst has thirty minutes. Which two outputs should the analyst ask Claude for? (Select two.)

  • A. A structured table or CSV artifact of all 40 exceptions with consistent field names matching the tracking system's import format, one exception per row.
  • B. A single long chat reply containing every exception in prose, so that neither the payments team nor the chief operating officer has to open a separate file.
  • C. A single slide deck artifact that groups the exceptions by category, intended to serve both the payments team's case loading and the chief operating officer's briefing before the call.
  • D. A short inline summary for the chief operating officer giving the total exposure, the three largest exceptions and what the payments team will do next.
  • E. The complete 40-row table for the chief operating officer as well, so that she works from the same record as the payments team rather than a separate summary that might leave something out.
Reveal the answer and the reason for every option

Answer: A and D. Format follows how each audience will use it.

A. Correct. Structured rows with the importer's field names are what the tracking system can load, and an artifact is the right container for a file the payments team will reuse rather than read.

B. Prose cannot be imported without someone re-keying 40 rows, and a long reply is the opposite of a two-minute read. It serves the analyst's convenience, not either audience's next action.

C. A deck cannot be loaded into a system that imports rows with fixed field names, and a briefing built to double as a data source is neither a two-minute read nor a record. One format for two next actions serves neither.

D. Correct. The chief operating officer's next action is a short call, so the summary gives her the total, the material exceptions and the plan, which is the information she needs in the time she has.

E. Working from the same record is not the same as needing all of it before a call. Forty rows take longer than two minutes to read and bury the total exposure and the three exceptions that matter; the summary carries those.

Organising information and selecting the output format means asking what happens to the output next. The payments team's next action is a system import, so the format is structured data with the field names the importer expects, and an artifact is the right container for a file that will be reused. The chief operating officer's next action is a two-minute read before a call, so the format is a short inline summary with the totals and the exceptions that matter. Two audiences with two next actions need two formats, and thirty minutes is enough for both. The distractors force one format on both: prose that cannot be imported and is too long to read, a slide deck that suits neither a system nor a hurried executive, and the raw table for an audience whose need is the shape of the data, not its volume.

When has one long conversation stopped being the right container?

Product and Model Selection includes an objective on context limits and when to restart, summarise or persist. The signature is a session that re-asks settled questions or contradicts agreed decisions. The item below adds two constraints that separate the plausible remedies: how long the work will run, and who else has to pick it up. Weigh every option against both.

Question 4 · Written for this post · Domain 3 · single response

A paralegal at Draycott Law has spent three hours in one Claude conversation building a 60-page due diligence checklist for an acquisition. Claude has started re-asking questions the paralegal settled an hour earlier and has twice contradicted an agreed exclusion. The matter will run for six more weeks, and two colleagues will need to pick up the work in the paralegal's absence. What should the paralegal do next?

  • A. Continue in the same conversation and paste the earlier decisions back into the chat each time Claude forgets one, so that the full three-hour history stays together in a single place.
  • B. Start a fresh conversation and paste the current 60-page checklist draft in as the opening message, so Claude works from the document itself instead of the three-hour history that has become unreliable.
  • C. Switch the current conversation to the strongest model on offer, since a more capable model will hold the three hours of history more reliably and stop contradicting the agreed exclusions.
  • D. Have Claude write a decision log and open-questions summary, verify it against the draft, then create a Project holding that summary and the current checklist for the team to continue from.
Reveal the answer and the reason for every option

Answer: D. Long context is a design problem, not a model problem.

A. Re-pasting decisions adds to the context that is already too long to be attended to reliably. The history staying in one place is the problem, not the solution, and the colleagues still cannot pick it up.

B. Restarting is right, but the draft is not the state. The agreed exclusions and settled questions live in the conversation, not in the checklist, so they are lost, and a single new conversation still gives the two colleagues nothing to continue from.

C. The failure is accumulated context, not insufficient capability. A stronger model in the same conversation inherits the same three hours and the same contradictions, at higher cost and with no persistence for the team.

D. Correct. A verified decision log turns three hours of history into a bounded, checkable state, and a Project gives the six-week matter and the colleagues a shared place to continue from.

The objective covers context limitations and knowing when to restart, summarise or persist. The symptom, re-asking settled questions and contradicting an agreed exclusion, is the signature of a conversation that has outgrown what the model can attend to reliably. The remedy is to compress the state into a verified summary and move it into a container built for persistence: a Project, because the work runs six weeks and two colleagues must pick it up. Verifying the summary against the draft matters, since a wrong decision log would persist an error. The alternatives each fail on one of the stated constraints. Pasting decisions back grows the very context that is failing. A fresh start with only the draft pasted in discards the decisions, which live in the conversation rather than the checklist, and does nothing for the colleagues. A more capable model in the same conversation still inherits the accumulated history, and the constraint is length and continuity, not reasoning power.

Where does Claude fit in a workflow, and how do you tell leadership?

Workflow Integration and Solution Design is 16 percent of the exam. Its items give an organisation a volume, a mix of task types and a deadline, then ask what an associate should propose. The first item below turns on segmenting a use case by consequence and decision authority. The second is a select-two on communicating value: the benefit and the control that makes it safe belong in the same sentence, and a number the pilot never measured belongs nowhere.

Question 5 · Written for this post · Domain 4 · single response

The customer service head at Marlow Transit Authority asks a business analyst to work out where Claude fits in complaint handling. The authority receives about 9,000 complaints a year: roughly half are information requests, three in ten are refund claims that require a payment action within a regulated timescale, and two in ten report service incidents, some involving passenger safety. The refund queue carries the longest backlog, staff are under pressure, and the head wants a proposal in a week that covers all three categories. Which analysis should the analyst put forward?

  • A. Propose automating replies for all complaint types, since 9,000 complaints a year justifies the investment and staff can review a random sample of the sent replies each week for quality.
  • B. Segment the complaints by structure, consequence and decision authority: Claude drafts information replies for agent review, prepares refund cases without deciding them, and only triages safety incidents.
  • C. Limit Claude to the information requests only, since refunds and incidents involve regulated payment actions and passenger safety, and leave those two categories entirely as they are handled by staff today.
  • D. Build one prompt that handles all three complaint types, launch it to the whole customer service team, and measure after a month which categories need adjustment or more human involvement.
Reveal the answer and the reason for every option

Answer: B. Analyse the use case by segment, not as one problem.

A. Automating refund decisions and safety incident replies puts a payment action and a safety judgement in the model's hands, with sampling after the fact. Volume justifies analysis, not blanket automation.

B. Correct. Each segment is matched to a role Claude can hold safely: drafting under review, augmenting a human decision, or triage only. That is what a requirements analysis is meant to produce.

C. Confining Claude to information requests leaves the refund queue, the backlog the head named, untouched, and fails a brief that asked for all three categories. Claude can prepare a refund case or triage an incident without deciding either.

D. One prompt for three categories with different consequences cannot hold the right controls for each, and launching to the whole team before measuring is the scale-first error. Segmentation is the analysis, not an afterthought.

Applying Claude to analyse requirements and use cases means recognising that 'complaint handling' is three different tasks with different structure, error detectability, consequence and decision authority. Information requests are stable and low consequence, so drafting under agent review is a good fit. Refund claims involve a payment action inside a regulated timescale, so the decision stays with a person and Claude augments the preparation. Safety incidents carry the highest consequence and need human judgement, so Claude's role is confined to triage that a person acts on. The segmented proposal captures value in every category while placing the human where authority cannot be delegated. Limiting Claude to information requests ignores the refund backlog the head named and the brief to cover all three categories, and the other two options treat the categories as identical, which is how a refund gets decided or a safety report gets answered by a template.

Question 6 · Written for this post · Domain 4 · select two

A consultant is preparing a briefing for the leadership team at Ridgeway Community College on using Claude in student services. In a six-week pilot in the enrolment office, first-draft time for email replies fell from an average of 25 minutes to 8, and 12 percent of drafts needed a factual correction about fees or deadlines before sending. Leadership has asked what the college will gain and what could go wrong. Which two points should the briefing make? (Select two.)

  • A. Present the time saving as a headcount reduction the college can bank in the next budget year, since 17 minutes per email across the office is a material sum.
  • B. Report the time saving alongside the 12 percent correction rate and the review step that catches it, so the benefit is stated net of the checking it requires.
  • C. Assure leadership that outputs are accurate because the enrolment office approved the pilot and staff there were satisfied with the drafts they received.
  • D. Commit that the correction rate will fall to zero once the fee schedules and academic calendar are uploaded to a Project, removing the need for review.
  • E. State the limitation that matters here: Claude can produce confident wrong statements about fees and deadlines, so controlling sources and a named reviewer are part of the design.
Reveal the answer and the reason for every option

Answer: B and E. Communicate the benefit and the control together.

A. Draft time is not headcount. The saving has to be netted against review and the office's other work before anyone can claim budget from it, and promising savings that then fail costs credibility.

B. Correct. Stating the benefit and the correction rate together, with the review step that catches errors, gives leadership the true net value and the resourcing the design depends on.

C. Staff satisfaction and the office's approval of the pilot are not evidence of accuracy. The pilot's own data shows that 12 percent of drafts were wrong on facts that matter to students.

D. Better sources lower the correction rate; they do not eliminate confident errors, and a promise of zero removes the review that protects students from a wrong deadline. Overstated certainty is a limitation of its own.

E. Correct. Leadership needs the specific failure mode for this use, wrong fees and deadlines stated fluently, and the controls that address it, so the design is understood as part of the proposal.

Communicating what Claude offers, and where it fails, to stakeholders means giving leadership a picture they can make a decision on: what was measured, what it cost to get, and what the failure mode looks like. The pilot measured a large time saving and a 12 percent factual correction rate, and those two numbers belong together, because the review step that catches the corrections is part of the cost of the benefit. The limitation that matters to a college is confident wrong statements about fees and deadlines, which is why controlling sources and a named reviewer are design features rather than optional extras. The distractors overstate the value or understate the limitation: converting minutes into headcount before any process redesign, treating staff satisfaction as evidence of accuracy, and promising that uploading documents removes the need for review, which no configuration can honestly guarantee.

Which written rules keep a Claude deployment inside its remit?

Configuration and Knowledge Management at 12 percent and Governance, Risk, and Responsible Use at 15 are separate domains whose items fail the same way: a rule that should exist in writing has been replaced by a hope. Item 7 is a Project instruction that exhorts instead of specifying. Items 8 and 9 move the rule outside the tool: who may make a consequential decision, and which route a new tool travels before anyone's data goes into it. Disclosure, emphasis and personal judgement are offered as substitutes, and none is one.

Question 7 · Written for this post · Domain 5 · single response

The customer success team at Nimbus Ledger, an accounting software company, runs a Claude Project whose only instruction is: 'You are a helpful, friendly expert assistant. Be accurate and concise and never make mistakes.' Over a month the drafted replies have varied widely in tone, three promised refunds the team cannot authorise, and several quoted prices that match nothing on the current price list uploaded to the Project. Which rewrite of the instruction addresses these failures?

  • A. Keep the existing instruction and add, in capital letters, that Claude must always be one hundred percent accurate, must double-check every reply before sending, and must never quote an incorrect price to anyone.
  • B. Paste twenty of the team's best past replies into the instruction so that Claude can match their tone and structure, without stating explicitly the rules those replies follow or the sources they used.
  • C. Define the role and audience, name the uploaded price list as the only pricing source, forbid commitments on refunds, credits or timelines, fix a format, and add an escalation rule for anything uncovered.
  • D. Add a line telling Claude to keep a warm, consistent tone and to consult the uploaded documents before replying to any customer, and leave the rest of the existing instruction exactly as it stands.
Reveal the answer and the reason for every option

Answer: C. An effective instruction states sources, limits and format.

A. Telling a model to be accurate is not a rule it can act on, and 'never quote an incorrect price' does not say which source is correct. The failures came from missing information, not missing emphasis.

B. Examples convey tone but not authority. Twenty past replies do not tell Claude that the uploaded list overrides its general knowledge or that refunds are outside its remit, so both failures continue.

C. Correct. Each observed failure is closed by a specific rule: a named pricing source, a prohibition on commitments, a defined audience and format, and an escalation path for anything the material does not cover.

D. Consulting the documents is not the same as answering only from them, so general knowledge still competes with the price list. Nothing says what Claude may not commit to, so refunds keep being promised, and no format or escalation rule is set.

Creating effective system-level instructions means writing the rules the outputs must follow, not exhorting the model to be good. Each failure maps to a missing rule. Tone varied because no audience or format was defined. Refunds were promised because nothing said what Claude may not commit to on the team's behalf. Wrong prices were quoted because the instruction never named the uploaded list as the authority, so the model's general knowledge competed with the source. The rewrite supplies each of those: role and audience, a named pricing source, an explicit prohibition on commitments, a format, and an escalation rule for questions outside the material. Emphasis in capitals adds pressure without information. Examples without stated rules teach style but not the boundaries that were crossed. A line about tone and consulting the documents softens the symptom without naming a source, a limit or a format, so the two costly failures continue.

Question 8 · Written for this post · Domain 6 · single response

The HR director at St Aldric Health, a hospital group, wants to use Claude to screen 600 applications for nursing posts, produce a ranked shortlist of 80, and send automatic rejections to the remaining 520 so that recruiters can focus on interviews. Screening is where recruiters lose most of their week, and the director's stated aim is to cut that time. Hiring decisions at the group are subject to equality law and to an internal policy requiring a documented human decision for every rejection. What is the most appropriate use of Claude here?

  • A. Implement the automated ranking and rejections exactly as proposed, adding a line to each rejection letter stating that an AI tool assisted the decision so that applicants are properly informed.
  • B. Use Claude only to draft the rejection letters once recruiters have completed all 600 screenings manually, keeping the tool away from anything that touches the assessment of candidates.
  • C. Have Claude rank all 600 applicants against the person specification and let a single recruiter approve the complete ranked list in one step before the rejections are released.
  • D. Use Claude to summarise each application against the published criteria as an aid, test the summaries for bias on a sample, and keep shortlisting and every rejection decision with recruiters.
Reveal the answer and the reason for every option

Answer: D. Keep the consequential decision with an accountable person.

A. A disclosure line does not make an automated rejection a documented human decision, and it does nothing about undetected bias in the ranking. Informing applicants is necessary but does not make the use appropriate.

B. This complies with the policy but leaves the screening time, the director's stated aim, untouched: recruiters still read all 600 applications by hand when structured summaries against published criteria would speed that without deciding anything.

C. One approval for 600 outcomes is a rubber stamp. No recruiter can review a full ranking in one step, so the decision is the model's in substance, which the policy and equality law do not permit.

D. Correct. Summaries against published criteria are an aid recruiters can check, bias testing on a sample addresses the hidden failure mode, and every consequential decision remains a documented human one.

Identifying appropriate and inappropriate use cases turns on consequence, detectability of error and where decision authority sits. Rejecting an applicant is a consequential decision with legal exposure, bias risk that is hard to detect after the fact, and an internal policy that requires a documented human decision for each one. Automated rejection is therefore inappropriate regardless of the disclaimer. Summarising applications against published criteria, tested for bias on a sample, is an appropriate augmentation: it makes recruiters faster and more consistent while every shortlist and rejection remains a recruiter's documented decision. Confining Claude to rejection letters satisfies the policy and misses the director's aim, since screening time is untouched. Approving a 600-name ranked list in one step is a human signature on a machine decision, which satisfies the letter of the policy and none of its purpose.

Question 9 · Written for this post · Domain 6 · select two

A dispatcher at Cobalt Freight, a logistics firm, has installed a browser extension that sends the contents of every open tab to a third-party service described as 'powered by Claude' and returns suggested replies. The firm's AI policy lists the approved tools, requires data to be classified before it enters any AI tool, and names a governance owner for approving new ones. The dispatcher's open tabs include customer manifests and driver rosters. Which two actions are consistent with the policy? (Select two.)

  • A. Keep using the extension for tabs that show no customer names, since the service is built on Claude and the policy already approves Claude for internal work.
  • B. Email the extension's vendor to ask whether they retain the data they receive, and continue using it if the vendor confirms that nothing is stored.
  • C. Stop using the extension and do the dispatch work in the Claude workspace the policy lists as approved, classifying the manifest data before it is entered.
  • D. Trial the extension for a further week and report any problems to the governance owner afterwards, so the decision is based on real experience.
  • E. Submit the extension to the governance owner through the approval route the policy describes, and treat it as unapproved until it appears on the list.
Reveal the answer and the reason for every option

Answer: C and E. Follow the policy's list and its route for additions.

A. A tool built on Claude is not the tool the policy approved, and driver rosters and manifests without names are still classified data. The dispatcher's own filtering of tabs is not the policy's classification step.

B. A vendor's assurance about retention is not an approval, and the policy names a governance owner for exactly that assessment. Continuing while the email is unanswered keeps the data flowing meanwhile.

C. Correct. The approved workspace is the tool the policy sanctions, and classifying the manifest data before entry is the step the policy requires, so the dispatch work continues within the rules.

D. A personal trial is the decision the governance owner is meant to make, taken without them, and a week of customer manifests reaches a third party while it runs. Experience is not authorisation.

E. Correct. The policy provides a route for new tools, and using it puts the extension in front of the person with authority to assess the vendor and the data flow, while nothing is used before approval.

Following organisational AI policies means using the tools the policy approves and using its stated route when something new looks useful. The extension is not on the list, it ships whole tabs including customer manifests and driver rosters to a third party, and being 'powered by Claude' says nothing about who receives the data or under what terms. The two consistent actions are to do the work in the approved workspace with the data classified as the policy requires, and to put the extension through the governance owner's approval route, treating it as unapproved until listed. The distractors each substitute a private judgement for the policy's mechanism: the dispatcher's own view of which tabs are safe, a vendor's assurance about retention, or a personal trial. None of those is the governance owner's decision, and the roster and manifest data are already flowing while they happen.

What do you check first when a Project that worked starts failing?

Troubleshooting and Optimization is the smallest domain at 10 percent, and its first objective is diagnosis before repair: locate where the failure enters before changing the prompt, the model or the knowledge. The item below gives you a date, a change and a symptom, and tests whether you connect them by evidence before touching anything.

Question 10 · Written for this post · Domain 7 · single response

For three months a Claude Project at Pellworth Insurance drafted claim acknowledgement letters that handlers accepted 96 percent of the time. Last week acceptance fell to 70 percent, and handlers report letters quoting the wrong excess amount. The only change was that the product team uploaded the 2027 policy wording to the Project alongside the 2026 wording, and both are still needed because existing policyholders remain on the earlier terms until renewal. What should the operations lead do first?

  • A. Rewrite the prompt to instruct Claude to always use the correct excess amount for the policyholder's terms and to double-check every figure carefully before it drafts the acknowledgement letter.
  • B. Trace two failing letters to the document each excess came from, then rename both wordings and add an instruction selecting the wording by the policy's inception date, and re-test on those cases.
  • C. Remove the 2027 policy wording from the Project immediately so the drafts return to the exact configuration that achieved 96 percent acceptance, and revisit adding the new wording at a later, quieter date.
  • D. Move the Project onto the strongest model on offer, since holding two policy wordings and choosing correctly between them is a harder reasoning task than the previous model was handling.
Reveal the answer and the reason for every option

Answer: B. Diagnose the source before touching prompt or model.

A. 'Use the correct excess' is the goal, not a rule. Without a way to identify which wording applies to a policyholder, the instruction gives the model nothing new to act on and the conflict remains.

B. Correct. Tracing the failing letters identifies the source conflict, the renamed documents and the inception-date rule give the model a basis for choosing the right wording, and re-testing on the same cases verifies the fix.

C. This restores the old acceptance rate by making letters for 2027 policyholders wrong instead, which the stem rules out: both wordings are needed. It also skips the diagnosis in favour of a rollback.

D. Whatever its capability, the model cannot tell which wording applies to a policyholder; the missing input is a selection rule, not reasoning power. Switching models adds cost and leaves the ambiguity untouched.

Diagnosing an underperforming output starts by locating where the failure enters: wrong or missing fact, wrong reasoning, wrong form, or inconsistency. A wrong excess amount that appeared the week a second policy wording was uploaded points to a source conflict, and tracing failing letters to the document each figure came from confirms or refutes that in minutes. The fix then follows the diagnosis: because both wordings are legitimately in force, the Project needs clearly named documents and a rule that selects the wording by inception date, and re-testing on the failing cases shows whether the rule works. The distractors change something before knowing what is wrong. Prompt emphasis cannot tell the model which document applies. Removing the new wording restores the old numbers and breaks letters for new policyholders, ignoring the stated constraint. A stronger model still has two indistinguishable documents and no rule for choosing between them.

Where did these CCAO-F practice questions catch you? A miss map by domain

Score the ten, then count misses by domain rather than in total. Ten CCAO-F questions cannot produce a percentage that means anything against a scaled cut of 720, and the score report gives percent-correct by domain, so that is the unit to practise reading. Each row names the items here, the objective to reread if you missed one, and the free mock's untimed filter for that domain.

DomainItems hereYour missesObjective to rereadUntimed drill
D1 Prompting and Task Execution (14%)1__ of 1O04: prompting strategy by task typeDomain 1 untimed
D2 Output Evaluation and Validation (21%)2 and 3__ of 2O06: inconsistencies as a signal about the source; O10: output format for its next useDomain 2 untimed
D3 Product and Model Selection (12%)4__ of 1O14: context limits: restart, summarise or persistDomain 3 untimed
D4 Workflow Integration and Solution Design (16%)5 and 6__ of 2O15: requirements and use cases by segment; O18: value and limitations stated togetherDomain 4 untimed
D5 Configuration and Knowledge Management (12%)7__ of 1O21: system-level instructions: sources, limits, formatDomain 5 untimed
D6 Governance, Risk, and Responsible Use (15%)8 and 9__ of 2O22: appropriate and inappropriate use cases; O24: the AI policy and its approval routeDomain 6 untimed
D7 Troubleshooting and Optimization (10%)10__ of 1O26: diagnose before changing anythingDomain 7 untimed
Count misses by domain: none means move on, one means reread the objective, two means run the domain filter untimedScore the ten itemsthen count misses by domain0 missesMove onThe full mock rechecks it1 missReread the objectiveThen redo the item cold2 missesRun the domain filterUntimed, in the free mock
A Cred Farmer reading of a ten-item score, not an official rule.

Domains with one item here cap at one miss, so treat a miss there as a reread, not a verdict. Two misses in the same domain is the pattern worth acting on: run that domain's filter untimed, then sit the full timed form. The two-week CCAO-F study plan schedules the rereads around two timed sittings, the CCAO-F cheat sheet compresses each domain to its decision rules, and why exam dumps are worth so little explains what a memorised answer letter costs.

Ten items are a sample, not a score

Two misses in the same domain is a pattern; one miss anywhere is noise until the full 60-item form confirms it. Nothing on this page is a scaled score, a pass verdict or a readiness figure, and it does not predict an official result.

Question bank dated . Checked against the CCAO-F Exam Guide v1.0 (July 2026) on .

Cred Farmer is not affiliated with Anthropic; the CCAO-F page records what the official sources say and when they were last read.

Key takeaways

  • Read the constraint before the options. Every item here turns on a stated fact: a deadline, a policy, a change last week.
  • Count misses by domain, not in total. Ten items give no percentage worth quoting against a scaled cut of 720.
  • The tidy answer is often the trap. Polishing, a disclaimer or a model switch each treats a symptom the scenario has already explained.
  • A rule that is not written down is a hope. Instructions, policies and decision authority all fail the same way when they live in one head.
  • Practice evidence is not a pass probability. Only measured answers on aligned timed forms move a Cred Farmer readiness score, and it does not predict an official result.

Sit the whole 60-item form under the clock

The free CCAO-F mock is the full 60 items in 120 minutes, no account, with a reason for every option and a raw count by domain beside the official weight. Signed in, five blueprint-aligned timed forms add 300 more items; only measured answers on those forms move a Cred Farmer readiness score, which does not predict an official result.

Start the free CCAO-F mock

Already scored the mock? Sign in for the five timed forms.

Questions

Frequently asked

The follow-up questions people search next.

Are these real CCAO-F exam questions?

No. All ten were written fresh for this post from the CCAO-F Exam Guide v1.0 and its published objectives; none comes from a live exam form or from the gated Cred Farmer bank. The guide's own three sample items are illustrative too, and it says they are not drawn from the live item bank.

How many questions are on the CCAO-F exam?

The CCAO-F has 60 items in 120 minutes, multiple-choice and multiple-response, and each item states how many responses to select. The result is a pass or fail with a scaled score from 100 to 1,000; the cut is 720. The fee is $99 and the credential is valid for 12 months.

Is there a free CCAO-F mock exam?

Yes. The free Cred Farmer mock at /practice/ccao-f runs the full 60-item form under a 120-minute clock with no account, then shows a raw count by domain and a reason for every option. It never reports a scaled score or a pass verdict, because no practice percentage converts to the 720 cut.

How should I score myself on these ten CCAO-F questions?

Count misses by domain, not in total, and treat a select-two item as a miss unless both picks match the key. One miss in a domain means reread the objective in the miss map and redo the item cold a day later; two means run that domain's untimed filter in the free mock first.

What score do I need on practice questions to pass the CCAO-F?

No practice percentage converts to the exam's scaled cut of 720, and Cred Farmer publishes none. Use practice to find the domains that need work. Signed in, only measured answers on blueprint-aligned timed forms move a readiness score, which is practice evidence; it does not predict an official result.

Keep reading

Related posts

Not affiliated with, or endorsed by, Anthropic or Pearson VUE. Details are summarised from publicly published program information and can change. Always confirm against the official exam guide before booking.

We use cookies and privacy-friendly analytics to understand usage and improve Cred Farmer. Essential features work either way. See our Cookie Policy.