Practice data

What learners get right and wrong

This is how people score on Cred Farmer’s own practice questions under a timed clock — it is not the real exam, it is not a pass rate, and where the numbers cannot tell two things apart we say so.

  • 746Timed practice papers sat
    Drills and study runs are not among them
  • 397Learners who sat one
    140 of 397 left their first paper unfinished
  • 2,515Answers in the measured set
    1,925 unanswered and 442 repeat encounters removed
  • 30.8%Papers left less than half answered
    230 of 746 papers
  • Read on . The third figure counts only answers that survive every filter below; the first two count everything, filtered or not. Each bar is drawn against the count printed above it, and none of them is an accuracy.

Read this first

What this page is not

  • It is not the real exam. Every figure here comes from our practice bank — questions we wrote, marked against our own answer key. The certification exams are written and delivered by other people, and nothing measured here transfers to them.
  • It is not a pass rate. We have verified outcomes for almost nobody, and we already published the reason we would not quote one even if we did, on our scoring methodology page. Nothing here names a score threshold, and nothing here should be divided by one.
  • It is not a claim about candidates. Everyone in this data chose a preparation site and then chose to sit a timed paper on it. That is a self-selected group and it is not a sample of anything wider.
  • It is not an ordering. The tables are sorted so they can be read, not to rank anything. Almost every gap between two domains here is smaller than the uncertainty around it, and we say which ones are not.

CCDV-F

Claude Certified Developer – Foundations, domain by domain

37 learners, 56 timed papers and 2,515 answers that survive every filter below.

Accuracy by exam domain

Pooled across every domain, 68.4% of answers were correct — 66.54% to 70.18% counting items, and 60.52% to 76.26% once the answers are grouped by the learner who gave them. Where two bands overlap, this data cannot tell those domains apart.

Every pairwise contrast

Every pairwise contrast. 28 contrasts between the 8 published rows. 1 of them separates after correction; the other 27 sit inside the noise, which means the measurement cannot tell those rows apart.

  • Model Selection and Optimization with Security and Safety: this contrast separates after correction.

One cell per pair. Filled survives correction for testing them all; open is inside the noise.

Of the 28 possible comparisons between these domains, exactly one survives once the intervals are grouped by learner and corrected for testing them all: learners answer Model Selection and Optimization about 11 points more accurately than Security and Safety. Every other gap in this table is inside the noise.

The two rows furthest apart, and why both intervals are printed
The two rows furthest apart, and why both intervals are printed
RowAccuracyPer-item intervalLearner-clustered intervalCorrect of answeredLearners
Claude Code59.4%49.37% to 68.66%46.75% to 72.0%57 of 9636
Model Selection and Optimization73.2%68.88% to 77.17%65.44% to 81.02%320 of 43737

Per item these two do not touch — the reading that would tempt you. Group the answers by learner and they overlap, so we do not claim they differ.

Claude Certified Developer – Foundations accuracy by exam domain, on timed practice papers taken on the current question bank.
DomainAnswered correctly95% interval per item95% interval clustered by learnerCorrect of answeredLearners
3Claude Code59.4%49.37% to 68.66%46.75% to 72.0%57 of 9636
6Prompt and Context Engineering61.8%55.72% to 67.49%49.7% to 73.85%160 of 25936
7Security and Safety62.5%55.61% to 68.92%52.13% to 72.87%125 of 20036
4Eval, Testing, and Debugging62.7%51.35% to 72.75%49.5% to 75.84%47 of 7535
8Tools and MCPs65.0%59.12% to 70.52%53.4% to 76.68%173 of 26636
2Applications and Integration70.0%66.7% to 73.01%62.97% to 76.93%568 of 81236
1Agents and Workflows73.0%68.22% to 77.25%65.03% to 80.92%270 of 37036
5Model Selection and Optimization73.2%68.88% to 77.17%65.44% to 81.02%320 of 43737
CCDV-F detail by objective — 12 shown, 13 withheldAn objective appears here once at least 20 learners have answered at least 100 of its questions. The rest are not shown, rather than shown small.
Claude Certified Developer – Foundations accuracy by exam objective, on timed practice papers taken on the current question bank.
ObjectiveAnswered correctly95% interval per itemCorrect of answeredLearners
O01Agent Architecture76.9%68.5% to 83.63%90 of 11736
O02Agent Construction with Claude64.7%55.78% to 72.71%77 of 11936
O03Agent Patterns and Frameworks76.9%69.03% to 83.2%103 of 13436
O06Claude API Mechanics81.6%75.19% to 86.67%142 of 17436
O07Software Engineering Foundations78.0%71.62% to 83.2%152 of 19536
O08Claude Application Design67.0%60.01% to 73.35%126 of 18836
O09Configuration Management64.7%55.78% to 72.71%77 of 11936
O12LLM Fundamentals78.8%71.25% to 84.84%108 of 13737
O13Technical Fundamentals72.5%65.48% to 78.51%129 of 17836
O19AI Application Security66.4%56.97% to 74.61%71 of 10733
O23Tool Implementation63.5%54.1% to 72.06%68 of 10736
O25Agentic Customization67.0%57.3% to 75.44%67 of 10036

Intervals here are per item only. These cells are far smaller than the domain cells above, and this page makes no claim that any objective differs from any other.

Not yet published

The other certifications

Same computation, same code path, not enough learners yet. A table of its own is earned at 30 distinct learners, and not before.

Learners per certification, against the floor
Learners per certification, against the floor
CertificationLearnersFloorTables published
CCDV-F — Claude Certified Developer – Foundations3730yes
CCAO-F — Claude Certified Associate – Foundations2930no
CCAR-F — Claude Certified Architect – Foundations1330no

Short of the line we publish no accuracy table and no objective detail — not a partial one, not a summary. The count is all this page says.

  • CCAO-FClaude Certified Associate – Foundations

    29 learners have sat a timed paper on the current question bank, across 36 papers. We publish accuracy by domain at 30 learners.

  • CCAR-FClaude Certified Architect – Foundations

    13 learners have sat a timed paper on the current question bank, across 24 papers. We publish accuracy by domain at 30 learners.

A certification with fewer than 5 learners is left off this page altogether, including its learner count. At that size a domain profile is close enough to one identifiable person’s results that publishing the shape of it — even as a card saying how few there are — is not something we are willing to do.

Hundreds of learners, not dozens

How people behave under the clock

Far more data than the accuracy tables, because these depend on what a learner did rather than on which question they were served.

Most papers that go wrong are never finished

30.8%230 of 746 timed papers, 27.62% to 34.24%

A paper counts as abandoned when fewer than half the questions answered. An unanswered question is marked wrong when a paper is scored, so this is not a low score — it is a paper that was left.

Every timed paper, and the median abandoned one

Every timed paper, and the median abandoned one. 230 of 746 timed papers were left less than half answered — 30.8%, 27.62% to 34.24%. 38 of those were left entirely blank. 140 of 397 learners did so on a first paper — 35.3%, 30.72% to 40.09%. The median abandoned paper had 3 of 60 questions answered, handed in at 10.3 minutes.

No outcome and no advice attached: this is a description of what happens, and the largest single pattern in everything we hold.

The flag button knows something

−27.0ppwithin-learner difference, −32.86pp to −21.22pp, 124 learners

Comparing each learner against themselves — items they flagged against items they did not — flagged items come back markedly less accurate. The difference spans every practice answer, not timed papers alone: the button is used far more often away from the clock.

The within-learner difference

The within-learner difference. flagged minus unflagged: −27.0pp, −32.86pp to −21.22pp, over 124 learners. Measured on 195 of 1,366 sittings — sittings where the flag button was used at all.

The two levels, on the current bank
The two levels, on the current bank
RowAccuracyPer-item intervalLearner-clustered intervalCorrect of answered
Flagged for review52.4%not publishednot published204 of 389
Not flagged68.0%not publishednot published2,504 of 3,683

Absolute levels are published on the current bank only, where they mean something. No interval is published for either, so none is drawn.

Only 195 of 1,366 practice sittings carry a flag at all, so this describes people who use the button. Read it as evidence that the button works — that learners can tell which of their own answers are shaky — and not as evidence that flagging changes anything.

People are less right than they feel when they are certain

Accuracy by what the learner said about the answer
  • Sure of the answer and Fairly confident: these two separate once answers are clustered inside learners.
  • Fairly confident and Guessing: these two sit inside the noise once answers are clustered inside learners, so the measurement cannot tell them apart.

A bracket says whether that step survives clustering the answers inside learners. A dashed one does not.

Accuracy by the learner's own confidence rating, recorded against each question before the paper was marked.
What the learner saidAnswered correctly95% interval per item95% interval clustered by learnerCorrect of answeredLearners
Sure of the answer78.3%71.44% to 83.91%69.85% to 86.77%130 of 16635
Fairly confident55.7%50.12% to 61.21%51.03% to 60.45%170 of 30552
Guessing43.0%35.85% to 50.5%33.48% to 52.57%74 of 17255

answers tagged “Sure of the answer” come back more accurate than answers tagged “Fairly confident”. “Fairly confident” and “Guessing” have overlapping intervals and are not told apart by this data.

Every self-rated answer on the current question bank, in any mode — a wider population than the tables above, and a small self-selected one: only learners who rate their own answers are in it. We publish it because our readiness model assumes far more of a “sure” than this measures.

Methodology

What we counted, and what we threw out

Six filters, each applied to whatever the one before it left. Nothing on this page was typed in by a person.

Six filters, in the order they are applied

The first three leave 340 timed papers on the current question bank

  1. Our own accounts2 papers

    Why

    We test the product on the product. Those runs are not learners, so they go first.

  2. Papers sat on the retired question bank675 papers

    Why

    The bank in use until the August release was far too easy to measure anything with. It is more than half of everything we hold, and it appears on this page only in the section that explains why it was dropped.

  3. Drills and untimed study papers351 papers

    Why

    Drills deliberately re-serve questions the learner previously got wrong, so drill accuracy measures the queue rather than the learner. Study mode has no clock, which makes it an open-book instrument. Three instruments reported as one number is not a measurement.

  4. Papers whose question bank moved underneath them214 papers

    Why

    For every remaining paper we rebuild, from today’s bank, exactly which domains and objectives it presented, and require an exact match with what was recorded at the time. Where it does not match, the question at a given position is no longer the question that was served, and the paper is dropped rather than guessed at.

  5. Questions that were never answered1,925 answers

    Why

    An unanswered question is marked wrong when a paper is scored, which is right for a score and wrong for a measurement. Every figure above counts only questions a learner actually completed.

  6. Second and later encounters with the same question442 answers

    Why

    Only a learner’s first answer to each question counts. Repeat sittings score several points above first sittings, which is memory rather than learning, and pooling them would flatter every figure on this page.

The last three leave 2,515 answers

Bars are scaled within their unit — papers against papers, answers against answers — never across the two. A step that removed nothing shows no bar.

Why there are two intervals

Answers are not independent of one another. They arrive in papers, and papers belong to learners, so one learner having a bad afternoon moves many answers at once. The per-item interval a reader expects — and which every published figure here carries — ignores that, and is measurably too narrow as a result. Grouping the answers by learner gives an interval several times wider, and that wider one is the honest measure of how much this sample can settle. Both are printed at the same weight, in the same row, so nobody has to take our word for which is which.

Interval bounds are printed at the precision they were computed to, and rounded outward rather than to the nearest tenth. Rounding a bound inward would publish a narrower claim than the data supports.

A learner is an account, not a person

Every count on this page counts accounts. Signing in with two providers makes two of them, and we do not link accounts across providers, so every learner figure here is an upper bound on how many distinct people are behind it. That is one of two reasons the threshold for publishing a table sits as high as 30.

The measured answers cover a short window

The banks were edited in place without a version change, so an answer recorded before that edit can no longer be matched to the question that was actually served. The fourth filter above drops those, and it drops a great many of them — which means the accuracy tables describe a narrow and recent slice of history rather than everything we hold. That is a real limit, and we would rather print it than leave it to be inferred.

The retired question bank

Why we threw away more than half the data

The only section where figures from the retired bank appear at all — the clearest evidence we have that a question bank, and not a learner, produces most of what looks like a finding.

Take one domain, “Prompting and Task Execution”, and put all 7 domains of that certification in order of accuracy. On the retired bank it sat in sixth place counting up from the lowest. On the current bank it sits in first place, 40 points lower than it was.

On the retired bank
92.7%90.84% to 94.18% · 874 of 943
On the current bank
52.9%45.8% to 59.9% · 100 of 189

Same platform, same domain, same kind of learner, opposite conclusion. Nothing about the people changed between those two figures. The questions did — and that is the entire argument for excluding the 2026-07-23 bank, and the 406 timed papers sat on it, from every accuracy figure on this page.

What these two figures do not license

Neither of those two orderings is one we would publish on its own. The gaps inside each of them are smaller than the uncertainty around them, and the certification this domain belongs to may not have enough learners to earn a table anywhere on this page. The movement between the two banks is the part that is not in doubt: the two intervals above do not come close to meeting.

A second symptom of the same problem. On the retired bank, of 209 papers with no unanswered questions, 209 cleared our own scoring threshold, the lowest of them scoring 745 on a thousand-point scale. An instrument that separates nobody is not measuring anything.

The same release also renumbered domains for CCAR-F, so papers sat before it carry domain numbers that no longer exist. There is no archive of the pre-release tree, which is why those papers cannot simply be remapped.

An earlier internal draft of this page read the first of those figures as a fact about people and drew the opposite conclusion to the one above. It was a property of the questions. This section is here so that mistake stays on the record instead of in the archive.

The negative results

What we measured and would not publish

Four things we computed, looked at, and left out — because a reader cannot otherwise tell a claim we cannot support from one we never thought to make.

Whether studying first improves your first paper
The differences do not survive on the current bank, and the striking version comes from the retired one.
Why

On the current bank the group with no study sections behind them and the group with several are not distinguishable, and the ordering is not even consistent. The striking version of this — a large gap in first-paper scores — comes almost entirely from the retired bank and from unfinished papers, and quoting a score against a threshold would be a back door to exactly the claim we refuse above.

A list of the concepts people get wrong
Structurally impossible rather than merely underpowered.
Why

Each question carries a concept label, and those labels are unique per question — one distinct value for every item in the bank. Any list built from them is a list of individual questions wearing a costume, so we do not build one.

How many learners pass
Off the table permanently, not withheld out of modesty.
Why

This is not a pass rate we are withholding out of modesty. It is off the table permanently, for the reason already published on our methodology page. Note the trap: the retired bank would have supported a spectacular-sounding version of it. That sentence would have been true, checkable, and a fact about a broken question bank.

Anything at all about the certification exams themselves
Our questions, our answer key, our self-selected learners.
Why

Every sentence on this page is scoped to our practice bank, and where you see a domain name it describes how our items for that domain were answered — nothing further.

We use cookies and privacy-friendly analytics to understand usage and improve Cred Farmer. Essential features work either way. See our Cookie Policy.