How we measure accuracy
Every number on our landing page comes from the measurement below — the papers, the scores and the misses, including our own.
The answer set: published, examiner-scored Cambridge 19 answers plus IDP’s published Band 9 sample answers. Every one is public — Cambridge 19 is in any bookstore with the examiner’s scores printed inside, and IDP publishes its Band 9 samples on its own site.
The protocol: every system received one identical grading prompt — same text, same format, no coaching — and graded every answer three times on 9 Aug 2026. Every number is that system’s median; ours run through the same scoring path our users get.
We publish misses as well as matches: our average miss is 0.22 of a band, we read the 8.0 answer 7.5 while Gemini hit it exactly, and no system — ours included — reaches IDP’s published 9. That is what honest measurement looks like, and we say so instead of promising “±0.5”.
Not one lucky essay.
Across all nine examiner-scored answers — our average miss: 0.22 of a band. Gemini 0.72 · ChatGPT 0.94 · DeepSeek 0.94 · Claude 1.50. Nine of nine we landed within half a band, with zero bias either way — every chatbot reads low.
Average miss vs the examiner’s band — smaller is better.
Worst single miss: IELTS Ace 0.5 · ChatGPT 1.0 · DeepSeek 1.5 · Gemini 2.0 · Claude 2.5 bands.
The papers, one by one
| Paper | Examiner band | ChatGPT | Gemini | DeepSeek | Claude | IELTS Ace |
|---|---|---|---|---|---|---|
| Cambridge 19 · Test 2 · Task 1 | 4.0 | 3.0 (−1.0) | 2.0 (−2.0) | 3.0 (−1.0) | 1.5 (−2.5) | 4.5 (+0.5) |
| Cambridge 19 · Test 1 · Task 1 | 5.5 | 5.0 (−0.5) | 5.0 (−0.5) | 5.5 (0.0) | 4.5 (−1.0) | 5.5 ✓ |
| Cambridge 19 · Test 1 · Task 1 | 6.0 | 5.0 (−1.0) | 5.0 (−1.0) | 4.5 (−1.5) | 4.5 (−1.5) | 6.0 ✓ |
| Cambridge 19 · Test 2 · Task 2 | 7.0 | 6.0 (−1.0) | 6.5 (−0.5) | 5.5 (−1.5) | 5.5 (−1.5) | 7.0 ✓ |
| Cambridge 19 · Test 1 · Task 2 | 8.0 | 6.5 (−1.5) | 8.0 (0.0) | 6.5 (−1.5) | 6.5 (−1.5) | 7.5 (−0.5) |
| IDP · published Band 9 sample · Task 1 | 9.0 | 7.0 (−2.0) | 8.0 (−1.0) | 6.5 (−2.5) | 5.5 (−3.5) | 8.0 (−1.0) |
The four answers shown on the landing page. The full nine-paper table with raw outputs ships with the symmetric re-run.
← Back to the landing page