Scorer calibration

How accurate is our AI IELTS Writing score?

Updated September 2026

We do not have a published accuracy figure right now, so we are not going to quote one. Calibration is in progress: the method, the corpus and every number are published on this page the moment a run completes on a large enough set of human-scored essays. Until then, treat every band the checker gives you as an estimate to work from, not a prediction of a test result.

How we measured it

We score a corpus of essays that already carry human ratings, through the exact prompt, schema and model the live product uses, and compare the AI's overall band with the human label essay by essay.

The scorer is a language model prompted with the publicly published IELTS band descriptors, plus a calibration block that corrects the known tendency of language models to mark IELTS Writing too generously. The calibration run imports that prompt from the same source file the live scoring route imports it from, so the number on this page measures the scorer you actually use rather than a copy of it. Every essay is scored once, the four criterion bands are averaged and rounded with the standard half-band rule, and the result is stored before any statistic is computed.

Run details
ItemValue
CorpusNot yet published
LicenceNot yet published
SourceNot yet published
Label scaleNot yet published
Corpus retrievedNot yet published
ModelNot yet published
Prompt versionNot yet published
Essays scoredNot yet published
Run dateNot yet published

The scripts that produce this page live in the repository at scripts/calibration/ — a fetch step that normalises the corpus and writes a manifest of ids, labels and per-essay hashes, a runner that scores them, and a report step that computes the tables below. The corpus text itself is not redistributed by us.

What this number does not mean

It is a measurement of agreement on one set of practice essays. It is not a prediction of your result, and it is not a claim that the AI marks like a certified examiner.

  • Not an examiner, and not affiliated. IELTS-Bank has no connection to the organisations that own or run IELTS. The scorer reads a text file; a certified examiner applies training, moderation and judgement that no prompt reproduces. Nothing here predicts or guarantees a test-day band.
  • The labels are human, not test scores. The corpus is practice writing rated by people, not scripts from a real sitting, and no test material is used. Human raters also disagree with each other, so part of the gap measured here is disagreement that exists between humans too.
  • Sample size and coverage. A few hundred essays is enough to detect a systematic lean, not enough to pin an agreement rate to the nearest percentage point. Bands at the edges of the range are represented by fewer essays than the middle, and the essays are not a random sample of the people who use this site.
  • It drifts. Model providers update models underneath a fixed model name, and we change the prompt. That is why each figure is stamped with a model and a prompt version, and why this page is regenerated rather than edited.

What the score is genuinely good for is diagnosis: which of the four criteria is holding your band down, and which sentences are costing you marks. That is what the report is built around, and it is useful even when the overall band is half a band off.

AI band score accuracy FAQ

How accurate is the IELTS-Bank AI writing score?

We are re-running our calibration and will not quote an accuracy figure until it is measured on a large enough set of human-scored essays. Until then, treat every band this tool gives you as an estimate only.

Is the AI score the same as an examiner score?

No. Our scorer is a language model prompted with the public IELTS band descriptors. It has no connection to the IELTS test partners, it does not see what an examiner sees, and it cannot award, predict or guarantee a test-day result. Use it to find what is holding a criterion down, not to decide whether you are ready.

What does "within 0.5 of a band" actually mean?

It is the share of essays where the difference between the AI overall band and the human label was no more than half a band. It is a measure of agreement on this set of essays. It is not a guarantee about your essay, and it says nothing about how two human raters would differ from each other on the same scripts.

Which essays were used?

A licensed, human-scored corpus of practice essays. The corpus, its licence and the date it was retrieved are published here as soon as the run completes. No test material is ever used.

Will this number change?

Yes. Model versions change, our prompt changes, and the corpus can be extended. Each published figure names the model and the prompt version it was measured on, and the page is regenerated from the stored results whenever we re-run the calibration.

See it on your own essay

The fastest way to judge a scorer is to give it writing whose band you already know.