Instant.
Accurate.
Reliable.

Does it mark as accurately as a human?

Evaluating a learner’s speaking involves subjective judgement and we know that human markers don’t all give the same score when marking a speaking test. That’s why test providers put a lot of effort into training their markers and moderating their work. They also analyse markers’ performance to measure what they call inter-rater reliability.

When it comes to operational tests, inter-rater reliability is expected to be over 0.8 – in other words, an 80% level of agreement between markers. IELTS, for example, in 2017 claimed levels of 0.83–0.86 for their speaking component, when each test was marked by a single person.1 If a test is marked by a second marker, the level rises to 0.90.2

When evaluating the reliability of the Dynamic Speaking Test, a calibration study was carried out to establish the reliability of the automated marking, by analysing its scores alongside those of six human markers. The conclusions of the study are given below.

The AI marking alongside human markers

DST calibration study 2024

  • Krippendorff’s alpha (a standard measure for inter-rater reliability) is α = .92 for all the scores across the seven markers (human and automated). This is well above the minimum acceptable level of 0.8, and also above typical operational levels. The values are statistically significant (p <0.01).
  • Comparing automated scores against individual human markers, the inter-rater reliability varies between 0.89 – 0.94. Comparing individual human markers with each other, the inter-rater reliability varies between 0.89 – 0.95.
  • A multi-facet Rasch analysis (MFRM) positioned the automated marking within the expected range of rater stringency. Two human markers were slightly more lenient than the automated marking, four were slightly more strict.
  • In conclusion, the automated scores show high correlation with those of the human markers, with no systematic differences or biases between AI and human ratings.

The full report can be found here.

AI and human markers giving the same scores

The report also points out that, unlike human markers, the automated system is completely consistent with itself. It does not suffer from tiredness or unconscious biases and applies the marking criteria in the same way every time. It therefore does not require ongoing moderation.

References

1 Quaid, E.D. ‘Reviewing the IELTS speaking test in East Asia: theoretical and practice-based insights.’ Lang Test Asia 8, 2 (2018). https://doi.org/10.1186/s40468-018-0056-5

2 IELTS (2026) Test statistics. Retrieved 15 June 2026. https://ielts.org/researchers/our-research/test-statistics

Key findings

  • Reliable for placement. AI scores fell within the same level of accuracy as a panel of expert human markers, with an inter-rater reliability > 0.9.
  • Humans set the standard. The CEFR boundaries themselves – what counts as a B1 response rather than an A2, and so on – are defined by human experts. The AI is calibrated to match those judgements, not to replace them.
  • Faster results, same quality. AI scoring delivers in seconds what a human panel takes a scheduled session to produce, without compromising the accuracy of the placement.
The assessment team reviewing calibration results