Evaluating a learner’s speaking involves subjective judgement and we know that human markers don’t all give the same score when marking a speaking test. That’s why test providers put a lot of effort into training their markers and moderating their work. They also analyse markers’ performance to measure what they call inter-rater reliability.
When it comes to operational tests, inter-rater reliability is expected to be over 0.8 – in other words, an 80% level of agreement between markers. IELTS, for example, in 2017 claimed levels of 0.83–0.86 for their speaking component, when each test was marked by a single person.1 If a test is marked by a second marker, the level rises to 0.90.2
When evaluating the reliability of the Dynamic Speaking Test, a calibration study was carried out to establish the reliability of the automated marking, by analysing its scores alongside those of six human markers. The conclusions of the study are given below.
The full report can be found here.
The report also points out that, unlike human markers, the automated system is completely consistent with itself. It does not suffer from tiredness or unconscious biases and applies the marking criteria in the same way every time. It therefore does not require ongoing moderation.
References
1 Quaid, E.D. ‘Reviewing the IELTS speaking test in East Asia: theoretical and practice-based insights.’ Lang Test Asia 8, 2 (2018). https://doi.org/10.1186/s40468-018-0056-5
2 IELTS (2026) Test statistics. Retrieved 15 June 2026. https://ielts.org/researchers/our-research/test-statistics