Plans to replace real examples of children’s work with AI created answers in SATs moderator tests will save the government at least £95,000 a year. Schools Week revealed the Standards and Testing Agency (STA) conducted a pilot last year, exploring whether large language models could help with the hefty standardisation exercise. Council moderators currently review a quarter of primary school writing scripts, to check teachers’ marking is consistent. But to complete their training, moderators must pass one of three standardisation exercises. One of these exercises is now using AI-generated sample answers including on creative writing, poetry and factual and fictional prose. In transparency records published this week, the Department for Education revealed it will reduce annual standardisation costs from more than £100,000 to less than £5,000. The use of ChatGPT 5 will “end the costly and burdensome practice of acquiring pupils’ scripts from schools,” it said. Since 2021-22, the Australian Council for Education Research (ACER) has sourced real samples of year 6 children’s work from schools to create the standardisation exercises. Two exercises under the pilot were still collected this way and will continue to be in for the 2026-27 and 2027-28 cycles. The STA will decide whether to continue this process beyond 2027-28 or procure a new supplier in spring 2027. AI mimics year 6 writing The “LLM-KS2 writing sample generator” is given a prompt, written by assessment researchers, to create responses that mimic the writing of year 6 pupils working to different standards. The prompts are “deliberately open ended and rely on prior knowledge on genres and topics taught in year 6, with some context provided”. The teacher assessment framework for key stage 2 writing is also shared. These AI-generated responses are “reviewed and heavily edited” by an assessment researcher, then reviewed externally by a panel of around 20 “vastly experienced” council moderation managers. DfE said collections can be “created quality and easily” and the LLM provides “greater flexibility to ensure each ‘pupil can statement’ with each specific standard is met securely”. About 2,000 moderators pass the overall test east year. Councils had the option to opt out of the trial, with their moderators sitting other exercises. It’s not clear if this will continue in the next two years. ‘Less authentic’ Schools Week contacted the largest 15 English councils for their experiences. Kent said only a third of the 24 moderators who completed the exercise achieved a pass, which was lower than the outcome for other exercises this year and comparable exercises from previous years. “Feedback from experienced moderators suggested the AI-generated writing often felt less authentic than genuine pupil work, with unusual patterns of errors and language choices. “Moderators also found it more difficult to make secure judgements about authenticity and independence. “While we cannot conclude that the AI-generated materials were solely responsible for the outcome, our feedback suggests they introduced additional complexity rather than simplifying the moderation process.” They added that further testing would be needed before AI-generated writing could replicate authentic work for moderation. Hampshire took part in STA research before the trial was launched. A council spokesperson said it identified “some subtle differences” between real and AI-generated scripts, but they “did not affect our ability to apply the assessment framework accurately and consistently”. The actual trial had “no noticeable impact on the overall standardisation process”. “It did not appear to affect outcomes for moderators, nor did it make the exercises noticeably simpler or more difficult to complete.” Lancashire took part in the trial “on a very limited basis by using a small number of moderators”, a spokesperson said. “The outcome showed no discrepancy with traditional methods of evaluating assessment.” Transparency must be maintained Sarah Hannafin, NAHT’s head of policy, said the school leaders’ union was “not particularly concerned… if complete transparency is maintained”. “A sensible approach was taken to trialling, with only one of the scripts created using AI; there is strong human oversight with expert review and editing. It was good to save money, but it was “a small saving in the context of the overall expenditure on testing”. Hannafin also stressed that the scripts should not become “too simplistic”. “Children’s work is often not that clear cut, and teachers/moderators need to be able to consider work which is borderline between two standards and make an informed decision on it. “STA need to evaluate and review the performance of the AI generated scripts compared to those created using pupil work as they consider whether to expand the trial.” Mick Walker, president of the Chartered Institute of Educational Assessors, said the trial saves schools from having to “dig out the samples and send them in”. “There are cost savings there without a doubt, and it saves schools having to do the legwork on that. “We still need that robust quality assurance process to make sure that the selected scripts are actually used for the training are at the right standards.” Walker added that the AI-generated responses needed to reflect the diversity of the population, including children with special educational needs. The government update said the tendency for the LLMs to exclude “atypical vocabulary and sentence structures” that may be used by neurodivergent or EAL pupils was “mitigated by thorough review”. The DfE was approached for comment.