Skip to main content

Standards and Testing Agency: LLM - KS2 Writing Sample Generator

Large Language model used to generate examples of pupils' writing at various standards at the end of KS2.

1. Summary

1 - Name

LLM - KS2 Writing Sample Generator

2 - Description

We use LLMs to generate English writing material including creative writing, factual and fictional prose, and poetry that is akin to that done by year 6 pupils at various standards within the teacher assessment framework, based on prompts written by Teacher Assessment (TA) Test Development Researcher (TDR). Outputs significantly edited before final use.

3 - Website URL

n/a

4 - Contact email

TAMOD.data@education.gov.uk

Tier 2 - Owner and Responsibility

1.1 - Organisation or department

Standards and Testing Agency (STA)

1.2 - Team

Teacher Assessment and Moderation Team

1.3 - Senior responsible owner

Deputy Director, Assessment Operations and Services

1.4 - Third party involvement

No

1.4.1 - Third party

Chat GPT 5

1.4.2 - Companies House Number

n/a

1.4.3 - Third party role

OpenAI

1.4.4 - Procurement procedure type

n/a

1.4.5 - Third party data access terms

n/a

Tier 2 - Description and Rationale

2.1 - Detailed description

The LLM - KS2 Writing Sample Generator (Chat GPT 5) uses specially designed prompts to generate year 6 writing, using the Teacher Assessment Framework for writing at the end of Key Stage 2 to ensure all outputs fit the specified standard. Prospective moderators must pass a standardisation exercise annually to achieve ‘approval to moderate KS2 writing’ status.The standard of pupils’ writing is assessed at the end of Key Stage 2 (year 6) by their teacher rather than a formal test as is the case with English and maths. A pupil will be judged to be working at the standard, working towards the standard, or working at greater depth. Created by the Standards and Testing Agency (STA) in DfE, teachers have a framework on which to base their assessment. They consider the pupil’s written work across year 6 to determine the standard they have achieved. Local authority moderators check teacher judgements to ensure national consistency. To complete their training, moderators must pass one of three standardisation exercises by assessing pupil writing portfolios against the framework and judging the standard correctly. Around 2,000 moderators pass each year.

2.2 - Benefits

Use of this LLM (Chat GPT 5) will end the costly and burdensome practice of acquiring pupils’ scripts from schools. Collections can be created quickly and easily; the LLM provides greater flexibility to ensure each ‘pupil can statement’ with each specific standard is met securely. Annual cost of standardisation has reduced from over £100,000 to less than £5,000.

2.3 - Previous process

Previously, STA was tied to a 5 year contract with a supplier, who would source pupil scripts directly from schools and create collections to represent a particular standard. Although some pupils’ work was occasionally edited, some collections presented were borderline between two standards. This has potential to vastly reduce the costs of creating standardisation exercises. For 2026/27 and 2027/28, the cost of producing standardisation exercises will be around £5k annually (for QA on 2 full exercises - with one full exercise comprising scripts from the old contract with our former supplier). A decision on whether to continue this method of production beyond 2027/28 or procure a new supplier will be made in spring 2027

2.4 - Alternatives considered

Various LLMs trialled, but others proved less effective. Other non LLM options would be, to continue to develop at cost, exercises in the current way.

Tier 2 - Deployment Context

3.1 - Integration into broader operational process

LLM- KS2 Writing Sample Generator (Chat GPT 5) is given a prompt, composed using knowledge of current year 6 topics or writing genres and the teacher assessment framework. Output is heavily edited by TA researcher before external review, before being copied into a word document for review by QA panel.

3.2 - Human review

Outputs are reviewed and heavily edited by TA researcher with in-depth knowledge of the TA framework and year 6 writing. Following internal review by TDRs, a panel of external moderation managers are utilised to further review authenticity.

3.3 - Frequency and scale of usage

At the moment the LLM (Chat GPT 5) is used to create 3 collections of pupils’ writing for one standardisation exercise.

3.4 - Required training

n/a

3.5 - Appeals and review

n/a

Tier 2 - Tool Specification

4.1.1 - System architecture

The LLM- KS2 Writing Sample Generator is a conversational AI system built on LLM technology (It was developed by OpenAI and is powered by the GPT (Generative Pre-trained Transformer) architecture flexos.work.-5) – essentially a Generative AI that produces human-like text responses from prompts.

4.1.2 - System-level input

Prompts used are deliberately open ended and rely on prior knowledge on genres and topics taught in year 6., with some context provided. The Teacher Assessment Framework for KS2 writing is also shared with the LLM.

4.1.3 - System-level output

Outputs generated are used as the basis upon which a piece of writing is created, with significant editing by TA researcher.

4.1.4 - Maintenance

n/a

4.1.5 - Models

n/a

Tier 2 - Model Specification

4.2.1. - Model name

Chat GPT 5 - self hosted

4.2.2 - Model version

Chat GPT 5

4.2.3 - Model task

The model is asked to provide a piece of writing similar to one a pupil in year 6 may write, which, once heavily edited, will be used to represent a particular standard within the Teacher Assessment Framework, to be used in a standardisation exercise in which aspiring moderators will need to correctly judge a collection.

4.2.4 - Model input

n/a

4.2.5 - Model output

Examples of pieces of Yr6 writing

4.2.6 - Model architecture

n/a

4.2.7 - Model performance

TA researcher (former primary school teacher with experience of assessing KS2 writing) responsible for editing outputs. QA process uses around 20 moderation managers from LAs to feedback on content of collections of writing, all of whom are vastly experienced in this field.

4.2.8 - Datasets and their purposes

n/a

2.4.3. Development Data

4.3.1 - Development data description

n/a

4.3.2 - Data modality

n/a

4.3.3 - Data quantities

n/a

4.3.4 - Sensitive attributes

n/a

4.3.5 - Data completeness and representativeness

n/a

4.3.6 - Data cleaning

n/a

4.3.7 - Data collection

n/a

4.3.8 - Data access and storage

n/a

4.3.9 - Data sharing agreements

n/a

Tier 2 - Operational Data Specification

4.4.1 - Data sources

Outputs heavily edited to add authenticity and to more accurately reflect writing of a year 6 pupil. Prompts made using knowledge of real year 6 writing topics and genres

4.4.2 - Sensitive attributes

No personal or sensitive data shared with LLM

4.4.3 - Data processing methods

Teacher Assessment researchers heavily edit outputs before they are shared for review.

4.4.4 - Data access and storage

n/a

4.4.5 - Data sharing agreements

n/a

Tier 2 - Risks, Mitigations and Impact Assessments

5.1 - Impact assessments

n/a

5.2 - Risks and mitigations

No sensitive information has been shared with this LLM. All information (ie Teacher Assessment Frameworks, training and exemplification materials) are publicly available. LLMs produce texts that often excludes atypical vocabulary and sentence structures that might be used by neurodivergent or non-native English speaking students Mitigated by thorough review. Accuracy of outputs is heavily reviewed and edited and subject to thorough internal and external review before being used in standardisation.

Updates to this page

Published 10 August 2026