Standards and Testing Agency: LLM - KS2 Writing Sample Generator
Large Language model used to generate examples of pupils' writing at various standards at the end of KS2.
1. Summary
1 - Name
LLM - KS2 Writing Sample Generator
2 - Description
We use LLMs to generate English writing material including creative writing, factual and fictional prose, and poetry that is akin to that done by year 6 pupils at various standards within the teacher assessment framework, based on prompts written by Teacher Assessment (TA) Test Development Researcher (TDR). Outputs significantly edited before final use.
3 - Website URL
n/a
4 - Contact email
Tier 2 - Owner and Responsibility
1.1 - Organisation or department
Standards and Testing Agency (STA)
1.2 - Team
Teacher Assessment and Moderation Team
1.3 - Senior responsible owner
Deputy Director, Assessment Operations and Services
1.4 - Third party involvement
No
1.4.1 - Third party
Chat GPT 5
1.4.2 - Companies House Number
n/a
1.4.3 - Third party role
OpenAI
1.4.4 - Procurement procedure type
n/a
1.4.5 - Third party data access terms
n/a
Tier 2 - Description and Rationale
2.1 - Detailed description
The LLM - KS2 Writing Sample Generator (Chat GPT 5) uses specially designed prompts to generate year 6 writing, using the Teacher Assessment Framework for writing at the end of Key Stage 2 to ensure all outputs fit the specified standard. Prospective moderators must pass a standardisation exercise annually to achieve ‘approval to moderate KS2 writing’ status.The standard of pupils’ writing is assessed at the end of Key Stage 2 (year 6) by their teacher rather than a formal test as is the case with English and maths. A pupil will be judged to be working at the standard, working towards the standard, or working at greater depth. Created by the Standards and Testing Agency (STA) in DfE, teachers have a framework on which to base their assessment. They consider the pupil’s written work across year 6 to determine the standard they have achieved. Local authority moderators check teacher judgements to ensure national consistency. To complete their training, moderators must pass one of three standardisation exercises by assessing pupil writing portfolios against the framework and judging the standard correctly. Around 2,000 moderators pass each year.
2.2 - Benefits
Use of this LLM (Chat GPT 5) will end the costly and burdensome practice of acquiring pupils’ scripts from schools. Collections can be created quickly and easily; the LLM provides greater flexibility to ensure each ‘pupil can statement’ with each specific standard is met securely. Annual cost of standardisation has reduced from over £100,000 to less than £5,000.
2.3 - Previous process
Previously, STA was tied to a 5 year contract with a supplier, who would source pupil scripts directly from schools and create collections to represent a particular standard. Although some pupils’ work was occasionally edited, some collections presented were borderline between two standards. This has potential to vastly reduce the costs of creating standardisation exercises. For 2026/27 and 2027/28, the cost of producing standardisation exercises will be around £5k annually (for QA on 2 full exercises - with one full exercise comprising scripts from the old contract with our former supplier). A decision on whether to continue this method of production beyond 2027/28 or procure a new supplier will be made in spring 2027
2.4 - Alternatives considered
Various LLMs trialled, but others proved less effective. Other non LLM options would be, to continue to develop at cost, exercises in the current way.
Tier 2 - Deployment Context
3.1 - Integration into broader operational process
LLM- KS2 Writing Sample Generator (Chat GPT 5) is given a prompt, composed using knowledge of current year 6 topics or writing genres and the teacher assessment framework. Output is heavily edited by TA researcher before external review, before being copied into a word document for review by QA panel.
3.2 - Human review
Outputs are reviewed and heavily edited by TA researcher with in-depth knowledge of the TA framework and year 6 writing. Following internal review by TDRs, a panel of external moderation managers are utilised to further review authenticity.
3.3 - Frequency and scale of usage
At the moment the LLM (Chat GPT 5) is used to create 3 collections of pupils’ writing for one standardisation exercise.
3.4 - Required training
n/a
3.5 - Appeals and review
n/a
Tier 2 - Tool Specification
4.1.1 - System architecture
The LLM- KS2 Writing Sample Generator is a conversational AI system built on LLM technology (It was developed by OpenAI and is powered by the GPT (Generative Pre-trained Transformer) architecture flexos.work.-5) – essentially a Generative AI that produces human-like text responses from prompts.
4.1.2 - System-level input
Prompts used are deliberately open ended and rely on prior knowledge on genres and topics taught in year 6., with some context provided. The Teacher Assessment Framework for KS2 writing is also shared with the LLM.
4.1.3 - System-level output
Outputs generated are used as the basis upon which a piece of writing is created, with significant editing by TA researcher.
4.1.4 - Maintenance
n/a
4.1.5 - Models
n/a
Tier 2 - Model Specification
4.2.1. - Model name
Chat GPT 5 - self hosted
4.2.2 - Model version
Chat GPT 5
4.2.3 - Model task
The model is asked to provide a piece of writing similar to one a pupil in year 6 may write, which, once heavily edited, will be used to represent a particular standard within the Teacher Assessment Framework, to be used in a standardisation exercise in which aspiring moderators will need to correctly judge a collection.
4.2.4 - Model input
n/a
4.2.5 - Model output
Examples of pieces of Yr6 writing
4.2.6 - Model architecture
n/a
4.2.7 - Model performance
TA researcher (former primary school teacher with experience of assessing KS2 writing) responsible for editing outputs. QA process uses around 20 moderation managers from LAs to feedback on content of collections of writing, all of whom are vastly experienced in this field.
4.2.8 - Datasets and their purposes
n/a
2.4.3. Development Data
4.3.1 - Development data description
n/a
4.3.2 - Data modality
n/a
4.3.3 - Data quantities
n/a
4.3.4 - Sensitive attributes
n/a
4.3.5 - Data completeness and representativeness
n/a
4.3.6 - Data cleaning
n/a
4.3.7 - Data collection
n/a
4.3.8 - Data access and storage
n/a
4.3.9 - Data sharing agreements
n/a
Tier 2 - Operational Data Specification
4.4.1 - Data sources
Outputs heavily edited to add authenticity and to more accurately reflect writing of a year 6 pupil. Prompts made using knowledge of real year 6 writing topics and genres
4.4.2 - Sensitive attributes
No personal or sensitive data shared with LLM
4.4.3 - Data processing methods
Teacher Assessment researchers heavily edit outputs before they are shared for review.
4.4.4 - Data access and storage
n/a
4.4.5 - Data sharing agreements
n/a
Tier 2 - Risks, Mitigations and Impact Assessments
5.1 - Impact assessments
n/a
5.2 - Risks and mitigations
No sensitive information has been shared with this LLM. All information (ie Teacher Assessment Frameworks, training and exemplification materials) are publicly available. LLMs produce texts that often excludes atypical vocabulary and sentence structures that might be used by neurodivergent or non-native English speaking students Mitigated by thorough review. Accuracy of outputs is heavily reviewed and edited and subject to thorough internal and external review before being used in standardisation.