Skip to main content
Transparency data

Technical Guide: Updated estimates of children with a parent in prison

Published 10 September 2026

Applies to England and Wales

1. Introduction

These estimates are of children under 18 relating to having a parent in prison.

This document provides a comprehensive guide to the estimates. It covers:

  • An explanation of the sources and quality of data used to produce the estimates;

  • The methodology adopted to compile the estimates (including data linking and Natural Language Processing);

  • Assumptions and limitations of the data and analysis

2. Data sources and quality

2.1 Summary of data sources

These estimates are based on administrative data held across His Majesty’s Prison and Probation Service (HMPPS). Five HMPPS data sources were used, comprising both structured datasets organised in formatted rows and columns, as well as unstructured free-text datasets. Historical records were retrieved from prior to 1st October 2024 to align with the end of the one-year study period, dating back up to twenty years in some cases.

Structured data sources:

  • Prison National Offender Management Information System (NOMIS) visitor lists

  • Probation National Delius (NDelius) personal contact lists

Unstructured data sources:

  • Offender Assessment System (OASys) assessments

  • NOMIS prison case notes

  • NDelius probation case notes

In addition, a data-matching exercise was undertaken with His Majesty’s Revenue and Customs (HMRC) using:

  • HMRC Child Benefit records

NOMIS contains information on individuals in custody, including their prison movements, activities and recorded contacts. It was used to identify the prison cohort (all individuals in prison during the period 1 October 2023 to 1 October 2024) and to obtain information on prisoner contacts.

OASys contains assessments completed for individuals at various stages of their criminal justice journey, including during custody and while under probation supervision. These assessments capture a wide range of information on offending-related risks, needs, and personal circumstances. Assessors may complete either a basic assessment (Layer 1) or a full assessment (Layer 3), depending on the individual’s circumstances and operational requirements.

NDelius is the probation case management system and contains information relating to individuals supervised in the community, including contact details, case management information and case notes.

The quality of information held within these operational systems is generally high, although it varies by data field depending on its operational importance and frequency of use.

On average an individual prisoner has around 600 probation and 300 prison case notes written about them across NDelius and NOMIS. These free-text records may include information disclosed directly by the prisoner, as well as details about dependants and wider family or social relationships obtained through contact with other agencies, such as social services. However, prisoners may be unwilling to disclose information about their children to staff for a range of reasons, such as concerns that it may trigger involvement from social services.

Further information on the quality and limitations of the specific data fields used in this analysis is provided below.

HMPPS data were linked to HMRC Child Benefit records for analytical purposes. Full Child Benefit accounts comprising claimants and dependants were returned to MoJ, which may include individuals not recorded in HMPPS data. All data were used solely to produce aggregated statistical findings. The circumstances or records of identifiable individuals were not examined.

2.2 Use of administrative data

The use of administrative data allows analysis to be conducted across the entire prison population rather than on a sample. As these datasets are routinely collected for operational purposes, their use also avoids the need for additional data collection.

Details of all administrative data sources used in the production of this release can be found in the MoJ Statement of Administrative Sources.

As with all large-scale administrative datasets, these datasets are subject to limitations. Errors may occur during data entry, updating or processing. In addition, the recording of dependants may differ across systems because definitions are not always applied consistently, and operational staff may interpret guidance differently.

2.3 Structured data fields used

NOMIS contact lists

NOMIS records details of individuals approved to visit prisoners. Where available, these records include the contact’s name, date of birth and relationship to the prisoner. However, not all children associated with a prisoner will necessarily be captured due to a range of reasons including:  

  • A child has not been registered as a visitor 

  • The parent has purposefully not disclosed parental status 

  • The prisoner is restricted from having contact with their children.

For this analysis, children were defined as individuals aged under 18 on 1 October 2024, the final day of the cohort period. Contacts aged 18 and over were therefore excluded.

NDelius contact lists

NDelius records key contacts associated with individuals under probation supervision, including next of kin and other significant personal contacts. These records may include a contact’s name, date of birth and relationship to the individual.

For this analysis, contacts who were aged under 18 on 1 October 2024 and whose relationship was recorded as son, daughter, stepson or stepdaughter were identified as children. The dataset may also include contacts who are carers or guardians of children associated with the person under supervision.

2.4 Unstructured HMPPS data sources

OASys assessments

OASys assessments contain information on family and personal relationships, including whether an individual has children and the nature of their contact with them. The recorded relationship may include categories such as child, grandchild or sibling. Whilst all prisoners should have an assessment, the availability and accuracy of this information depend on the individual’s disclosure of their circumstances and the quality of information captured and recorded by the assessor at the time of assessment.

The three components included in this study are the Basic Custody Screening Tool (BCST) conducted within 72 hours of arrival into custody, the full assessment, and Child at Risk records. BCST focuses on the number of children in the household and does not mandate the recording of individual children’s details. Child at Risk identifies children who may be at risk of serious harm from the offender, where records contain detailed information on individual children.

NOMIS case notes

Case notes are used by prison staff to record information about a prisoner’s behaviour, progress and other relevant events or circumstances. They provide a detailed and regularly updated record of an individual’s situation while in custody.

Staff with contact responsibilities are expected to record information in case notes regularly, and management processes are in place to monitor the frequency and quality of entries.

As case notes are unstructured free-text records, there is no mandatory format or requirement to record information about dependants. Consequently, the presence of relevant information depends on staff identifying it as sufficiently important to document.

NDelius case notes

NDelius case notes record interactions between probation practitioners and individuals under supervision, as well as relevant contact with third parties.

Guidance requires that contacts are recorded within one working day and that records clearly distinguish between factual information and professional opinion.

As with NOMIS case notes, NDelius case notes are unstructured free-text records. Information about children or dependants is therefore only available where practitioners consider it relevant to record.

2.5 HMRC Child Benefit data

Child Benefit is provided to families responsible for bringing up a child under 16, or under 20 if they are in approved education or training. Only one person can claim Child Benefit for a child, but there is no limit to how many children they can claim for. This benefit is not to be confused with Universal Credit or Tax Credits administered by DWP which has been capped at two children for certain time periods. For HMRC Child Benefit, age is part of the qualification criteria, and so benefit records also contain the numbers and ages of children as well as the address.

As the data relates directly to children, it should be a reliable indicator of whether an individual has dependants under 18.

However, Child Benefit is claimed by the person with primary responsibility for the child, who may not always be the child’s parent. Although most claimants are parents, they could also be grandparents, guardians and others if the child is not in parental care. As of April 2026, 87% of all child benefit records related to children aged under 18. However for the study period used to produce the estimates in this publication, Child Benefit take-up rates were at 90% as of May 2023.

Findings indicated that the use of AI enabled more complete information to be extracted from HMPPS unstructured data, reducing the number of children identified solely from HMRC Child Benefit.

3. Data governance

The BOLD programme has established procedures for the effective governance of data it uses across all pilot projects. The BOLD programme’s privacy notice is available on gov.uk.

3.1 Governance

Analysis and research using data collected in operational systems across HMPPS is covered by pre-existing Data Protection Impact Assessments (DPIA). This work is covered by these under the research and analysis purpose. Statistics are in aggregate form only for the purposes of understanding offenders; information about any specific individual is not of interest.

3.2 Confidentiality

This statement sets out the arrangements in place for protecting persons’ confidential data when estimates are published or otherwise released into the public domain. The Code of Practice for Statistics states that:

Organisations should look after people’s information securely and manage data in ways that are consistent with relevant legislation and serve the public good.

To comply with this and with the Data Protection Act of 2018 and to maintain the trust and co-operation of those who use these estimates, the following provisions have been put in place:

  • Private information collected by MoJ is stored in line with our data security policies.

  • Electronic data is held on password-protected networks.

  • All new staff undergo security vetting before receiving access to data systems and all staff undertake mandatory training on information responsibility annually.

Some counts may have been removed for Statistical Disclosure Control purposes. In line with MoJ and GSS guidance, assessment of the risk of disclosure considers the following:

  • Level of aggregation (including geographic level) of the data;

  • Size of the population;

  • Likelihood of an attempt to identify; and

  • Consequences of disclosure.

3.3 Engaging the public

Public trust around how data is shared is critical for BOLD, and we partnered with the Centre for Data Ethics & Innovation (CDEI), and the research company Britain Thinks, to undertake extensive engagement with affected groups, trusted intermediaries, and the general public. The results of this exercise, and what we have learnt from listening to the public, have tangibly informed the design of the BOLD programme and has been published by the CDEI.

3.4 Use of AI models and data security

This work uses Large Language Models (LLMs) to analyse and extract information from free-text contact note data. LLMs are a type of artificial intelligence (AI) which can process and analyse human language and can be instructed to carry out specific tasks through prompts and examples.

The LLMs used by MoJ in this publication have been hosted by AWS via Bedrock . The model weights are securely hosted and managed entirely within the AWS system. All requests are processed within the AWS environment with no third-party API calls. Input prompts, data and outputs are kept private and not used for model training. Bedrock architecture has secure end-to-end encryption. All data is hosted and processed within the EU-West region.

4. Methodology

The estimates in this report were produced using a combination of administrative data linkage, large language models (LLMs), and matching to HMRC Child Benefit records. In summary, the process involved:

  1. Extract social relationships of adult prisoners (of all ages) from structured HMPPS data

  2. Extract social relationships of adult prisoners from unstructured data fields, primarily case notes

  3. Identifying relationships likely to represent a prisoner’s children or carers of those children

  4. Linking these records to HMRC Child Benefit claimant and child records

  5. Adjusting estimates to account for under-identification in the linked data using external validation sources

  6. Comparing local authority-level estimates with independent data sources to assess potential bias or coverage issues.

4.1 Learnings from previous methodology

Previous estimates relied on matching prisoners and their partners to HMRC Child Benefit records. This work found that male prisoners were often not the Child Benefit claimant, with successful matches more likely through female partners. The revised methodology therefore uses a wider range of social relationship information to identify children and their carers, including partners, ex-partners, grandparents, siblings and other family members.

4.2 Linking of HMPPS prisoner records

The main HMPPS data sources (NOMIS, NDelius and OASys) do not share a common identifier, making direct linkage difficult. To overcome this, records were linked and deduplicated using Splink, a Ministry of Justice probabilistic matching tool. This enabled data relating to the same individual to be combined across systems despite inconsistencies or missing information.

Further details about Splink are available in Data First: An Introductory User Guide and Data First: Criminal Courts Linked Data.

4.3 Information extraction from unstructured data

Free-text data was filtered to identify references to social relationships using regular expressions (RegEx) and keyword searches, while system-generated notes unlikely to contain useful information were excluded. A full list of the keywords are given in Section 6.1.

The extracted text was then processed using an LLM to identify relationship types and associated personal details such as name, addresses, relationship type, date of birth. An example prompt and some sample input and output examples are given in Section 6.2.

4.4 Model selection

Models were tested taking account of accuracy, processing time and cost as well as data source. Prompts (instructions for information extraction given to the model) were refined to improve data extraction, minimise unnecessary personal data collection and reduce potential bias in the identification of family relationships.

4.5 Data minimisation and filtering

Extracted relationship data underwent further processing to establish how individuals were connected to a prisoner. Concretely, canonical relationship types were defined to capture basic relations such as partner and child. Subsequently, multi-step relationships were standardised into relationship chains where possible to support consistent filtering. For instance, the phrases “prisoner’s partner’s child” and “child of prisoner’s partner” are both standardised into the chain prisoner → partner → child.

Records were retained only where there was evidence of a potential parent-child relationship. This ensured that only information relevant to identifying prisoners’ children was included in subsequent matching.

Filtering was based on:

  • Relationship type

  • Child shared addresses with the prisoner

  • Child age (under 18 on 1 October 2024).

4.6 Linking to HMRC Child Benefit data

Prisoner and family relationship data were linked to HMRC Child Benefit records from 2019 to 2024 using probabilistic matching via Splink. Child Benefit data contains information on claimants, children and household characteristics, providing an additional source for identifying dependent children. Nested matching was carried out in two stages:

  1. Family-level matching, comparing household information across datasets.

  2. Individual-level matching within likely family matches.

All children on the Child Benefit account are sent back to MoJ if a match is found to a prisoner thought to have children.

Individuals were flagged if they matched to more than one child or more than one claimant. Personal details were not received back from HMRC in these instances of multiple matches.

An additional 14% of children were identified through linkage with HMRC Child Benefit claimant records. This proportion is lower than might be expected because the improved AI-based methodology was able to identify a much larger number of children directly from prison records.

In total, 37% of children identified were either confirmed or supplemented through HMRC Child Benefit data. This figure is also lower than anticipated because compliance with General Data Protection Regulation (GDPR) data minimisation principles, meant that not all prisoner social network data could be sent to HMRC in the event of finding a possible match. This reduced the success of matching due to being unable to test multiple address combinations to pinpoint which of these linked to prisoners’ families. This has been mitigated by adjusting for undercount (see section 4.8).

4.7 Creating a final dataset

Information from HMPPS and HMRC sources was deduplicated and combined to create a single dataset of identified children affected by parental imprisonment. Duplicate records of the same child arising from multiple data sources or years of Child Benefit data were combined using probabilistic matching techniques. Where information conflicted, HMRC Child Benefit data was prioritised where available.

4.8 Validation and adjustment for undercount

Coverage of the linked dataset was assessed against two independent sources:

Operation Paramount, which identifies children affected by parental imprisonment in the Thames Valley region.

The Children’s Commissioner’s census of schools and colleges in England. Schools were asked to provide an estimate of the number of pupils with a parent or carer in prison in summer 2024.

Operation Paramount is a police-led initiative within the Thames Valley Police area that connects children of imprisoned parents with third-sector organisations that can provide support. The initiative relies on children being recorded within police systems, for example through contact with the police as members of a household involved in an incident such as a burglary. However, it has been running for several years, with many process improvements over that period, and is taken to be a robust source of information on children with a parent in prison for that area.

A detailed analysis compared a sample of children from the linked MoJ-HMRC data with Operation Paramount. Of a sample of 24 children (and their families and associated household members taking the sample to around 100 individuals), 20 children were identified as well as most of their other household members (siblings under 18). To account for this undercount, the final estimate was adjusted by a factor of 1.2. No evidence of systematic underestimation was identified through comparison with Children’s Commissioner data. The Children’s Commissioner’s estimated counts were compared with local authority counts of identified children. The resulting Spearman’s rank correlation coefficient was 0.61, indicating a good level of agreement between the two datasets.

4.9 Caveats and limitations

Several limitations should be considered when interpreting these estimates.

Use of Large Language Models:

LLMs may occasionally misclassify or infer relationships incorrectly. To mitigate this risk, extracted information was tested against source records and supplemented with structured data where available. Prompt design and validation exercises were also used to minimise bias and improve accuracy.

Relationship identification:

LLMs performed less well at identifying more complex relationships, such as stepchildren or a partner’s children, than simpler family relationships. Establishing current relationships was particularly challenging where family circumstances had changed over time.

Data quality:

Some addresses extracted from unstructured records were incomplete or invalid. Where possible, address validation and standardisation tools were used to improve quality.

Definition of a parent:

Different data sources define parents, carers and dependants differently. For example, Child Benefit data identifies legal claimants, while HMPPS records may include biological, step and other parental relationships. The use of multiple sources helps improve coverage but does not remove these differences entirely.

Coverage of data sources:

Although Child Benefit has high coverage, not all eligible families are represented in HMRC records. In addition, HMPPS sources may not capture all children connected to a prisoner. Children who turned 18 before 1 October 2024 were excluded, even if they spent part of the study period aged under 18.

5. Contact details

You can send enquiries and feedback on these estimates to the team at CAPI@justice.gov.uk

6. Appendix

6.1 Free-text filter keywords

grandma, nan, grandpa, mum, mother, dad, father, parent, aunt, uncle, brother, sister, sibling, cousin, husband, wife, ex, partner, spouse, fiancée, fiancé, fiance, boyfriend, girlfriend, friend, son, daughter, child, children, niece, nephew.

The prefixes half, step, ex, and grand can be added to these where the resulting relation is sensible, e.g. half sibling, ex-partner. The use of language analysis means any similar keywords to those listed are also picked up for example nan and grandmother as well as grandma.

6.2 LLM prompt

<role> 

You are a probation officer collating data about an offender. 

</role>

You will be given the name of an offender and a series of case notes relating to that offender. 

Your task is to extract the personal relationships of an offender from notes written by probation officers. 

Extract the offender’s relationships and corresponding names. 

If the note details multiple names, create a row for each of them. Ensure each row is unique.  

Your output must be in the following JSON format: 



  {{“relationship_type”: “relationship_type”, “name”: “name”, “address”: “address”, “date_of_birth”: “date_of_birth”}}, 

  {{“relationship_type”: “relationship_type”, “name”: “name”, “address”: “address”, “date_of_birth”: “date_of_birth”}} 



The output should only be this list of relationships within the square brackets specified. There should be no preamble. 

If any of the fields cannot be determined then they should be left as null. 

“relationship_type” should be a family relationship such as “son”, “daughter”, “wife”, “ex-husband”; or a social relationship such as “friend”. 

“date_of_birth” should be formatted as “DD/MM/YYYY”. 

“name” should be a proper name and not a synonym for the relationship. For example,  

from the sentence “Fred’s mother visited him on Monday.” we can infer the relationship 

{{“relationship_type”: “mother”, “name”: null, “address”: null, “date_of_birth”: null}} 

“Fred’s mother” or any other synonym for a relationship is not suitable as a name. 

“address” can be in any format. 

There should be only one entry for each person. Consider whether multiple entries 

relate to the same person and if so combine them. 

The following example gives an example of input and expected output. Note that  

this contains dummy data that should not be repeated.