Quality and methodology information (QMI) for healthcare-associated infections (HCAI) reports
Updated 8 October 2026
About this report
This report outlines the quality and methodology information (QMI) relevant to the healthcare-associated infections (HCAI) statistics which are published by the UK Health Security Agency (UKHSA). HCAI statistics are published monthly, quarterly and annually and include:
Accredited official statistics:
- HCAI monthly data sets
- Quarterly epidemiological commentary: Mandatory Gram-negative bacteraemia, MRSA, MSSA and C. difficile infections: hereafter ‘Quarterly epidemiological commentary’
- Annual epidemiological commentary: Gram-negative, MRSA, MSSA bacteraemia and C. difficile infections: hereafter ‘Annual epidemiological commentary’
- HCAI annual data sets
Management information:
- Annual commentary on MRSA, MSSA and Gram-negative bacteraemia and Clostridioides difficile infections from independent sector healthcare organisations in England: hereafter ‘Independent sector report’
This QMI report supports users in understanding the strengths and limitations of these statistics, ensuring UKHSA is compliant with the quality standards stated in the Code of Practice for Statistics. The report covers:
- the strengths and limitations of the data used to produce the statistics
- the methods used to produce the statistics
- the quality of the statistical outputs
About the statistics
These statistics present trends in the counts and rates of the 6 infections which comprise the mandatory surveillance of bacteraemia and Clostridioides difficile (C. difficile) infections (CDI): Escherichia coli (E. coli) bacteraemia, Pseudomonas aeruginosa (P. aeruginosa) bacteraemia, Klebsiella species (Klebsiella spp.) bacteraemia, meticillin-resistant Staphylococcus aureus (MRSA) bacteraemia, meticillin-susceptible Staphylococcus aureus (MSSA) bacteraemia and CDI. The data is summarised by key epidemiological and clinical characteristics.
Geographical coverage: England, integrated care board (ICB) level, sub integrated care board (SICB) level and NHS acute trust level
Publication frequency:
- monthly (monthly data tables)
- quarterly (quarterly epidemiological commentary)
- annual (annual epidemiological commentary, annual data tables, independent sector report)
Change log
8 October 2026: Updated the Routine SGSS-DCS audit section.
24 September 2026: New versions of lower layer super output areas (LSOA) 2021 and Index of Multiple Deprivation (IMD) 2025 in deprivation analysis. Ethnicity analyses now use age-sex standardised rates. CDI-specific denominators have been added.
22 January 2026: New calculation of Quarterly Mandatory Lab Returns (QMLR) data reported in the QEC to reduce the impact of outliers: England-level blood culture positivity and CDI positivity is computed as a median of all available trusts’ values and minor revisions of old terminology.
25 September 2025: Revised the standard population for age- and age-sex standardised rates which now use ONS mid-year estimate populations (for ethnicity and deprivation analyses, respectively) rather than the 2013 European Standard Population. Revised CDI hospital-onset-healthcare-associated (HOHA) definition in the AEC in line with bacteraemia and the QEC. New QMLR blood culture and CDI stool specimen positivity metrics added to AEC. Revised sampling rates which now use England population denominators for England-level metrics instead of bed-days denominator for AEC.
10 July 2025: Added calculations for rate of blood culture sets, overall stool specimens examined, rate of stool specimens examined for CDI diagnosis, pooled E. coli, Klebsiella spp., P. aeruginosa, MRSA and MSSA blood culture positivity and CDI positivity in the Data set production subsection.
10 April 2025: Acknowledged a change of definition in date of admission in the Sound methods section.
5 February 2025: Revised the definition of CDI HOHA, reviewed unlocks and HCAI DCS-SGSS sections with analysis in 2023/2024.
9 January 2025: Added the link to a relevant R package and seasonality considerations to the Sound methods section.
18 December 2024: QMI report first published.
Contact
To contact the team responsible for producing these statistics, please email mandatory.surveillance@ukhsa.gov.uk
Suitable data sources
Statistics should be based on the most appropriate data to meet intended uses.
This section describes the data used to produce the statistics.
Data sources
The primary data source for the numerator data is the mandatory surveillance of bacteraemia and CDI which is collected through UKHSA’s HCAI DCS (data capture system) covering 6 data collections. Mandatory surveillance began in response to increasing rates of MRSA bacteraemia across NHS trusts and was subsequently rolled out for other HCAIs when concern arose. Independent sector providers were also mandated to submit data from April 2009. There is a complete timeline of changes to mandatory surveillance available.
The inclusion criteria for reporting each bacteraemia or infection to the surveillance system are:
- for MRSA bacteraemia, positive blood cultures caused by S. aureus resistant to meticillin, oxacillin, cefoxitin or flucloxacillin
- for MSSA bacteraemia, positive blood cultures caused by S. aureus which are susceptible to meticillin, oxacillin, cefoxitin, or flucloxacillin, and not subjected to MRSA reporting
- for E. coli bacteraemia, all laboratory-confirmed positive blood cultures cases of E. coli bacteraemia
- for CDI (patients 2 years and older):
- diarrhoea stools (Bristol Stool types 5 to 7) where the specimen is C. difficile toxin positive
- toxic megacolon or ileostomy where the specimen is C. difficile toxin positive
- pseudomembranous colitis revealed by lower gastro-intestinal endoscopy or computed tomography
- colonic histopathology characteristic of CDI (with or without diarrhoea or toxin detection) on a specimen obtained during endoscopy or colectomy
- faecal specimens collected post-mortem where the specimen is C. difficile toxin positive or tissue specimens collected post-mortem where pseudomembranous colitis is revealed, or colonic histopathology is characteristic of C. difficile infection
For the annual epidemiological commentary, the mandatory surveillance data is linked to the data sources below to obtain key epidemiological characteristics.
Additional patient characteristics such as patient postcode, GP code and date of death are obtained from linkage with the NHS Spine Summary Care Records data set.
NHS acute trust-level population data does not currently exist in England as NHS acute trusts do not treat patients within defined geographical boundaries. Therefore, a suitable proxy for population is required to calculate hospital-onset (HO) and HOHA rates. The occupied overnight bed-days from the national hospital bed occupancy data set data set (also known as ‘KH03’) provides the daily average overnight bed occupation per year for a specific period: annually from financial year 2007 to 2008 to financial year 2009 to 2010 and per quarter from financial year 2010 onwards. This open access data set is published by NHS England and provides a measure of clinical activity in each trust, which is used as a proxy measure of the patient population. The latest published bed occupancy data may include revisions to previous publications. When there are missing trust quarters in the data we apply a proxy from the same quarter in the previous year.
As C. difficile infection numerators only include people aged 2 years and over, denominators for all outputs are similarly restricted to this age group.
The high-level ethnic group for each case is identified by linking our surveillance records using NHS number and date of birth to NHS England Hospital Episode Statistics Admitted Patient Care (HES APC), HES Outpatient Care (HES OP), HES Urgent and Emergency Care Activities (HES AE), Secondary Uses Service (SUS) records. This approach extends the COVID-19 Health Inequalities Monitoring for England (CHIME) tool.
The Office for National Statistics (ONS) mid-year population estimates up to and including 2024 (the latest at the time of query) are at lower layer super output area (LSOA) level (2021 version) then aggregated for integrated care board (ICB)-level and national analyses too, for the calculation of total, community-onset (CO), community-onset-community-associated (COCA) (and for CDI community-onset-indeterminate-associated (COIA)) incidence rates. For calendar year populations 2025, 2026 and 2027 (required to produce financial year populations up to 2026 inclusive) we assumed they stayed the same as the proxy year 2024.
Data quality
The data that we use to produce statistics must be fit for purpose. Poor quality data can cause errors and can hinder effective decision making.
We have assessed the quality of the source data against the data quality dimensions in the Government Data Quality Framework.
This assessment covers the quality of the data that was used to produce the statistics, not the quality of the final statistical outputs. The quality summary assesses the quality of the final statistical outputs.
Strengths and limitations of the mandatory surveillance data
The strengths of the data
The surveillance is at patient-level including both risk factor data and date of positive specimen, date of inpatient admission and date of recent discharge, which allow for onset location and prior trust exposure to be ascertained. These enhanced data provide a platform to identify potential interventions, which could not be garnered from other surveillance schemes in England.
In addition, the surveillance scheme is a census of all microbiologically-confirmed episodes of the 6 bacteraemias and CDI, which provides up to 2% greater ascertainment than comparative voluntary surveillance schemes (excluding CDI cases, due to issues with voluntary surveillance described in the Routine SGSS-DCS audit sub-section with voluntary laboratory surveillance data).
Well-completed patient identifiers allow for direct linkage with other data sources which make fuller data sets and reduce data entry burden for trusts. For example, data can be linked from the mandatory surveillance scheme with data from the voluntary laboratory reports to access antimicrobial susceptibility information, HES APC for comorbidity information and prior healthcare interactions.
Reporting from the live mandatory surveillance database, HCAI DCS, for registered users such as healthcare professionals provides current statistics and other tabulations or graphical representations of their data.
The limitations of the data
Despite the ability to link the mandatory surveillance data with other data sets, the completion of the data return takes time which leads to variable field completion for the non-mandatory fields and restricts the data’s utility.
There is a potential conflict between the use of these data for epidemiological purposes by UKHSA and performance management or audit by others.
While the effect on data validity is not currently of great concern, as discussed in ‘Mandatory HCAI Surveillance Data in NHS performance management’, the emphasis on performance management surrounding reductions in MRSA bacteraemia and CDI could lead to an emphasis on the infection prevention and control of these infections over others which have not had similar attention.
Summary
The mandatory data completion requirements, combined with long-term surveillance and enriched data, make the mandatory surveillance data the most reliable and suitable source for these statistics.
Accuracy
Accuracy is about the degree to which the data reflects the real world. This can refer to correct names, addresses or represent factual and up-to-date data.
The accuracy of the case-level data submitted to the mandatory surveillance of healthcare associated infections scheme is assured by the chief executive officer (CEO) of all the reporting acute trusts via the monthly sign off process, which was mandated by Chief Medical Officer (CMO) from October 2005. To add or delete cases after the sign-off, the reporting organisation CEO needs to request an unlock to the mandatory surveillance team in a formal process described in the Accuracy and reliability sub-section.
Completeness
Completeness describes the degree to which records are present.
For a data set to be complete, all records are included, and the most important data is present in those records. This means that the data set contains all the records that it should and all essential values in a record are populated.
Completeness is not the same as accuracy as a full data set may still have incorrect values.
Routine SGSS-DCS audit
We undertake routine comparison and quality assurance of HCAI DCS data with voluntary laboratory surveillance data and the Second-Generation Surveillance System (SGSS), which is used by laboratories to report cases of microbial infection from various samples like blood, urine and faeces. Information on antibiotic and antifungal susceptibility is also submitted where relevant. Although primarily an internal system used by healthcare professionals, the data reported via this system is routinely compared to the mandatory data collected via HCAI DCS. This routine comparison between surveillance systems provides a data quality check of case ascertainment on the HCAI DCS.
It is not currently possible to include C. difficile data in the routine HCAI DCS and SGSS comparison as information on these cases is not comparable due to data quality and reporting issues in the SGSS. C. difficile testing is a 2-stage process where the second stage identifies the C. difficile toxin. As only C. difficile toxin-positive cases are reportable to the mandatory surveillance system, it is not currently possible to differentiate reported C. difficile cases which have tested positive for C. difficile toxins from those which have not with an acceptable degree of accuracy from the SGSS.
In general, more cases are captured via the HCAI DCS than the SGSS. Meticillin resistance in the mandatory surveillance is reported by NHS acute trusts after susceptibility testing but meticillin resistance in the SGSS is determined by selecting the most severe susceptibility results from patients’ blood cultures within a 14-day period. This difference explains some of the apparent over-ascertainment of the voluntary MRSA reports in some financial years, and therefore this should be considered when comparing the case numbers for MRSA and MSSA bacteraemia for SGSS versus HCAI DCS.
Not all cases in the SGSS would be reported to the HCAI DCS. SGSS cases for each bacteraemia and infection reported to the HCAI DCS are defined as:
- for MRSA bacteraemia, the earliest S. aureus blood isolate per patient within a 14-day period with resistant or indeterminate result to meticillin, oxacillin, cefoxitin or flucloxacillin result within the 14-day period
- for MSSA bacteraemia, the earliest S. aureus blood isolate per patient within a 14-day period with a susceptible result to meticillin, oxacillin, cefoxitin or flucloxacillin result within the 14-day period
- for E. coli bacteraemia, the earliest E. coli blood isolate per patient within a 14-day period
- for Klebsiella spp. bacteraemia, the earliest Klebsiella spp. (including K. aerogenes, formerly known as Enterobacter aerogenes) blood isolate per patient within a 14-day period
- for P. aeruginosa bacteraemia, the earliest P. aeruginosa blood isolate per patient within a 14-day period
Ordered matching between cases from HCAI DCS and SGSS show only about 5% of infections captured by the HCAI DCS cannot be found in the SGSS.
As part of routine laboratory data checks, laboratories with cases reported to the SGSS but not identified in the HCAI DCS are contacted for feedback on the discrepancy. The cases are closed if:
- the unmatched case is subsequently identified in the HCAI DCS
- the unmatched case is added to the HCAI DCS as a new record
- there is a legitimate reason for not reporting it to the HCAI DCS
Accounting for the open cases identified in the SGSS, the HCAI DCS captures an estimated 99%, 97%, 97% and 97% of E. coli, Klebsiella spp., P. aeruginosa and S. aureus bacteraemia cases, respectively, which are eligible for mandatory reporting, suggesting the HCAI DCS provides an accurate national picture of the overall burden of infection which is under mandatory surveillance in England.
Uniqueness
Uniqueness describes the degree to which there is no duplication in records. This means that the data contains only one record for each entity it represents, and each value is stored once.
Some fields, such as National Insurance number, should be unique. Some data is less likely to be unique, for example geographical data such as town of birth.
The HCAI DCS has a de-duplication algorithm where cases reported by the same reporting organisation with matching NHS number and date of birth are flagged to the reporting trust to determine whether the case is a true duplicate. The CEO of the reporting organisation is required to sign-off data monthly which provides an additional verification of the uniqueness of the data. Technically speaking the data is inputted on an episode basis, however following these self deduplication checks we refer to this data in reports as ‘cases’ rather than ‘episodes’.
Consistency
Consistency describes the degree to which values in a data set do not contradict other values representing the same entity. For example, a mother’s date of birth should be before her child’s.
Data is consistent if it does not contradict data in another data set. For example, if the date of birth recorded for the same person in 2 different data sets is the same.
The HCAI DCS includes various validation rules which prevent the entry of invalid dates such as disallowing a specimen date preceding a patient’s date of birth. In such cases, a meaningful validation error message will be displayed to the data-entry user to correct the input before proceeding.
Timeliness
Timeliness describes the degree to which the data is an accurate reflection of the period that it represents, and that the data and its values are up to date.
Some data, such as date of birth, may stay the same whereas some, such as income, may not.
Data is timely if the time lag between collection and availability is appropriate for the intended use.
Cases entered in the HCAI DCS require sign-off by the CEO of the reporting organisation by the end of the 15th day of each month for the previous month’s data. Hence there is minimal delay between the data collection and availability. Organisations can input the data at any time but are mandated to report cases for a given month within 15 calendar days following it; for the majority of organisations that meet these reporting deadlines, this leads to a maximum reporting delay of 6 weeks for those cases with a specimen date at the start of the month; and a maximum publication delay of these count data of 2 months for our monthly reports.
Validity
Validity describes the degree to which the data is in the range and format expected. For example, date of birth does not exceed the present day and is within a reasonable range.
Valid data is stored in a data set in the appropriate format for that type of data. For example, a date of birth is stored in a date format rather than in plain text.
HCAI DCS applies data validation for many fields to prevent incorrect information being entered to ensure accuracy. For example, the ‘specimen date’ must be in the correct date format, otherwise an error message will be displayed to resolve before progressing further. Whenever possible, drop-down lists are used to minimise data entry errors.
Sound methods
Statistical outputs should be made using the best available methods and recognised standards.
This section describes how the statistics were produced and quality assured.
Data set production
All cases of bacteraemia and CDI originate in the hospital (HO) or community (CO).
A case of bacteraemia or CDI is classified as HO if it meets all the following criteria:
- the patient is an in-patient, day-patient, emergency assessment patient or unknown location
- the specimen was taken at an acute trust or at an unknown location and
- the specimen was taken on or after day 3 of the admission (admission date is considered day 1)
Cases that do not meet all the above criteria are categorised as CO.
Note that the HO and CO definitions only exist for cases reported by acute trusts. A similar stratification is not available for independent sector provider organisations.
Any publications prior to 23 January 2025 will use the previous HO definition for CDI which included specimens taken on or after day 4 of the admission (admission date is considered day 1).
It is not possible for UKHSA to change the onset status of a case as it is determined by the above criteria based on the data provided by the reporting organisation. A case may change from one category to another only if the relevant case details are incorrect and this would require an amendment by the trust. Reports published before September 2017 used the term ‘trust-apportioned’ for HO cases and ‘not trust-apportioned’ for CO cases which was simply a change in terminology.
A SICB for each case is attributed in the following order:
- if the patient’s GP practice code is available (and is based in England), the case will be attributed to the SICB at which the patient’s GP is listed, or
- if the patient’s GP practice code is unavailable but the patient is known to reside in England, the case is attributed to the SICB catchment area in which the patient’s residential postcode lies, or
- if both the patient’s GP practice code and residential postcode are unavailable or if a patient has been identified as residing outside England, then the case is attributed to an SICB based on the postcode of the headquarters of the acute trust that reported the case
From the attributed SICB all cases are attributed to a commissioning ICB. ICBs cover a specific geographical area and are NHS organisations that are responsible for planning health services for their local population.
All cases of bacteraemia and CDI are attributed to a SICB and then commissioning ICB regardless of onset. UKHSA’s HCAI DCS does not need the NHS organisations to record a patient’s SICB details for a case as an automated trace is performed to obtain this information—an extract comprising patient NHS number, date of birth, patient forename, patient surname and sex are submitted to NHS Digital via Demographics Batch Services (DBS) tracing service twice daily and matched with a multi-stage algorithm.
As of 1 April 2026, changes were made to the SICB and ICB organisational structures. These changes included the closure of one SICB and 12 ICBs, and the creation of one new SICB and 6 new ICBs. Therefore, for the AEC 2025-26, organisational tables were presented using both the ‘old’ organisational structures from before 1 April 2026 (as of 1 January 2026), and the ‘new’ organisational structures as of 1 April 2026.
Cases are categorised into one of the following prior trust exposure groups, which also detail the denominators they use when computing incidence rates:
- HOHA: NHS patient specimens taken from day 3 of admission onwards, where day 1 is the day of admission, with Patient Location of ‘NHS Acute Trust’ or ‘Unknown’. Applies to patients with a Patient Category of ‘Inpatient’, ‘Day patient’, ‘Emergency Assessment’, or ‘Unknown’. Includes records with an unknown admission date, where the Patient Location is ‘NHS Acute Trust’ or ‘Unknown’. Hospital overnight bed-days are used as a denominator as the patient has already been admitted to hospital.
- Community‑onset healthcare-associated (COHA): any case reported by an NHS acute trust not determined to be hospital-onset healthcare associated but where the patient was discharged from the reporting organisation within 28 days prior to the current specimen date (for bacteraemia and CDI), where date of discharge is day 0). Hospital overnight bed-days and hospital day-only are used as a denominator. The addition of ‘day only’ accounts for community cases who have not been admitted and may initially present as day-only.
- Community-onset indeterminate association (COIA) for CDI cases only: any case reported by an NHS acute trust not determined to be hospital-onset healthcare Associated but where the patient was discharged from the reporting organisation between 28 and 84 days prior to the current specimen date (where date of discharge is day 01). ONS mid-year populations are used for the denominator.
- Community‑onset community-associated (COCA): any case reported by an NHS acute trust not determined to be hospital-onset healthcare associated but and where the patient has not been discharged from the reporting organisation within the past 28 days (or 84 days for CDI), to the current specimen date (where date of discharge is day 0). ONS mid-year populations are used for the denominator.
- Missing: any case reported by an NHS Acute Trust not determined to be hospital-onset healthcare-associated but where the information on the patient’s prior discharge was not reported. As of April 2019, it is no longer possible to leave these questions blank.
- Unknown: any case reported by an NHS acute trust not determined to be hospital-onset healthcare-associated but where the information on the patient prior discharge was reported as ‘Don’t know’.
An R package for working with data downloaded from the HCAI DCS can be found on GitHub.
From 1 April 2024 following discussions with NHSE partners, trusts were required to enter the decision-to-admit date as the admission date rather than the inpatient admission date for patients who were admitted to A&E. This is to account for patients spending increasingly longer periods in A&E and thereby increasing the risk of infection who would otherwise be categorised as CO cases rather than as HOHA cases following a positive specimen.
Monthly data tables
These data tables include a monthly count of total reported cases for each data collection as well as a breakdown by prior trust exposure for the last 13 months. The counts are reported at England, ICB, NHS acute trust, UKHSA centre and NHS region levels. The data tables also record whether the data was signed off.
Quarterly epidemiological commentary (QEC)
The incidence rate of total and CO cases is calculated using their quarterly count and the mid-year population for England. It is converted to an annualised incidence rate to allow comparisons with annual incidence.
It is calculated as the count of reported episodes in England in a given quarter divided by the mid-year population of England in that year, multiplied by the number of days in that year, divided by the number of days in that quarter and multiplied by 100,000.
The incidence rate of HO cases is calculated using their quarterly count and the bed occupancy average bed-day activity for England.
It is calculated as the count of reported episodes in a given quarter in England divided by the daily average number of occupied overnight beds in that quarter in England, then divided by the number of days in the same quarter and multiplied by 100,000.
Percentage changes in rates when presented in the commentary have been rounded to one decimal place. Conversely, graphs in our reports and accompanying data (like in our QEC) use unrounded calculation numbers.
Annual epidemiological commentary (AEC)
The Office for National Statistics (ONS) mid-calendar year population estimates are used to calculate the financial year population. For instance, for financial year 2025 to 2026 mandatory surveillance data, we use the mid-calendar year 2024 population estimate. At the time of query the population estimates for calendar years 2025 onwards were unavailable at the small area population (SAPE) level, so the most recent (mid-2024) estimate was used. Most analyses rely on the SAPE level extracted at LSOA level which is then aggregated up to SICB, ICB or England level.
Rates at the SICB, ICB or England level are calculated as the number of new cases within that area, divided by their population and multiplied by 100,000, expressed as ‘per 100,000 population’.
Hospital bed occupancy data from NHS England is used as an indicator of the total activity in each trust during the relevant periods and is used in the denominator of acute trust rates. It has been published quarterly since April 2010.
Hospital bed occupancy data is sometimes missing for some trust quarters’ ‘overnight occupied beds’ and/or ‘day-only admissions’. In these cases we generate the proxies for these using the most recent non-missing quarter from their previous years. For example, due to 5 trusts missing 2025 to 2026 FQ4 ‘overnight occupied beds’ data the AEC 2025 to 2026 used proxies as described. The denominator of acute trust rates for all cases or HOHA, HO or CO ones use the overnight occupied beds metric. However, for COHA cases, the denominator is the ‘overnight occupied beds’ plus ‘day-only admissions’ metrics then multiplied by the number of days in the relevant period.
The acute trust rate is then the number of new cases reported by the trust, divided by the relevant denominator multiplied by 100,000: the rates of all, HO and HOHA is expressed as ‘per 100,000 bed-days’ while for COHA it is ‘per 100,000 bed-days and day admissions’.
Prior to trust apportioning, the rates for all cases were calculated per acute trust using ‘per 100,000 bed-days’. Therefore, to retain the historical time series, we continue to provide an all-cases rate per acute trust. ‘All reported cases’ refers to all bacteraemias or C. difficile infections that are detected by the acute trust that processed the specimen. It does not necessarily imply the infection was acquired there.
To calculate time-to-onset of an episode (bacteraemia or CDI) among inpatients, the number of days between the date of admission to an NHS acute trust and the date of positive specimen are used. This was performed for only patients who were admitted to an NHS acute trust and for those whose specimen was taken on or after the date of admission also at an acute trust. The number of days between the date of admission and the date of specimen is then grouped into meaningful categories by the number of days. To assess seasonal trends, case data were aggregated by financial quarter, organism and onset type (HO, CO and total cases) while observations from years before financial year 2010 to 2011 were excluded. Percentages were calculated by dividing the number of cases for each organism, financial quarter and onset type by the total cases for that organism and onset type within the corresponding financial year and multiplied by 100.
To calculate the age-sex stratified denominator for incidence rates by deprivation, ONS mid-year populations of 1-year age-sex bands (from 0 to 89 years, then 90 years plus) at LSOA level from financial year 2017 onwards (and calendar year proxies from 2025 onwards) and linked to IMD 2025 (Index of Multiple Deprivation) deciles from the Ministry of Housing Communities and Local Government. IMD deciles are converted to quintiles. The IMD quintile for each case is assigned using firstly the patient’s postcode on the HCAI DCS at the time of infection and if missing then their GP’s postcode. The postcode is then linked to LSOA and thus IMD quintile. Calendar year populations are converted to financial year populations.
Age-sex stratified denominators for incidence rates by ethnicity are calculated similarly to deprivation apart from the following details. The 2021 Census populations provides 1-year age-sex bands by 20 minor ethnic groups. The 20 minor ethnic groups are aggregated up to the 5 major ethnic groups (Asian, Black, Mixed, Other and White). The proportions of each age, sex, and ethnic group observed in calendar year 2021 were applied to the respective mid-year ONS populations for calendar years 2017 to 2020 and 2022 to 2024 and subsequent proxy years to generate age-sex-ethnicity proxy denominators—this rescaling using the 2021 census population to achieve the age-sex stratified ethnic populations in the calendar period 2017 onwards assumes that England’s age, sex, and ethnic distributions have not changed substantially since 2017 and that the 2021 census year is a good representation of this period.
The observed (crude) incidence rates were calculated for each organism and financial year as the number of infections in a financial year in each IMD quintile or major ethnic group divided by the population in that IMD quintile or ethnic group, multiplied by 100,000.
LSOA-level age-sex stratified deprivation denominators and national age-sex stratified ethnicity population denominators are a prerequisite for computing their directly standardised rates (DSRs). DSRs (age-sex standardised) by deprivation or ethnicity are estimated using the PHEindictormethods R package (according to national guidance) with confidence intervals using Byar’s method.
DSRs are not calculated for counts less than 10 where the method becomes unreliable, whereas crude rates (labelled as the ‘Observed’ rate in the deprivation and ethnicity analyses) do not have this restriction. For statistical consistency with the DSRs, confidence intervals for crude rates in deprivation and ethnicity analyses use Byar’s method throughout the time series; most of the data points of these time series involve sufficiently large counts to warrant this method even though a few data points will have low counts that are normally incompatible with Byar’s method or will be unable to compute.
Cases without a known IMD quintile or ethnic group (8.2% and 5.8% of all cases, respectively) are excluded from the calculation of deprivation and ethnicity rates. This means that rates stratified by deprivation or ethnicity are slightly underestimated.
A missing IMD quintile value may be because:
- the patient’s residence was not in England
- the patient’s residence was in a new area that was not assigned to the IMD 2025 decile data set or
- the patient was homeless
A missing ethnic group value may be because:
- the patient had chosen not to state their ethnic group on admission to hospital
-
the trust did not record a valid NHS number, date of birth, name or sex on the HCAI DCS which are important for a successful trace or
-
HES data sets did not contain an ethnicity value
- Age-sex DSRs for all cases (without deprivation nor ethnicity stratification) are also provided at the ICB level using similar methods as described above.
All case-level data (not just the most recent year) is sent for retracing on the DBS tracing service to correct date of birth, sex and date of death values in case they have been corrected since initial trace. This will cause small changes in historical age-sex estimates, mortality, deprivation and ethnicity analyses, and CDI counts (if the revised age would drop a patient due to being aged under 2 years on the specimen date).
Mortality rate is used for assessing risk of death and is calculated by dividing the number of deaths by the population at risk. This reflects the incidence of all-cause deaths following these infections in the population. Case fatality rate is a measure for comparing survivability of different infections and is expressed as the number of deaths as a percentage of all reported cases. Data is presented on all-cause mortality, and therefore includes deaths that may not be directly attributable to the infections.
When using mortality sub-analyses such as by age and sex, onset or region (accompanying tables S10 to S12), aggregate totals should not be calculated as they are expected to differ from the official value quoted in the main analysis (accompanying table S9). This is caused during the calculation of the thirty-day all-cause deaths which inherently has rounding error due to a ceiling function to provide an integer result and further amplified by subsequent aggregation.
Since inclusion of C. difficile Ribotype Network (CDRN) data in financial year 2023 to 2024, we updated the analysis to focus on the perspective of the HCAI DCS case as the subject and thus common denominator of most of the metrics (accompanying tables S17 and S18). For example, in Table S18 we produce row percentages instead for the ribotype-specific proportion; this supports the data structure since an HCAI DCS case linked to a CDRN sample can in rare situations return multiple ribotypes per sample.
Three QMLR indicators measuring the total number of blood culture sets examined, total number of stool specimens examined, and total number of stool specimens examined for diagnosis of CDI are included, and feature in the QEC and AEC.
The quarterly population of England is the mid-year population of England in that year, divided by the number of days in that year and multiplied by the number of days in that quarter.
The blood culture sampling rate per 1,000 population is calculated as the number of blood culture sets examined divided by the quarterly or annual population of England and multiplied by 1,000.
The overall stool specimens examined rate per 1,000 population is calculated as the number of stool specimens examined, divided by the quarterly or annual population of England and multiplied by 1,000.
The stool specimens examined for CDI diagnosis rate per 1,000 population is calculated as the number of stool specimens examined for diagnosis of CDI divided by the quarterly or annual population of England and multiplied by 1,000. Note we do not restrict the population denominator to those aged 2 years or over like other CDI rates because the labs who collect the numerator do not restrict by age when collecting the data.
The pooled blood culture median positivity percentage is calculated for each trust by dividing the total number of cases of E. coli, Klebsiella spp., P. aeruginosa, MRSA and MSSA bacteraemia reported per trust via the HCAI DCS by the total number of blood cultures examined in the respective quarter or financial year per trust and multiplying by 100. The median of all trusts’ positivity percentages is then calculated.
The CDI median positivity percentage is calculated for each trust by dividing the total number of CDI cases reported per trust via the HCAI DCS by the total number of stool specimens examined for CDI diagnosis in the respective quarter or financial year per trust and multiplying by 100. The median of all trusts’ positivity percentages is then calculated.
It is worth noting that QMLR indicators are affected by variation in trust reporting. For example, delayed or incomplete reporting of QMLR indicators by trusts can lead to apparent decreases in the rates of blood culture sampling or stool specimens examined, particularly in the most recent quarter at that time.
Independent sector (IS) report
Counts and rates (per 100,000 bed-days and discharges) of MRSA, MSSA, E. coli, Klebsiella spp., P. aeruginosa bacteraemia and CDI are presented by IS organisation for the latest 12-month period with comparison of rates to the previous year.
An IS organisation can comprise a group of private hospitals owned by one company or a single private hospital. It is possible to identify a group versus a hospital using the ‘number of hospitals in organisation’ field in the HCAI DCS.
The modified inpatient bed-days (bed-days plus discharges) are provided for the most recent financial year available as an indication of the size of each facility.
Hospitals are categorised as ‘large’ (50 beds or more) or ‘small’ (fewer than 50 beds). NHS treatment centres and diagnostic centres seeing mainly day case patients are listed for the hospitals within a group. All types are listed where a group comprises more than one hospital type. IS organisations are requested to submit their bed-day plus discharge denominators by survey. The calculation for the bed-day plus discharge denominator for shorter stay hospitals is the sum of the number of bed days in a year and the number of discharges in a year.
Instead of counting the number of midnights the patient was resident for, this counts the number of different days on which they were in the hospital. A day case will count as 1, a one-night stay in the year will count as 2.
Bed-days in the financial year April 2023 to March 2024, for example, is the sum of the number of beds occupied each midnight during the year. So this sum starts with the number of bed occupants at midnight for the day ending 1 April 2023 and ends with the number of bed occupants at midnight on 31 March 2024. Alternatively, in that financial year if the bed-days is being derived from admission dates and discharge dates, the calculation is the discharge date or 1 April 2024 (whichever is earlier) minus the admission date or 1 April 2023 (whichever is later).
Only patients who are admitted to hospital before 1 April 2024 say, and discharged on or after 1 April 2023 are counted towards a bed-day in that financial year. That is, the latest date they could have been admitted was 31 March 2024 and the earliest date they could have been discharged was 1 April 2023. If the patient is still in hospital and does not yet have a discharge date then, 1 April 2024 should be used as discharge date. The sum of the days for all the patients then provides the total number of bed-days. Discharges in that financial year include the number of patients with a discharge date between 1 April 2023 and 31 March 2024. It is the sum of the number of patients discharged on 1 April 2023 and the number discharged for each subsequent day up to and including 31 March 2024. It should include any day cases that took place during the year.
Quality assurance
Statistical processing for official statistics is performed independently by 2 scientists and final data cross-checked to verify that the data are correct. In addition, when rates are calculated for our QEC and AEC reports, infographics and data tables. We also cross-check the processing of NHS England’s hospital bed occupancy data (occupied overnight bed days and day admissions) (promptly feeding back any errors to NHS England) and ONS’ population denominators.
Confidentiality and disclosure control
Personal and confidential data is collected, processed, and used in accordance with the UKHSA privacy notice. All UKHSA staff with access to personal or confidential information must complete mandatory information governance training, which must be refreshed every year. Information is stored on computer systems that are kept up-to-date and regularly tested to make sure they are secure and protected from viruses and hacking. UKHSA staff do not store data on their own laptops or computers. Instead, data is stored centrally on UKHSA servers.
No personally identifiable information is included in the published data. The structure of the published tables prevents them from being broken down in ways that could compromise individual privacy through cross-referencing. Additionally, when small numbers are reported in the data, a careful assessment is conducted to balance the need for detailed reporting with the potential risk of statistical disclosure, ensuring privacy is maintained without compromising the usefulness of data.
Geography
Mandatory surveillance includes data from all NHS trusts in England. Each report contains data for overall counts and rates (except for monthly tables which include counts only) at different geographic levels. Monthly tables are published at national (England), ICB, UKHSA centre, NHS region, and NHS acute trust levels. The QEC is published at national (England) level. The AEC is published at national (England), ICB, SICB and NHS acute trust level and features crude and age-sex standardised rates at the ICB level.
Quality summary
The Code of Practice for Statistics defines quality in statistics as:
- fitting their intended uses
- based on appropriate data and methods
- not materially misleading
Quality requires skilled professional judgement about collecting, preparing, analysing, and publishing statistics and data in ways that meet the needs of people who want to use the statistics.
This section assesses the statistics against the European Statistical System dimensions of quality.
Relevance
Relevance is the degree to which the statistics meet user needs in both coverage and content.
These mandatory surveillance outputs are critical to tracking progress towards controlling key HCAIs. In particular, the National Action Plan for AMR 2024–2029 and 2019–2024 before it, set out the ambition to control Gram-negative bacteraemia including the 3 infections covered by this surveillance (E. coli, Klebsiella spp. and P. aeruginosa).
The data also allows for NHS acute trusts to monitor their infection rates, and benchmark against peers and nationally.
The different statistics published are used in the following ways:
- mandatory HCAI surveillance outputs help monitor progress on controlling key HCAIs and for providing epidemiological evidence to inform action to reduce them
- mandatory surveillance outputs are routinely used to appraise local/regional NHS management of infection levels within their area
- data provide unique case-level information
- data are used to support the NHS objective of improving the quality and safety of health services and promoting patient choice by providing access to information on NHS performance
- data are used nationally for benchmarking purposes and for the performance management of MRSA bacteraemia and CDI objectives set by NHS Improvement
- data and outputs are routinely used to answer relevant Parliamentary Questions
- data are used to inform patient choice via the NHS Choices website
- NHS acute trusts and SICB locations use these data to monitor progress against these objectives and to inform action on reducing these infections locally
- the E. coli, Klebsiella spp. and P. aeruginosa bacteraemia surveillance outputs are an integral part of NHS Improvement’s strategy to prevent increase in Gram-negative bloodstream infections by 2029 compared to the 2019 to 2020 financial year baseline, as part of the UK National Action Plan for AMR 2024 to 2029 which superseded the previous UK National Action Plan for AMR 2019 to 2024
We have continued to make changes to the publications to meet user needs:
- from the financial year 2025 to 2026 AEC, we have added age-sex standardised incidence rates for ethnicity in addition to deprivation. Ethnicity linkage method also improves on CHIME specification on the multiple data sets we link to. All CDI rates use CDI-specific population denominators which are restricted to people aged 2 years and older
- from the financial year 2024 to 2025 AEC, we have included age-sex standardised deprivation rates, improved linkage for deprivation and ethnicity analyses and updated our standard population from the European Standard Population of 2013 to England standard populations by financial year. We have added positivity metrics for blood culture and CDI, the latter will be a useful additional metric for use during CDI outbreaks
- from the financial year 2023 to 2024 AEC, we have also included analysis of 2 QMLR indicators: total blood culture sets and stool specimens examined for CDI diagnosis tests
- from the financial year 2022 to 2023 AEC, we have added age-standardised incidence rates by deprivation and ethnicity
Accuracy and reliability
Accuracy is the proximity between an estimate and the unknown true value. Reliability is the closeness of early estimates to subsequent estimated values.
Infection cases are reported by NHS acute trusts. As part of the verification process, the CEO of the acute trust signs off infection data reported each month by the end of day 15 of the following month. This sign-off process provides formal assurance that the data are accurate and complete.
Published statistics should include details of all cases for the reported period. However, on occasion, an amendment is required for the following reasons:
- full laboratory results are pending at the time of original sign-off
- deletions are required as an acute trust has entered case information incorrectly or it is a duplicate of a case reported from another trust. In that scenario, a CEO must request the deletion of the wrong information to be replaced with the correct one. NHS acute trusts or external agencies like The Care Quality Commission may also perform audits of local infection data. This can result in requests to add infection episodes that had not previously been entered.
- NHS acute trusts may request to alter their data to improve the SICB attribution of a given infection record. This process is undertaken via an ‘unlock’ of the HCAI DCS. A log of the number of unlocked cases by data collection and unlock reason is maintained.
A total of 72 (53%) acute trusts requested an unlock of at least 1 case across all organisms affecting data in the financial year 2023 to 2024 which totalled 269 unlocked cases. 38.8% of those unlocks were additions, 5.5% were amendments and 61.7% were deletions to a locked period. Comparing the previous financial year 2022 to 2023 with 2023 to 2024, there was an increase of 8.3% in the number of trusts which requested unlocks to change their data but a 9.4% decrease in the total number of unlocks. The number of unlock requests to add and delete a new case declined by 22.4% and 38.0%, respectively, while the number of requests to amend a case increased by 63.5% (63 to 103 requests).
The HCAI DCS includes functionality for acute trusts to identify duplicate infection episodes within their trust. A pop-up for potential duplicates at case entry appears to prompt the trust user that no duplicates have been entered for a designated period. Following sign off, as the CEO of an acute trust has verified their data as being accurate, data used for statistical publications are not altered by the UKHSA mandatory HCAI surveillance team to remove potential duplicate records. This may result in multiple listings of the same infection episode in the data set.
There is a possibility that some cases may not be reported to the HCAI DCS, resulting in under-coverage. To ascertain the level and to rectify this, a consistency study (the ‘routine SGSS-DCS audit’) is performed comparing SGSS with the HCAI DCS.
Data changes between releases are highlighted in each publication, so that users are made aware of any changes to historical data between publications. Further information on this process is available on the caveats page of each routine publication.
Not all IS organisations sign off their data or submit data for the reporting period, potentially leading to unfinalised and inaccurate data.
Measurement error
All mandatory HCAI surveillance data is collected via the HCAI DCS. The appendices of the mandatory HCAI surveillance protocol detail definitions and guidance on each field in the data collection. Therefore, there should be little concern over the interpretation of the questions by different users, although it should be noted that some questions are subjective in nature such as asking the clinical opinion of the treating physicians.
There is a low number of records with ‘item’ non-response errors as the bulk of data used to produce the mandatory HCAI surveillance outputs are from mandatory questions in the HCAI DCS, i.e. a response is required to save the infection episode. The exceptions are in the data collected on risk factors for bacteraemias presented in the AEC because the risk factor or source of bacteraemia questions are not mandatory fields. However, there are accompanying statements in the relevant sections of the AEC on the level of response for these data. However, ‘unit’ non-response where individual NHS acute trusts who have not entered data and/or signed off data does exist. All trust-level outputs highlight such non-responders. Consistent non-responders are further referred to NHS England for follow-up.
Processing error
Processing errors may occur during the data entry stage. The data collected via the HCAI DCS is either entered by hand or partially uploaded (key responses to questions required to save an infection episode) using the HCAI DCS data upload wizard. Data entry errors may occur because the source data at the acute trust is incorrect or missing or in the transcription process.
While it is not possible to provide a level or direction of bias through processing errors for the entire data collection, it is possible to estimate the collective level of processing errors for 2 key variables - date of birth and NHS number. These can be used as an indicator for the full data collection. Assessing the percentage of all cases which could not be attributed via a match with the NHS Spine provides an indication of data entry errors.
There is the potential for bias in the statistics as organisations aim to meet performance targets. Therefore, there is a conflict between the use of statistics for both epidemiology and public health, and for performance management.
Timeliness and punctuality
Timeliness refers to the time gap between publication and the reference period. Punctuality refers to the gap between planned and actual publication dates.
Mandatory HCAI surveillance data is published in as timely a manner as possible. Data should be signed off by acute trusts’ Chief Executives by the end of day 15 of the following month; note we still use unsigned-off data in our reports but mention in data tables its signed-off status while also chasing trusts for unsigned off data. Data is published on a monthly, quarterly and annual basis and are pre-announced at least 28 days in advance, in line with the Code of Practice for Statistics.
The UKHSA official statistics publication calendar includes mandatory HCAI surveillance-specific announcements.
Monthly data tables
Monthly data is processed and analysed before being published on the first Wednesday of the following month. This occurs between 2 and 6 weeks following the end of a given month, depending on how the month falls. For example, January 2017 data was signed off on 15 February 2017 and then published on 1 March 2017. This is 2 weeks from sign-off to publication.
QEC
The QEC is published approximately 3 months following sign-off of the last full month of data for inclusion. Publication of this report occurs on the first Thursday of the fourth month after the quarter covered in the reported. For instance, data up to and including December 2021 were signed off on 15 January 2022 and published on 7 April 2022.
Annual data tables and AEC
Annual data tables and the accompanying AEC is usually published in September each year. The annual data tables include counts and rates for both acute trusts and SICBs. The AEC represents the most substantial HCAI mandatory surveillance output during the year. The lead time necessary for analysis and compilation of data is considerable. Decreasing the amount of time between sign off and publication of these reports has been considered. However, doing so would not allow enough time to undertake relevant data quality checks on either the data used for preparing the report or the report itself. Hence the benefit of using the current publication schedule far outweighs any minor benefit that might be achieved in reducing the lead time for the AEC publication.
Accessibility and clarity
Accessibility is the ease with which users can access the data, also reflecting the format in which the data is available and the availability of supporting information. Clarity refers to the quality and sufficiency of the metadata, illustrations and accompanying advice.
All HCAI outputs have been reviewed for accessibility requirements, with several changes made to ensure they are accessible. Since QEC 2021 financial quarter 4 and AEC financial year 2021 to 2022, they have been published in HTML format which provides the accessibility features mentioned in the GOV.UK accessibility statement. This format enhances accessibility by supporting screen readers and allowing easy navigation using a keyboard, ensuring the content is accessible to a wider audience. Additionally, HTML allows for text resizing and media alternatives such as alt text which help to improve the overall user experience. The publications have also been reviewed for clarity, incorporating plain English language, main messages and data visualisations.
The reports include data visualisations which help users to understand the data. These have been reviewed and updated to ensure colours used provide sufficient contrast to be distinguished and are colour-blind friendly. The accompanying data tables are published in ODS format and follow accessibility guidelines. Each sheet contains only one table and no nested tables. Each data table contains a Contents worksheet and a Notes worksheet to describe the subsequent data tables.
Coherence and comparability
Coherence is the degree to which data that are derived from different sources or methods, but refer to the same topic, are similar. Comparability is the degree to which data can be compared over time and domain.
The mandatory HCAI surveillance scheme aligns closely with surveillance processes and definitions of the European Centre for Disease Prevention and Control (Europe) and the Centre for Disease Control (USA), to allow comparability where possible.
There are, however, some differences between the English mandatory HCAI surveillance programme and the surveillance undertaken by others, including the UK devolved administrations and internationally. These include some case definitions and protocols for diagnosing the infections, definitions on inpatient episode versus trust apportioned or assigned episodes, age groups included in the surveillance schemes and the way in which data are presented by periods. As the population sizes of the other devolved administrations are different to England, crude counts of infections cannot be compared amongst countries in the UK. Furthermore, as the population demographics amongst the devolved administrations differ, the denominators used to calculate infection rates are not directly comparable. Our introduction of age-sex standardised rates by ICB since financial year 2024 to 2025 has at least helped to account for demographical differences across England.
Uses and users
Users of statistics and data should be at the centre of statistical production, and statistics should meet user needs.
This section explains how the statistics are used, and how we understand user needs.
Appropriate use of the statistics
These statistics present information on cases reported to mandatory surveillance since the start of surveillance for each infection. Data is presented for transparency and to allow tracking of these mandated infections across multiple settings.
The onset algorithm, and since 2017 the prior trust exposure algorithm, helped us align more closely with the European Centre for Disease Control (Europe) and the Centres for Disease Control (USA) sought to attribute cases to either a hospital or community setting and whether a case was healthcare-associated to identify where the infection may have occurred. Prior healthcare exposure definitions require trusts to enter the patient’s exposure in only the reporting trust which may lead to an underestimation of healthcare association as patients may have had contact with other healthcare settings which is not captured. Despite this, this definition has been consistently applied across all surveillance years. Also, when comparing data with other UK nations or countries, it is important to consider any differences in infection definitions, onset and prior trust exposure algorithms, and the deduplication window used.
The IS report does not provide a basis:
- for comparisons between different IS organisations due to their variable size and range (case mix) of patients admitted
- for reliable comparison of these infections between the NHS and IS organisations
Known uses
We are aware that the statistics have been used for:
- monitoring progress on controlling key HCAIs and for providing epidemiological evidence to inform reduction actions
- education and training
- strategy and resource allocation
- benchmarking purposes and for the performance management of MRSA bacteraemia and CDI objectives
- research
- informing patient choice
Known users
National users
UKHSA use the data to:
- undertake epidemiological analyses at national, regional and local level
- provide, on request, relevant response to parliamentary questions
Department of Health and Social Care (DHSC) use the data to:
- routinely brief ministers on national and regional incidence of MRSA, MSSA, E. coli, Klebsiella spp. and P. aeruginosa bacteraemia and CDI
- inform and identify national-level targets for interventions and reduction strategies
NHS England and NHS Improvement use the data to:
- identify and establish performance management
- set national and local-level performance management targets
- assess performance against objectives
Regional or local users
ICBs use the data to assess NHS trust and SICB performance against targets and objectives at a local level.
UKHSA Field Service and UKHSA regions use the data to:
- assist in outbreak investigations
- inform public health initiatives at a local level
NHS acute trusts use the data to:
- inform trust boards of their position on key HCAIs (MRSA, MSSA and E. coli bacteraemia and CDI)
- monitor progress against performance management objectives
SICBs use the data to:
- monitor progress against performance management objectives
- assist in the commissioning of services from relevant acute level providers
User engagement
Since AEC 2023 to 2024 we have provided an online survey form at the top of the report to collect readers’ feedback.
A routine ‘Stakeholder Engagement Forum’ is held every 6 months. This meeting includes representation from a wide range of national and local level stakeholders as such as SICB and acute trusts. The meeting’s agenda includes recent publications, experiences, improvements and future developments.
Following the meeting, a summary of the stakeholder engagement forum discussion is available.
Meeting feedback is used to improve ongoing engagement. It is also used to inform future development and to ensure that data users remain central to the process.
Related statistics
Most health protection functions in the UK are devolved to the other UK nations’ public health agencies. Here are examples of their output, including from neighbouring European countries: