Office of the Public Guardian: Investigations Assistant
AI‑powered workflows that support OPG Investigators with tasks such as financial transaction categorisation and financial analysis, to reduce manual workload and speed up investigations.
1. Summary
1 - Name
OPG Investigations AI Assistant
2 - Description
The Office of the Public Guardian (OPG) plays a vital role in protecting individuals who may lack the mental capacity to make decisions about their health and finances and have a lasting or enduring power of attorney (LPA/EPA) in place. When concerns are raised about the potential misuse of an LPA/EPA, OPG has the power to investigate concerns.
The OPG Investigations AI Assistant is a digital tool that supports investigators when analysing and categorising financial data provided to OPG (inc the content of bank statements provided). It is used to reduce manual effort, improve consistency, and help investigators focus on judgement‑based decisions while maintaining full human oversight.
3 - Website URL
N/A
4 - Contact email
customerservices@publicguardian.gov.uk
Tier 2 - Owner and Responsibility
1.1 - Organisation or department
The Office of the Public Guardian (OPG), an executive agency of the Ministry of Justice (MoJ).
1.2 - Team
OPG Investigations Team
1.3 - Senior responsible owner
Chief Operating Officer, Operations, OPG
1.4 - Third party involvement
Yes
1.4.1 - Third party
Microsoft Limited
1.4.2 - Companies House Number
00984275
1.4.3 - Third party role
Microsoft Limited supported the development of the Investigations Assistant by providing technical expertise and configuring the underlying AI services used in the tool. This included setting up secure Azure‑based components, helping to building the AI workflows such as document processing, financial transaction categorisation and ensuring the system is deployed safely within the OPG environment. All operational control and decision‑making remain with the OPG. Microsoft has no access to OPG data.
1.4.4 - Procurement procedure type
Framework agreement call-off
1.4.5 - Third party data access terms
N/A
Tier 2 - Description and Rationale
2.1 - Detailed description
The Investigations Assistant is an AI‑enabled workflow designed to support the Office of the Public Guardian (OPG) investigations processes when concerns raised with OPG relating to financial misappropriation or abuse trigger an investigation.
The automated assistant supports routine, time‑consuming tasks that investigations and support staff currently carry out manually. The time consuming tasks involve manual assignment of a category to each financial transaction, e.g. utility bill, pension payments, care fees, cash withdrawal etc.
For investigations involving financial categorisation, the tool categorises financial transactions obtained from bank statements provided in the course of financial investigations concerning the actions of an attorney or deputy. Initial testing and evaluation found the time to categorise transactions has almost halved. These time savings will now be further validated in a controlled live pilot.
2.2 - Benefits
Reduces manual effort in the time‑consuming stages of an investigation, especially financial categorisation and analysis.
Improves consistency by generating structured first‑pass outputs that investigators review, amend, and approve as appropriate.
Supports clearer audit trails by recording investigator updates and AI‑generated reasoning in a way that can be monitored and reviewed.
Frees investigators time to focus on higher‑value investigatory actions and more considerative activities, rather than routine administrative tasks. Provides more sustainable ways of working in response to an increasing demand for OPG investigatory work.
2.3 - Previous process
Investigators gathered and reviewed case evidence manually, often stored across multiple files and formats. Financial transactions were categorised manually, with investigators identifying potential concerns through detailed spreadsheet work.
The overall process required substantial manual effort, created variability in outputs, and contributed to an extended timeframe for completion of an investigation.
2.4 - Alternatives considered
Non‑algorithmic alternatives
- Increasing the investigator resource and use of overtime. This would add capacity but would not address the underlying manual nature of the work and increase operating costs.
- Continuous improvment of existing processes and existing templates in isolation. While this would provide some benefits, it would not meaningfully reduce the workload for financial categorisation.
- Improving case‑management systems without AI. Better storage and workflow tools would support organisation, but the core analytical tasks would remain manual.
Algorithmic alternatives
- Rules‑based categorisation. A simple deterministic approach was considered but would struggle with variation in transaction descriptions, require ongoing maintenance, and be less adaptable.
- RPA-based automation. Useful for repetitive actions but unable to interpret financial transactions or produce meaningful content.
- Modern AI‑assisted workflows (chosen approach). This offers the clearest path to reducing manual effort in the stages that dominate investigator workload, while keeping investigators fully in control. It also supports secure operation, consistent outputs, and ongoing monitoring of performance.
Tier 2 - Deployment Context
3.1 - Integration into broader operational process
Concerns OPG receive relating to financial misappropriation or financial abuse involve analysis of individual financial transactions. The AI tool is used to support categorisation of financial transactions on bank statements provided to OPG. Examples of the categories applied to specific transactions are included in section 2.1. The purpose of categorising financial transactions is to support the identification of spending which may not be in the best interests of the person who lacks capacity to manage their own financial affairs.
The tool:
- supports investigators by producing an initial categorisation of financial transactions.
- highlights transactions that may need further attention.
- reduces time spent on routine categorisation so investigators can focus on areas of concern.
The tool provides the following information:
- A categorised view of financial transactions, line by line.
- Flags on transactions where the AI is uncertain or where further review may be required.
How that information is used:
- Investigators complete a full review to check and correct the categorised transactions before completing their analysis.
3.2 - Human review
Investigators review all AI outputs before anything is progressed or shared. They check the categorisation of transactions for accuracy, correct any errors and complete the categorisation for any flagged transactions. Supervisors or senior practitioners may carry out additional spot‑checks. Human review is the formal assurance step: nothing produced by the tool is used without an investigator’s oversight and sign‑off.
3.3 - Frequency and scale of usage
The tool is designed for day‑to‑day use across investigations, and it is expected to be used to support investigations that involve financial categorisation and later analysis. Roll-out training, and use will scale gradually through the pilot phase and expand as staff become familiar with the workflow and ongoing evaluation of the tool/workflows.
For current information about the volume of investigations within the OPG, visit https://www.gov.uk/government/collections/opg-annual-reports
3.4 - Required training
As the tool is made available to investigators they will receive practical introduction to the tool and its workflow. Training covers:
- how to access the files
- how to review, check and correct AI‑generated outputs
- how to use AI responsibly, including limits and appropriate use
- how to escalate issues or report unexpected behaviour
Training is delivered in simple sessions supported by guidance materials.
3.5 - Appeals and review
The tool does not make decisions. All decision making remains with OPG investigators. Members of the public can raise a complaint to dispute OPG decisions relating to Investigations, not the AI output itself. Details of the OPG complaints procedure is published on gov.uk. https://www.gov.uk/government/organisations/office-of-the-public-guardian/about/complaints-procedure
The usual OPG complaints processes remains unchanged. If a concern is raised about how AI‑supported information contributed to a decision, the case can be reviewed through existing supervisory and complaints channels, involving an investigator examining all original evidence and the final judgement.
Investigations that result in an application to the Court of Protection will also involve a review of the evidence by the court.
Tier 2 - Tool Specification
4.1.1 - System architecture
The system follows a modular architecture with four main components:
- AI interface (frontend). A browser‑based interface that investigators use to upload documents, review AI‑generated outputs, and make corrections. It provides secure sign‑in, case navigation, and access to processed materials.
- Backend processing service. A secure API service that handles all business logic, including calling AI models, applying rules, managing workflow steps, and preparing outputs for users.
- Document‑processing pipeline. When files are added to an investigation, an automated pipeline extracts text, converts documents where necessary, classifies file types, and prepares content for analysis. This ensures investigators work from consistent, searchable material.
- Data store. Processed documents, extracted text, and structured outputs are stored in a secure, encrypted database. This allows the interface to retrieve case materials quickly and ensures consistent versioning. Security foundations
The architecture uses encrypted storage, encrypted network connections, identity‑based access, and containerised hosting to ensure that all data remains protected and within controlled environments.
4.1.2 - System-level input
The tool accepts an Excel spreadsheet containing financial transaction data, and case metadata, such as case reference numbers or contextual fields provided at the point of upload.
All input is uploaded via SharePoint‑linked folders.
4.1.3 - System-level output
The tool generates an updated Excel spreadsheet with predicted categories per financial transaction and confidence flags for that categorisation. The output is returned to the investigator for review, editing, and final approval.
The tool also produces logs and structured evaluation of the tool outputs vs human corrections which support performance monitoring and review during the pilot.
4.1.4 - Maintenance
Routine technical maintenance, such as updating container images, reviewing logs, and ensuring the document‑processing pipeline remains stable. Monitoring model performance, using investigator corrections to identify accuracy issues or drift. Model refinement, carried out periodically based on quality insights gathered during real‑world use. Dependency updates, including security patches, library updates, and infrastructure upgrades.
There is currently no automatic model retraining; any retraining occurs through planned, controlled cycles.
4.1.5 - Models
The tool uses a combination of model types:
- gpt-4o (version 2024-11-20, classification) — Azure AI Foundry
- text-embedding-3-large (version 1, embeddings) — Azure AI Foundry
- LightGBM (confidence scoring) — local ML model running in-process, no external calls
Tier 2 - Model Specification
4.2.1. - Model name
A locally hosted LightGBM model running in‑process within the Azure Function App. It does not make external API calls and is not provided as a managed service by a third‑party API provider.
4.2.2 - Model version
The model is an internally trained LightGBM model. There is no externally defined vendor version number.
4.2.3 - Model task
The model is part of the AI workflow and predicts whether an upstream transaction classification made by the AI workflow is likely to be correct or incorrect. It outputs a continuous confidence score used to decide whether a transaction should be accepted or flagged for human review. OPG Investigators can then see which transactions have been flagged by the model and are expected to look at this as part of their review of the AI workflow outputs, as well as reviewing all other categorised transactions.
4.2.4 - Model input
Inputs include:
Reduced‑dimension transaction embeddings, Retrieval similarity statistics, Transaction metadata, Signals derived from the classifier’s structured output, Embeddings of the classifier’s reasoning text.
All inputs are numerical or boolean features assembled into a fixed‑length feature vector during in‑memory processing
4.2.5 - Model output
Confidence score (0.0–1.0): a single numeric score indicating how likely the model believes the upstream classification is correct. Scores below a defined threshold (currently 0.91) result in the transaction being flagged for human review.
4.2.6 - Model architecture
Gradient‑boosted decision tree classifier (LightGBM)
Model type: Binary classifier using LightGBM.
Objective: Negative log‑likelihood, optimised to discriminate between correct and incorrect classifications.
Feature engineering:
- Transaction embeddings reduced using PCA,
- Normalisation via standard scaling,
- Combination of semantic, statistical, metadata, and language‑based features into a 387‑dimensional feature vector.
Calibration: Platt (sigmoid) calibration applied post‑training to improve score calibration.
Weighting: Feature importance is learned automatically by the model; no manual feature weighting rules are applied.
Deployment: runs locally, in‑process, within the Azure Function App. No external calls are made.
This architecture is used strictly for confidence estimation, not for primary transaction classification.
4.2.7 - Model performance
Evaluation focused on error detection and safe auto‑acceptance.
Validation approach: trained and evaluated on a labelled dataset of approximately 54,000 correct and incorrect classification outcomes.
Metrics used: Precision, recall, and accuracy on unflagged (auto‑accepted) transactions.
Operational threshold: Confidence threshold set at 0.91, selected empirically.
Key findings:
- On higher‑quality (“golden”) data, around 23% of transactions are flagged for review, while approximately 98.6% of unflagged transactions are correctly classified.
- The model captures the majority of classification errors while keeping review volumes manageable.
Fairness, privacy, and security:
- All data is stored in line with OPG’s published Records, Retention and Disposal Schedule.
- All processing is in‑memory and confined to the UK Azure region.
- The model does not retrain on live user data.
No third‑party testing has been performed; evaluation is conducted internally by the project team which is a combination of OPG and JusticeAI Unit experts.
4.2.8 - Datasets and their purposes
Internally managed datasets only
Training dataset:Pre‑labelled transaction classifications (correct vs incorrect) derived from MOJ/OPG data based on historical datasets with an additional human review, which is used to train the LightGBM model.
Validation and calibration dataset: Held‑out labelled data used for threshold selection and Platt calibration.
Operational artefacts: the trained LightGBM model file and associated PCA/scaler artefacts stored in Azure Blob Storage and loaded into memory at application startup.
No external or third‑party datasets are used, and no live user data is retained for training or retraining.
2.4.3. Development Data
4.3.1 - Development data description
The development data for the confidence scoring model was created in stages involving both OPG/MOJ and Microsoft.
- OPG/MOJ created an initial dataset of approximately 900 fully synthetic financial transactions generated using Copilot and manually curated to support early experimentation and design validation.
- OPG/MOJ provided Microsoft with an anonymised dataset of approximately 8,000 transactions derived from a random sample of OPG cases, with an emphasis on cases containing larger or unusual transactions. This data originated from real transaction records but was fully anonymised prior to sharing, with all personally identifiable information replaced by synthetic identifiers and transactions aggregated to appear as a single synthetic case.
- A further 14,000 real transactions were anonymised in the same way as the second stage and provided to strengthen accuracy.
- Microsoft augmented the anonymised OPG dataset by scraping publicly accessible transaction examples and using AI to generate additional synthetic transactions. This augmentation produced a larger labelled dataset of approximately 54,000 transaction outcomes, marked as correct or incorrect, which was used to train, validate, and calibrate the LightGBM confidence scoring model. There are no publicly accessible datasets associated with this development data.
Larger datasets were initially used to replicate what occurs during a typical investigation and to improve the model from an operational lens, then a targeted curation of lower frequency unusual categories was completed for targeted improvements. The use of precision, recall and F1 scores throughout evaluation also ensured an approach that sought quality outcomes on all categories. The target curation approach has improved model accuracy and allowed less frequent occuring categories to improve accuracy, however there is still work to be done to improve this.
4.3.2 - Data modality
All datasets used for development consist of tabular financial transaction data. This includes structured fields such as transaction descriptions, payment types, debit and credit amounts, dates, and reference fields. All features used during model development are numerical or categorical representations derived from this tabular data. No image, audio, video, sensor, geospatial, or graph data is used.
4.3.3 - Data quantities
The development data comprises approximately 900 fully synthetic transactions created by OPG/MOJ, approximately 22,000 (two phases of 8,000 and 14,000) anonymised transactions, which was used to build an augmented dataset of approximately 54,000 labelled transactions using a combination of the anonymised OPG data, publicly accessible transaction examples, and AI‑generated synthetic transactions. The final dataset used for model training, validation, and calibration therefore consists of the around 54,000 labelled samples.
4.3.4 - Sensitive attributes
The raw transaction data contained personal and financial information, including names, account details, and identifiers. Prior to sharing with Microsoft, OPG/MOJ fully anonymised the data by replacing all personal names with synthetic identifiers, removing National Insurance numbers, masking bank account numbers, replacing sort codes with synthetic values, and removing case references and source identifiers. Some non‑personal institutional names, invoice numbers, and location reference numbers were retained where they were not considered personal data. The augmented dataset created by Microsoft contains only anonymised or synthetic data. As a result, the development data does not contain direct personal data or protected characteristics, although some transaction patterns may act as indirect proxies.
4.3.5 - Data completeness and representativeness
The combined dataset reflects realistic operational financial transaction patterns, including common and uncommon transaction types, high‑value and unusual transactions, and transaction descriptions that may be incomplete or ambiguous.
Representativeness is addressed by:
- Including both correct and incorrect classifications,
- Training the model to detect uncertainty rather than optimise for perfect coverage,
- Using confidence thresholds to manage residual uncertainty via human review.
No claim is made that the data is statistically representative of the UK population as a whole.
4.3.6 - Data cleaning
Data cleaning and preparation included full anonymisation of historic OPG data prior to sharing, aggregation of transactions across multiple cases into a single synthetic case to remove case‑level identifiability, removal or masking of sensitive identifiers, and standardisation of transaction fields. During model development, additional preprocessing steps included feature engineering, normalisation, and dimensionality reduction. No original unamended historic dataset was retained once anonymisation was completed.
4.3.7 - Data collection
The data provided by OPG/MOJ was collected from cases for the explicit purpose of testing and improving an AI‑assisted financial transaction analysis tool. Its use and sharing were formally approved following anonymisation and Information Assurance review. Microsoft‑generated data was created using publicly accessible transaction examples and AI‑generated synthetic transactions to expand coverage and scale model training. The data was not repurposed from an unrelated context and is directly aligned with the intended operational use of confidence scoring in financial investigations.
4.3.8 - Data access and storage
Prior to sharing, access to the anonymised dataset was restricted to authorised OPG/MOJ staff. The dataset was shared with Microsoft following Information Assurance and OPG Senior Information Risk Owner (SIRO) approval. The augmented and anonymised development dataset is stored and managed by Microsoft within controlled development environments and is used solely for model development, validation, and calibration. Security and privacy measures include formal anonymisation, access controls, purpose limitation, and controlled storage environments. OPG/MOJ retains governance responsibility for the original historic data and the conditions under which it was shared.
4.3.9 - Data sharing agreements
The anonymised dataset was shared by OPG/MOJ with Microsoft under explicit internal approvals for the purpose of developing and testing the AI tool. There is no onward sharing beyond Microsoft, no use of the data outside the agreed development scope, and no sharing under the Digital Economy Act. Use of the data is restricted to the approved testing and development activities and subject to the anonymisation and governance conditions set by OPG/MOJ.
Tier 2 - Operational Data Specification
4.4.1 - Data sources
The tool uses data provided directly through operational use and comprises Excel Spreadsheets uploaded by OPG employees containing digitised financial transaction data from bank statements requested and provided in the course of an investigation. It also includes limited contextual metadata associated with investigation cases, supplied at the point of upload or through case folders. The tool does not ingest data from external public APIs or third party data sources
4.4.2 - Sensitive attributes
The operational data may contain personal and sensitive information, including:
a) Personal data relating to donors and attorneys involved in investigations. b) Financial information, such as transaction histories and account activity.
The tool does not explicitly use protected characteristics as inputs for decision making. All sensitive data is processed only for the purpose of supporting investigations and is subject to the same strict access controls, data minimisation, and secure handling requirements already in place within OPG prior to the tool’s development. Investigators remain responsible for assessing relevance and appropriateness of information used in decision making.
4.4.3 - Data processing methods
Operational data undergoes several preprocessing steps before being analysed:
- Document text from bank statements provided to OPG is extracted and converted into a standard, machine readable format (Excel format).
- Files are classified by type to ensure they are processed appropriately.
- Financial data is normalised to ensure consistent column structures and formats.
- Missing or incomplete values are handled conservatively, with uncertainty flagged for human review.
- Outputs the model cannot confidently interpret are clearly marked for investigator attention. No automated decisions are taken based on preprocessing alone.
4.4.4 - Data access and storage
Operational data is stored to support investigation workflows.
- Uploaded documents, extracted text and structured outputs are stored in secure systems controlled by the organisation.
- Access is restricted to authorised users involved in investigations, using role based access controls and identity based authentication.
- Data is encrypted both in transit and at rest.
- Data retention follows existing organisational retention and records management policies for investigation materials.
- Responsibility for data governance and access rests with the OPG. User interaction data is limited to what is necessary for operational use, monitoring and evaluation.
4.4.5 - Data sharing agreements
Operational data is not shared with external organisations for secondary use. Operational data relevant to an investigation can be shared with the Court of Protection as part of the evidence to support and application to the Court. In serious cases of financial abuse, data may be shared with the police.
Tier 2 - Risks, Mitigations and Impact Assessments
5.1 - Impact assessments
The following impact assessments have been completed or are in place for this tool:
Data Protection Impact Assessment (DPIA) - OPG Investigations – AI assisted Financial Categorisation was completed and approved prior to pilot deployment; updated as the scope moved from trial to pilot. Key findings:
- The tool processes personal and financial data that is already required for investigations in accordance with statutory provisions.
- Risks relating to data protection can be mitigated through strong access controls, UK based processing, data minimisation and mandatory human review of outputs by an investigator.
- No automated decision making is taking place; investigators remain responsible for all decisions.
For the Financial Categorisation Workflow, evaluation was completed prior to pilot deployment involving two steps:
- A comparison of AI outputs to human outputs of cases conducted by the project team. This was conducted in Jan-Feb 2026.
- A comparison of AI outputs to human outputs conducted by the investigators based on current (live) cases, but where the AI comparisons were conducted after the human review and the results were not used in investigation outcomes. This evaluation considered precision, recall and F1 scores by categories and reached a level acceptable for operational deployment as a live pilot, with continued monitoring and evaluation. This was conducted in Feb-March 2026.
5.2 - Risks and mitigations
Risk: Incorrect or misleading AI outputs
- Description: The tool may generate incorrect categorisation or draft text that could mislead if relied on without review.
- Mitigation: All outputs must be reviewed and approved by investigators. Nothing is used without human sign off.
Risk: Over reliance on AI suggestions
- Description: Users may place undue confidence in AI generated outputs.
- Mitigation: Training emphasises that the tool provides assistance only. Investigators retain responsibility for judgement and decision making.
Risk: Bias or unfair outcomes
- Description: Patterns in historical or operational data could lead to biased suggestions.
- Mitigation: The tool does not use protected characteristics as inputs. Human review allows investigators to challenge and correct outputs, and performance is monitored during the pilot.
Risk: Privacy and data protection
- Description: The tool processes sensitive personal and financial information.
- Mitigation: Data is processed within secure systems, access is restricted to authorised staff, and data handling follows existing OPG policies and DPIA controls.
Risk: Security or unauthorised access
- Description: There is a risk of unauthorised access to investigation data.
- Mitigation: Role based access controls, identity based authentication, encryption, and audit logging are in place.
Risk: Reduced trust if AI use is not transparent
- Description: Stakeholders may be concerned about the use of AI in investigations.
- Mitigation: Transparency is addressed through clear internal guidance, investigator accountability, and publication of an ATRS record.