Unofficial. This site is an experimental reformatting of data published by NHS England. It is not endorsed by NHS England. Always check the official Data Uses Register before relying on anything here.

QuantiCode: Admitted Patient Care Data

University of Leeds · Academic

In term In term in the September 2026 edition: the latest version runs to 30 November 2026.

Reference
DARS-NIC-49164-R3G5K
Current version
v3.7
Term of current version
1 December 2023 to 30 November 2026
Start date
1 October 2017
Data controller
Sole Data Controller
Commercial purposes
Yes
Sublicensing
No
Files released to date
1

Why the data was released

Objective for processing

The University of Leeds are running a project called QuantiCode, to develop novel data mining & visualization tools/techniques that will transform people’s ability to analyse event sequence data such as electronic health records, social care data, and retail data. These data often exist in databases that contain millions of records, each comprising hundreds of numerical and categorical variables. This project was funded by the Engineering and Physical Sciences Research Council (EPSRC; March 2016 - February 2020) and the Alan Turing Institute (ATI; May 2019 - December 2021). This DSA is funded by the EPSRC (July 2023 - June 2026), and is being continued with the support of funding from the University of Leeds (ongoing).

QuantiCode project involves seven external organisations (Bradford Teaching Hospitals NHS Foundation Trust, Consumerdata Limited, NHS England [formerly NHS Digital], Sainsbury’s Supermarket Limited, Leeds City Council, Leeds North Clinical Commissioning Group, AQL Limited), who are interested in the research, contribute through meetings, identifying analysis use cases and requirements for the tools/techniques, and receive reports and copies of the tools as outputs of the project. Three of the organisations (NHS England, Sainsbury’s, Leeds City Council) provided datasets, which are neither linked in any way nor used together in any analysis. None of the organisations have access to each other’s data. No record-level data from NHS England will be shared, or be in any way accessible, to third party organisations. Similarly, no aggregated data including small numbers (as defined in the HES Analysis Guide) may be shared with (or be in any way accessible to) third party organisations.

Under UK GDPR article 6(1), the lawful basis for the processing is (e) “public task” (research is a task in the public interest), and under article 9(2) the lawful basis for processing special category data is (j) Archiving, research and statistics (with a basis in law). Like most British universities, the University of Leeds has charitable status because its primary purposes of advancing education and research are deemed to deliver a public benefit. The QuantiCode project is conducting fundamental academic research to acquire new knowledge about data mining & visualization techniques for analysing event sequence data. To help ensure that this research generalises to the scale and complexity of the real world, QuantiCode has been provided with health (from NHS England), social care (Leeds City Council) and retail data extracts (Sainsbury’s). As is clear from the Yielded Benefits (see 5d(iii) below), the project’s applied focus has been on health and the vast majority of the tangible benefits to date are in that domain. Those benefits include identifying an issue that affects the NHS’s Payment by Results system for hospitals, showing how the project's tools can be used to identify the source of data quality problems so hospitals are able to improve the future quality of their own data, and revealing previously unknown issues in the primary and secondary care data extracts that researchers were using to study melanoma/diabetes survival and acute myocardial infarction (heart attack) outcomes/hospitalization rates, respectively.

QuantiCode is multi-disciplinary, with academic researchers from computer science, maths, ethics, geography, and health. The NHS England data is only used in the part of the project that is developing novel tools/techniques for investigating data quality. Data quality is an issue that is particularly important for NHS England – it impacts directly on areas such as running the NHS (invoice validation, etc.), as well as some indirect benefits (research and analysis to develop policy/clinical guidelines). Specifically, QuantiCode’s objectives are to develop tools/techniques to: (1) comprehensively investigate data quality (number of missing values, percentiles, outliers, etc.), (2) make detailed investigations of patterns of missing data and, where possible, to identify factors associated with their source, and (3) investigate factors that affect diagnostic persistence (the most recent NHS England report shows that this affects 42% of HES data episodes for patients with conditions that do not remit https://digital.nhs.uk/data-and-information/data-tools-and-services/data-services/data-quality). By collaborating with the QuantiCode project, it is expected that there will be benefits for NHS England and to third parties who use NHS England data (e.g., see 5d (iii) Yielded Benefits, below).

To perform this research into data quality, QuantiCode requested pseudonymised admitted patient care (APC) data from hospital episode statistics (HES). A single year of HES Admitted Patient Care data was being requested in order to fulfil these objectives. The factors in (2) and (3) are currently unknown and may involve any variable in a dataset, which is why all variables (except those deemed sensitive or identifiable) were requested.

To address the GDPR principle of data minimisation, the University of Leeds has only requested a single year of pseudonymised data, this the minimum amount of data needed in order to still be able to achieve the aims stated within this agreement. Data concerning individuals admitted to any English NHS hospital in a whole year is required for the purposes listed in (1), (2) and (3) to ensure that the tools/techniques scale to the volume and complexity of real NHS data. Additionally, while some data issues are common (3) others such of those in (2) can be rare, so potentially missing if only a part-year dataset, a subset of APC variables or limited geography was requested.

The University of Leeds is the sole controller who also processes the data for the purposes described within this Agreement

Processing activities

The University of Leeds do not flow any data to NHS England. Under a previous iteration of this agreement, NHS England flowed pseudonymised HES APC (2015/16) to the University of Leeds. This data was sent via the Secure Electronic File Transfer System (SEFT). There are no subsequent flows of data.

The data will only be accessed by substantive employees of the University who are contributing to the project, all of whom have been trained in data protection and confidentiality. The data is in a category that the University classifies as IRC-Confidential. The data will be stored on University computer systems that are NHS DSPT and ISO 27001 accredited and accessed in accordance with ISO 27001 and University policies, including the information protection policy.

The patient data from NHS England will not be linked with any other dataset, including those received from Quanticode Partners.

There will be no attempt to re-identify individuals from the pseudonymised data.

The Quanticode project is developing tools and methodologies for investigating data quality across a number of datasets. The aim is to test these tools and methodologies across a variety of datasets, including health data.

The QuantiCode project will process datasets provided by organisations involved in the project (including NHS England), and the methodologies/tools will be developed to apply across datasets as much as possible. One of the aims of this is to ensure that techniques are developed which are generically applicable, rather than each sector needing to develop their own tools. The datasets from the different organisations are never linked and never used together in any analysis. None of the other organisations will have access to the NHS England data.

More specifically, the NHS England dataset will be used to help design and test new data analysis tools (called QCprofiling, ACE and RFviz), which allow users to investigate data quality in tabular data such as electronic health records. The following types of computation will be performed: (a) calculate descriptive statistics (number of missing values, percentiles, outliers, etc.), (b) calculate sets and set intersections, and (c) data mining (using well-established methods such as random forests, gradient boosting, entropy and information gain). To put these in the context of achieving QuantiCode’s purpose, (a) is central to allowing users to comprehensively investigate data quality (as enabled by the QCprofiling software), and (b) is central to allowing users to go further with detailed investigations of patterns of missing data (the ACE software). Data mining techniques such as entropy and information gain (c) are essential for pinpointing the origin of patterns of missing data in ACE, and techniques such as random forests and gradient boosting are essential for investigating diagnostic persistence (the RFviz software).

Output from the computations will be visualized using a wide range of techniques, as appropriate to the type and scale of data. For comprehensive data quality investigations and investigating diagnostic persistence (QCprofiling and RFviz output, respectively) those techniques span bar charts, histograms, scatterplots, box plots and character maps. ACE uses a smaller suite of techniques (bar charts, histograms and heatmaps). Real-world data is essential both for designing the above tools and testing them. That testing is important for informing design choices (e.g., what are the speed/accuracy trade-offs of data mining methods such as random forests vs. gradient boosting) and testing the scalability of the tools to reduce bottlenecks and improve performance.

The NHS England dataset will only be used where necessary, and will only be used for the purposes described in this agreement. Live data will not be used for early stages of development/testing when it would be more appropriate to use test data.

The data from NHS England may only be processed in order to produce the outputs detailed below (in the Specific Outputs Expected section).

Microsoft Limited provide Cloud Services for the University of Leeds as part of the LASER platform (Leeds Analytic Secure Environment for Research) and are therefore listed as a processor. They supply support to the system, but do not access data. Therefore, any access to the data held under this agreement would be considered a breach of the agreement. This includes granting of access to the database[s] containing the data.

The Data will be accessed by authorised personnel via remote access.

The Controller(s) must confirm and provide evidence upon audit by NHS England that access via any remote device complies with the data security obligations within this DSA and the Data Sharing Framework Contract.

For remote access:

- Remote access will only be from secure locations situated within the territory of use (as further restricted elsewhere within the DSA if so done) stated within this DSA;

- Access controls granting users the minimum level of access required are in place;

- Remote access is only via secure connections (e.g., VPNs or secure protocols) to protect data;

- Multifactor authentication (MFA) is required for remote access;

- Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for the remote access;

- All remote access is undertaken within the scope of the organisation’s DSPT (or other security arrangements as per this DSA) and complies with the organisation’s remote access policy.

The above applies in addition to any condition set out elsewhere within the DSA (e.g. who may carry out processing, and for what purpose).

Expected output

The intended outputs relating to health were:

1) Data analysis tool Version 1: This output will be a visual analytics tool, which allows users to gain an overview of missing data patterns and investigate data integrity in health datasets.

2) Research report 1: This report will describe the application of the tool to health data, and the benefits that the tool provides. The report will be submitted to a high-impact outlet such as the Journal of the American Medical Informatics Association (the pre-eminent journal for research into methods for analysing health data).

3) Data analysis tool Version 2: This version of the visual analytics tool will allow users to investigate bias caused by data quality issues in health datasets.

4) Research report 2: This report will describe the application of the tool for bias investigations, and the benefits that the tool provides. The report will be submitted to the Journal of the American Medical Informatics Association.

Outputs 1 & 3 will not contain any data – a user will load their dataset into the tool to analyse ‘missingness’. Outputs 2 & 4 will contain only aggregate level data with small numbers suppressed in line with HES analysis guide. That data will be shown in figures that illustrate the usage of the tool.

The ultimate beneficiaries of this work will be the general public. For example, Integrated Care Boards (ICBs) and local authorities analyse data to generate business intelligence for operations and investment, with the aim of providing us all with improved and more cost-effective services. Businesses similarly require business intelligence for operations and investment, which translates to jobs and other economic benefits. To bridge the gap between these indirect benefits from the project and popular interest in big data, the University of Leeds will conduct a range of public engagement activities which include live demonstrations at the annual Leeds Festival of Science, a short film, an on-line tutorial about the ethics surrounding data analytics, and publishing articles in the popular scientific press.

The following outputs have been produced as of June 2023:

Note: all outputs either did not contain any NHS England data or were aggregated with small numbers supressed in line with HES Analysis Guidance.

Software (see Outputs 1 & 3, above)

• A data analysis tool called ACE. Version 1 was released to the QuantiCode project partners on 20/7/18. Version 2 was released to the QuantiCode project partners on 6/11/19. This version of ACE contained substantial new functionality to provide set visualization functionality that is generic, so that sets of both missing and present data can be analysed. Some parts of the ACE were re-engineered to speed up processing and make it more scalable. As part of that, ACE was tested on a 64 GB desktop PC with datasets that contained up to 21 million records, 417 fields (i.e., sets) and 141,000 unique combinations of field (set intersections).

• A Python/Pandas library (“QCprofiling”) for computing a suite of data quality checks and visualizing the output in matrices of miniature visualizations has been developed. The checks include data type, missing values, unique values, value lengths, character pattern, percentiles, outliers and example values.

• The ACE software has been made available. DOI (https://doi.org/10.5518/1150). The ACE software will be made widely available for free non-commercial use, initially leveraging the existing Alan Turing Institute Health Programme (includes researchers from 13 partner universities) and the HDRUK Health Data Research Hubs (includes NHS-accredited cloud-based IT platforms). This involves publicity (Q4 2020 onward) and follow-up support of interested researchers, providing additional evaluation material that is needed for peer reviewed academic papers (see below)

• A Python version of the software has been released on the well-known PyPI repository (https://pypi.org/project/setvis/).

Peer reviewed academic papers (*indicates outputs that concern “Non-health benefits” outlined in 5d; all other outputs are directly related to the NHS England data request)

• Research reports (see Outputs 2 & 4, above): These “reports” are in fact academic papers. The first journal manuscript was submitted to the IEEE Transactions on Visualization and Computer Graphics on 24/7/2018, and unfortunately not accepted. The second was submitted to the Journal of the American Medical Informatics Association on 11/6/2020 and also not accepted. Revisions and resubmissions are planned, and in both cases further work is required (see 5a. Objective for processing, and 5b. Processing activities).

• *Adnan, M., & Ruddle, R.A. (2018). A set-based visual analytics approach to analyze retail data. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA).

• Adnan, M., Nguyen, P. H., Ruddle, R. A. & Turkay, C. (2019). Visual analytics of event data using multiple mining methods. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA).

• Ruddle, R. A., & Hall, M. S. (2019). Using miniature visualizations of descriptive statistics to investigate the quality of electronic health records. Proceedings of the 12th International Joint Conference on Biomedical Engineering Systems and Technologies - Volume 5 HEALTHINF: HEALTHINF, ISBN 978-989-758-353-7, pages 230-238.

• Ruddle, R. A., Adnan, M., & Hall, M. (2022). Using set visualisation to find and explain patterns of missing values: a case study with NHS hospital episode statistics data. BMJ Open, 12(11), e064887. DOI http://dx.doi.org/10.1136/bmjopen-2022-064887.

Reports and Presentations (*indicates outputs that concern “Non-health benefits”)

• Feedback to NHS Digital, correcting an error in the APC Data Dictionary (ref: NIC-309404-B7V2J; 20/6/2019)

• “QuantiCode: Benefits to Health and/or Social Care” report for NHS Digital (17/7/2019)

• *Understanding customers’ missions from the products they purchase. Technical report for Sainsbury’s. June 2019.

• *Briefing to Leeds City Council (LCC) at Adult Social Care Information Management and Technology (IM&T) Team Meeting (8/11/16)

• “Exploring missingness patterns in Admitted Patient Care (APC) data using ACE” to NHS Digital Head of Information Utilisation and colleagues (16/10/2017)

• Demo of ACE to NHS Digital and NHS England staff, with discussion of application to issues such as diagnostic persistence (14/11/2017)

• Demo of ACE to NHS Digital Data Quality and Casemix teams (1/12/2017)

• “Visualizing data profiles and analysis pipelines”. Presentation at Visualization for Data Science and AI. Alan Turing Institute, London, 13/9/2019.

• *Discussion of strategies for investigating data quality for LCC’s adult social care business intelligence dashboard (29/10/19)

• Ruddle, R. A., Hama, L., Wochner, P., & Strickson, O. T. SetVis: Visualizing Large Numbers of Sets and Intersections. Submitted to EuroVA. Includes an exploratory investigation of diagnostic persistence for autism.

• AI UK presentation “Does AI help visualization or visualization help AI?” (23rd/24th Mar 2021 (https://www.turing.ac.uk/ai-uk).

• ATI REG Tech Talk to provide overview of the setvis software and hands-on demo/tutorial (14/12/21).

• “Visualizing the Quality of Data” film on YouTube (https://tinyurl.com/VizDataQuality) has now had 2900 views.

• HES missing data case study (see BMJ Open paper above) used as a complex, real-world example for approximately 1000 MSc Data Science students to date.

Keynote conference presentation

• “Visualizing health data – from fundamental research to successful applications”. 13th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC), Malta, 26/2/2020

Public engagement

• Exhibition and hand-on demo of prototype of ACE at the University of Leeds Be Curious research festival (25/3/17). 847 adults and children attended.

• Release a film about “Visualizing the Quality of Data” on YouTube https://tinyurl.com/VizDataQuality (20/9/2019)

Other

• Training workshops for ACE were run in July and December 2018

The following new outputs are expected:

Software

• Release Python data profiling/data quality library on public PyPI repository (Q2 2024).

• Release workflow for systematically investigating structures of missing data, implemented in Jupyter Notebook/Python (Q3 2024).

Peer reviewed academic papers

• Paper(s) describing the investigation of diagnostic persistence in autism, dementia, diabetes, IHD, Parkingsons and schizophrenia. For submission to a high-impact health-focussed journal such as BMJ Open (Q2 2024).

• Paper(s) describing the scalability of visualizations for explainable AI as applied to diagnostic persistence scenarios. For submission to high-impact outlets such as the IEEE Transactions on Visualization and Computer Graphics (Q4 2024; Q4 2025; Q4 2026)

Reports & presentations

• Publish a freely available “Practitioner’s Guide to Visual Data Profiling”. This Guide will describe best practice for the use of visualization to characterise data and investigate data quality, across a range of application domains.

• Develop and run a continuing professional development (CPD) course on data profiling, which will be made generally available.

• Six-monthly reports/presentations to NHS England Data Quality team (similar to previous presentations; see above) to stimulate adoption.

• Submit presentation about diagnostic persistence for AIUK 2024.

• Public engagement via the annual Be Curious research festival.

Expected measurable benefits

It is expected that there will be benefit to health from the outputs provided to NHS England. The benefits are hoped to occur as benefits to the health system, either directly (through improvements to data collections) or indirectly (better quality data used in analysis/research).

The following headline benefits are hoped to arise from NHS England's usage of the tool:

1) Allow NHS England to conduct integrity checking to identify occurrences of poor data quality in routinely collected data, for feedback to data providers.

2) Allow NHS England to cross-reference variations in data quality issues with external influences (e.g. change in policy, change in priorities, change in resource, change in service provision or structure, coding improvement initiatives).

3) Allow NHS England to conduct bias checking to identify occurrences of poor data quality in routinely collected data, for feedback to data providers.

4) Support NHS England in the development of new business rules for automated data profiling and feedback to data providers.

5) Improve NHS England’s understandings of biases across time, geography and activity; NHS England’s focus has necessarily been on data quality by provider (in order to report back to providers so that data quality issues can be addressed ), with less emphasis given to time, geography and activity.

6) Allow NHS England to quantify the impact of data quality issues - e.g., whether the degree of missingness for certain conditions supports linkage match rates.

7) Improve the quality of the data used by NHS England when conducting analysis.

8) Improve the quality of the data provided by NHS England to third parties when conducting research/analysis.

NHS England manages many of the nation’s critical health and care data assets. It collects data from a range of care providers and provides secure and controlled access to those data by legally authorised bodies. Better use of health and care data will help those involved to:

- manage the system more effectively;

- commission better services;

- understand health and care trends in more detail;

- develop new treatments; and

- monitor the safety and effectiveness of care providers.

Understanding the quality of data is essential in deciding whether it is fit for these uses. The benefits above will help NHS England develop appropriate methods to monitor, challenge and highlight data quality issues to various audiences, giving the them the ability to correct the data if submitting or adjust their findings accordingly if just analysing.

NHS England also has a statutory responsibility, enacted in the Health and Social Care Act 2012 (section 266), to assess the quality of the data it receives against nationally published standards and to publish the results of those assessments. These benefits are likely to lead to new measurements in the report covering this statutory responsibility.

The tools and methodologies developed through the QuantiCode project may also benefit other organisations in other sectors, but the data from NHS England is only disseminated on the basis of there being a benefit to health.

The following benefits are expected to arise during the time-period of this data sharing agreement extension:

1) Allow NHS England to identify factors that affect a data quality issue called “diagnostic persistence”, which is widespread but poorly understood.

2) Support NHS England in identifying new business rules for automated data profiling and feedback to data providers.

3) Improve the quality of the data used by NHS England when conducting analysis.

4) Improve the robustness of data quality checks carried out by third parties when conducting research/analysis with NHS England data.

To put the importance of (1) in perspective, NHS England has asked QuantiCode to investigate six conditions (autism, dementia, diabetes, ischemic heart disease, Parkinson’s and schizophrenia). Initial investigation has shown that, in the 21 million episode APC dataset, 1 million episodes from 300,000 patients exhibit that data quality issue (diagnoses not persisting when they should for one or more of the conditions). By identifying factors that affect this issue, the aim is to provide insights that allow NHS England to design interventions (coding policy, validation, feedback, etc.) that improve diagnostic persistence and, therefore, directly improve data quality. As a consequence, the hospitals would have more accurate information for treating their own patients, and health analysts (3) and researchers (4) would have more robust data for epidemiological studies. In a production environment, data quality checks need to be automated. QuantiCode has already yielded benefits by showing how effective interactive visualization is for discovering previously unknown issues (e.g., in diagnosis code, operation and date fields; see 5d(iii)). Understanding the patterns exhibited by those issues, makes it possible for NHS England to design new business rules to automate the checking of future data (2), and exploit QuantiCode’s visualization techniques for further, continuous quality improvement.

Note: The external organisations (NHS England and six others; see (5a) for details) that are involved in this project received copies of the reports and software tools that the project produced up until 29/2/2020 (the end of the project’s original EPSRC funding). Additionally, and as described below in (iii) Yielded Benefits, Leeds City Council were shown how ACE’s set visualization techniques could be used to investigate patterns of missing data in an anonymized adult social care service provision dataset. Sainsbury’s were shown how frequent itemset mining techniques could be combined with ACE’s set visualization to analyse customer transactions.

University of Leeds have developed a promising analysis workflow for the analysis of lost autism diagnoses. The current aim is to generalise that workflow to account for the vast majority (90+ %) of lost diagnoses across all of the above conditions that do not remit. That is hoped to yield benefits for the understanding of diagnostic persistence and, more generally, for combining visualization with machine learning for other types of complex health data analysis.

Benefits reported so far

To date there have been three main types of benefit. In the following, the eight original Expected Measurable Benefits (see 5d(ii)) are referenced as [1], [2], etc.

Benefits to NHS England and hospitals themselves: The ACE tool revealed unexpected gaps in sets of fields for some Admitted Patient Care (APC) data records, and pinpointed the source of 90+% of the gaps as being a particular admission method at a single hospital [1]. Some of the gaps were in the diagnosis code fields (DIAG_nn) and others were in the operation fields (OPERTN_nn). The latter are of particular concern to NHS England because they affect the process that the Casemix Team use to calculate Healthcare Resource Groups (HRGs), which are one of the building blocks of the NHS’s Payment by Results system for hospitals [6]. These findings from ACE allow NHS England to provide feedback to the provider to rectify the problem, via the existing HES data quality lifecycle. ACE has subsequently revealed similar gaps in the date fields for operations (MYOPDATE_nn) and inconsistencies between the operation fields and their corresponding date fields. The speed (a few minutes) with which users were able to interactively find these unexpected missing data issues with ACE, contrasts with the years that the issues have existed yet gone undetected. It would be straightforward to implement new business rules to automatically check for these gaps [4], to complement ACE’s power in exploratory analysis for revealing new issues. The results would be improved data quality for NHS England’s own analysis [7], and in data and/or data quality notes provided to third parties [8].

Benefits to third party users of electronic health records: QuantiCode has concentrated its effort on assisting researchers working on two other projects – one using APC data (DARS-NIC-17649-G0X4B) to study long-term outcomes and hospitalization rates for survivors of acute myocardial infarction (heart attack), and the other using ResearchOne primary care data to study the survival from melanoma of patients with type 2 diabetes.

ACE revealed two main issues with the APC data. One concerned missing admission and episode start dates. In principle they can be estimated from each other if only one is missing, but ACE showed that that was hardly ever the case (99.8% of records that were missing one of those dates was also missing the other). However, ACE also let the researchers localize the origin to maternity records, and therefore make an informed decision to discard the records because the project was conducting heart attack research [8]. The other issue concerned survival time, which was missing in 12 million records and led the researchers to comment that they had asked NHS England to derive the survival time for all patients based on the study census date, regardless of their mortality status, but clearly that had not been done. This finding made the researchers realise that they had to estimate survival time for those patients who did not die prior to the census date, losing precision and precluding analyses of short-term survival patterns within 30 days of admission. In future, ACE would allow the researchers to check their data on receipt, so that errors can be rectified (NHS England’s time-limit is one month from receipt) [8].

Complementing ACE, the QCprofiling software revealed a number of longitudinal data quality issues [5] across nine years of APC data extracts (150 million records in total; 116 different fields). Some were the result of policy change (e.g., all values of Clinical Commissioning Group fields missing in the early extracts but not the latest, and longitudinal differences in coding depth for the operation fields) [2], others from improvements to business rules (the primary diagnosis was sometimes missing in the early extracts but never from the latest ones), some were prevalent across all of the extracts (e.g., widespread use of punctuation characters in diagnosis codes, which complicates data cleaning), one affected data linkage [6] (the use of two different (16 vs. 32 character) pseudonymized patient identifiers in one of the extracts).

The QCprofiling software also revealed data quality issues with the primary care data (90 million records; 14 tables; 41 different fields). These included validation error text appearing in a birth year field, clinical codes being padded out with ‘.’ characters (i.e., inconsistent coding precision between different data tables), and the distribution of blood pressure measurement dates leading the researchers to realise that they had not been provided with the historical data that they had expected [8]. All of these issues were new to the researchers, despite the fact that they had had the data for several months. As with the APC data, the QCprofiling software was a game-changer for investigating data quality because it allowed the researchers to see patterns that spanned a large number of records/tables/fields at the glance of an eye.

Non-health benefits: ACE was also evaluated and used for more exploratory analysis with data from two of QuantiCode’s other external partners (Leeds City Council (LCC) and Sainsbury’s), to help generalise the design of ACE. LCC were shown how ACE’s set visualization techniques could be used to investigate patterns of missing data in an anonymized adult social care service provision dataset. The main benefit lies in making explicit the patterns for discussion by analysts and stakeholders prior to use of the data for other purposes such as modelling future service demand. In some cases explanations could be offered for unexpected patterns (e.g., tacit knowledge about historical changes in which data were recorded), and in others further investigation would be required.

Sainsbury’s were shown how frequent itemset mining techniques could be combined with ACE’s set visualization to analyse customer transactions (e.g., What items are bought together? How to highlight differences and similarities between stores, or over time?). This was exploratory research, indicating a possible approach that would allow customer transactions to be analysed at a much finer level of detail than is possible with Sainsbury’s current methods.

Important note: As stipulated in the project’s data agreements, the datasets from the different organisations (NHS England, LCC, Sainsbury’s, ResearchOne) are never linked and never used together in any analysis. None of the other organisations have access to the NHS England data.

Some health conditions (e.g., autism, dementia, diabetes, IHD, Parkinson's and schizophrenia) do not remit, so the relevant ICD code should always appear (“persist”) as one of the diagnoses in APC data, even if the patient is being treated for something else (e.g., a broken arm). However, for years the NHS has been concerned that diagnostic persistence is poor, with up to 50% of episodes missing those diagnoses for non-remitting conditions. NHS England asked University of Leeds to investigate this because the reasons are not known, and the investigations have proved challenging. Machine learning models based on decision trees, random forests and gradient boosting gave little in the way of insights because they could not handle the extreme complexity of health records data. However, a promising analysis workflow has been developed that combines set visualization techniques with on-the-fly computation, and this has been used to identify two previously unknown patterns that account for the 60% of lost autism diagnoses in this data agreement’s Admitted Patient Care dataset.

Datasets on the current version

Legal basis for provision: Health and Social Care Act 2012 – s261(2)(a)

Datasets approved under DARS-NIC-49164-R3G5K-v3.7
DatasetType of dataSensitivity FrequencyConfidential data
Hospital Episode Statistics Admitted Patient Care (HES APC) Anonymised - ICO Code Compliant Non-Sensitive One-Off Does not include the flow of confidential data

Files released

Files released counts only files released externally by DARS. Access granted in NHS England's own systems, such as its Secure Data Environment, is not included.

Patient opt-outs were not applied to the one file released under this agreement. About opt-outs

No files recorded as released under the current version. 1 was released under earlier versions, shown in the version history.

Version history

The register lists each renewal of this agreement as a separate row. This site has 4 versions.

DARS-NIC-49164-R3G5K-v3.7 1 December 2023 to 30 November 2026
Title
QuantiCode: Admitted Patient Care Data
Commercial
Yes
Sublicensing
No
Datasets
1
Files released
0

Datasets: Hospital Episode Statistics Admitted Patient Care (HES APC)

What changed from DARS-NIC-49164-R3G5K-v2.1

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-49164-R3G5K-v2.1
FieldWasBecame
Start date2021-05-192023-12-01
End date2023-09-302026-11-30
Hospital Episode Statistics Admitted Patient Care (HES APC): legal basisHealth and Social Care Act 2012 - s261 - 'Other dissemination of information'Health and Social Care Act 2012 – s261(2)(a)

Objective for processing

The University of Leeds are running a project called QuantiCode, to develop [11 words unchanged] to analyse event sequence data such as electronic health records, social care data data, and retail data. These data often exist in databases that contain millions [11 words unchanged] project was funded by the Engineering and Physical Sciences Research Council (EPSRC; Mar March 2016 – Feb 2020), - February 2020) and the Alan Turing Institute (ATI; May 2019 - December 2021). This DSA is funded by the EPSRC (July 2023 - June 2026), and is being continued with the support of funding from the Alan Turing Institute (ATI; May 2019 – Dec 2021) and the University of Leeds (ongoing). QuantiCode project involves seven external organisations (Bradford Teaching Hospitals NHS Foundation Trust, Consumerdata Limited, NHS Digital, England [formerly NHS Digital], Sainsbury’s Supermarket Limited, Leeds City Council, Leeds North Clinical Commissioning Group, AQL [25 words unchanged] the tools as outputs of the project. Three of the organisations (NHS Digital, England, Sainsbury’s, Leeds City Council) provided datasets, which are neither linked in any [10 words unchanged] organisations have access to each other’s data. No record-level data from NHS Digital England will be shared, or be in any way accessible, to third party [17 words unchanged] shared with (or be in any way accessible to) third party organisations. Under UK GDPR article 6(1), the lawful basis for the processing is (e) “public task” (research is a task in the public interest), and under [83 words unchanged] of the real world, QuantiCode has been provided with health (from NHS Digital), England), social care (Leeds City Council) and retail data extracts (Sainsbury’s). As is [93 words unchanged] study melanoma/diabetes survival and acute myocardial infarction (heart attack) outcomes/hospitalization rates, respectively. QuantiCode is multi-disciplinary, with academic researchers from computer science, maths, ethics, geography, and health. The NHS Digital England data is only used in the part of the project that is [6 words unchanged] quality. Data quality is an issue that is particularly important for NHS Digital England – it impacts directly on areas such as running the NHS (invoice [54 words unchanged] and (3) investigate factors that affect diagnostic persistence (the most recent NHS Digital England report shows that this affects 42% of HES data episodes for patients [11 words unchanged] QuantiCode project, it is expected that there will be benefits for NHS Digital England and to third parties who use NHS Digital England data (e.g., see 5d (iii) Yielded Benefits, below). [1 paragraph unchanged] To address the GDPR principle of data minimisation minimisation, the University of Leeds has only requested a single year of pseudonymised [84 words unchanged] part-year dataset, a subset of APC variables or limited geography was requested. The University of Leeds is the sole data controller who also processes the data for the purposes described within this Agreement

Processing activities

The University of Leeds do not flow any data to NHS Digital. England. Under a previous iteration of this agreement agreement, NHS Digital England flowed pseudonymised HES APC (2015/16) to the University of Leeds. This data [5 words unchanged] Electronic File Transfer System (SEFT). There are no subsequent flows of data. [1 paragraph unchanged] The patient data from NHS Digital England will not be linked with any other dataset, including those received from Quanticode Partners. [2 paragraphs unchanged] The QuantiCode project will process datasets provided by organisations involved in the project (including NHS Digital), England), and the methodologies/tools will be developed to apply across datasets as much [44 words unchanged] analysis. None of the other organisations will have access to the NHS Digital England data. More specifically, the NHS Digital England dataset will be used to help design and test new data analysis [50 words unchanged] (using well-established methods such as random forests, gradient boosting, entropy and information gain ). gain). To put these in the context of achieving QuantiCode’s purpose, (a) is central to allowing users to comprehensively investigate data quality (as enabled by our the QCprofiling software), and (b) is central to allowing users to go further [40 words unchanged] and gradient boosting are essential for investigating diagnostic persistence (the RFviz software). [1 paragraph unchanged] The NHS Digital England dataset will only be used where necessary, and will only be used [16 words unchanged] of development/testing when it would be more appropriate to use test data. The data from NHS Digital England may only be processed in order to produce the outputs detailed below (in the Specific Outputs Expected section). Microsoft Limited provide Cloud Services for the University of Leeds as part of the LASER platform (Leeds Analytic Secure Environment for Research) and are therefore listed as a data processor. They supply support to the system, but do not access data. [17 words unchanged] agreement. This includes granting of access to the database[s] containing the data. The Data will be accessed by authorised personnel via remote access. The Controller(s) must confirm and provide evidence upon audit by NHS England that access via any remote device complies with the data security obligations within this DSA and the Data Sharing Framework Contract. For remote access: - Remote access will only be from secure locations situated within the territory of use (as further restricted elsewhere within the DSA if so done) stated within this DSA; - Access controls granting users the minimum level of access required are in place; - Remote access is only via secure connections (e.g., VPNs or secure protocols) to protect data; - Multifactor authentication (MFA) is required for remote access; - Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for the remote access; - All remote access is undertaken within the scope of the organisation’s DSPT (or other security arrangements as per this DSA) and complies with the organisation’s remote access policy. The above applies in addition to any condition set out elsewhere within the DSA (e.g. who may carry out processing, and for what purpose).

Expected output

[1 paragraph unchanged] 1) Data analysis tool Version 1:This 1: This output will be a visual analytics tool, which allows users to gain an overview of missing data patterns and investigate data integrity in health datasets. [1 paragraph unchanged] 3) Data analysis tool Version 2:. 2: This version of the visual analytics tool will allow users to investigate bias caused by data quality issues in health datasets. [2 paragraphs unchanged] The ultimate beneficiaries of this work will be the general public. For example, CCGs Integrated Care Boards (ICBs) and local authorities analyse data to generate business intelligence for operations and [79 words unchanged] ethics surrounding data analytics, and publishing articles in the popular scientific press. As of September 2020 the following outputs have been produced, or are expected to be produced: The following outputs have been produced as of June 2023: Note: all outputs either did not contain any NHS Digital England data or were aggregated with small numbers supressed in line with HES Analysis Guidance. [3 paragraphs unchanged] Peer reviewed academic papers (*indicates outputs that concern “Non-health benefits” outlined in 5d; all other outputs are directly related to the NHS Digital data request) • The ACE software has been made available. DOI (https://doi.org/10.5518/1150). The ACE software will be made widely available for free non-commercial use, initially leveraging the existing Alan Turing Institute Health Programme (includes researchers from 13 partner universities) and the HDRUK Health Data Research Hubs (includes NHS-accredited cloud-based IT platforms). This involves publicity (Q4 2020 onward) and follow-up support of interested researchers, providing additional evaluation material that is needed for peer reviewed academic papers (see below) • A Python version of the software has been released on the well-known PyPI repository (https://pypi.org/project/setvis/). Peer reviewed academic papers (*indicates outputs that concern “Non-health benefits” outlined in 5d; all other outputs are directly related to the NHS England data request) [4 paragraphs unchanged] Reports (*indicates outputs that concern “Non-health benefits”) • Ruddle, R. A., Adnan, M., & Hall, M. (2022). Using set visualisation to find and explain patterns of missing values: a case study with NHS hospital episode statistics data. BMJ Open, 12(11), e064887. DOI http://dx.doi.org/10.1136/bmjopen-2022-064887. Reports and Presentations (*indicates outputs that concern “Non-health benefits”) [3 paragraphs unchanged] Presentations (*indicates outputs that concern “Non-health benefits”) [6 paragraphs unchanged] • Ruddle, R. A., Hama, L., Wochner, P., & Strickson, O. T. SetVis: Visualizing Large Numbers of Sets and Intersections. Submitted to EuroVA. Includes an exploratory investigation of diagnostic persistence for autism. • AI UK presentation “Does AI help visualization or visualization help AI?” (23rd/24th Mar 2021 (https://www.turing.ac.uk/ai-uk). • ATI REG Tech Talk to provide overview of the setvis software and hands-on demo/tutorial (14/12/21). • “Visualizing the Quality of Data” film on YouTube (https://tinyurl.com/VizDataQuality) has now had 2900 views. • HES missing data case study (see BMJ Open paper above) used as a complex, real-world example for approximately 1000 MSc Data Science students to date. [8 paragraphs unchanged] Note: all outputs all outputs either did not contain any NHS Digital data or will be aggregated with small numbers supressed in line with HES Analysis Guidance. [1 paragraph unchanged] • The ACE software will be made widely available for free non-commercial use, initially leveraging the existing Alan Turing Institute Health Programme (includes researchers from 13 partner universities) and the HDRUK Health Data Research Hubs (includes NHS-accredited cloud-based IT platforms). This involves publicity (Q4 2020 onward) and follow-up support of interested researchers, providing additional evaluation material that is needed for peer reviewed academic papers (see below) • Release Python data profiling/data quality library on public PyPI repository (Q2 2024). • Release version 1 of the RFviz software (Q4 2021) for free use by the QuantiCode project partners and free non-commercial use by others. This software will allow a data-driven approach to be taken for the investigation of diagnostic persistence and similar data quality issues. • Release workflow for systematically investigating structures of missing data, implemented in Jupyter Notebook/Python (Q3 2024). • New workflow software, to make the QCprofiling library more accessible to users for comprehensively investigating data quality and developing new business rules for automated data profiling (Q2 2022). [1 paragraph unchanged] • Technical paper Paper(s) describing the ACE tool. investigation of diagnostic persistence in autism, dementia, diabetes, IHD, Parkingsons and schizophrenia. For submission to a high-impact outlet health-focussed journal such as Information Visualization (Q3 2021). BMJ Open (Q2 2024). • Paper describing the application of ACE to investigating electronic health records. For submission to a high-impact outlet such as the Journal of the American Medical Informatics Association (Q1 2022). • Paper(s) describing the scalability of visualizations for explainable AI as applied to diagnostic persistence scenarios. For submission to high-impact outlets such as the IEEE Transactions on Visualization and Computer Graphics (Q4 2024; Q4 2025; Q4 2026) • Paper describing the design and evaluation of RFViz to the investigation of diagnostic persistence. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q2 2022). • Paper describing the application of the new workflow software. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q4 2022). [1 paragraph unchanged] • Six-monthly reports/presentations to NHS Digital Data Quality team (similar to previous presentations; see above) to stimulate adoption. • Publish a freely available “Practitioner’s Guide to Visual Data Profiling”. This Guide will describe best practice for the use of visualization to characterise data and investigate data quality, across a range of application domains. Public engagement • Develop and run a continuing professional development (CPD) course on data profiling, which will be made generally available. • A demonstration of the QuantiCode software has been accepted for the AI UK showcase (23rd/24th Mar 2021; postponed from 2020 due to Covid-19 https://www.turing.ac.uk/ai-uk). • Six-monthly reports/presentations to NHS England Data Quality team (similar to previous presentations; see above) to stimulate adoption. • Submit presentation about diagnostic persistence for AIUK 2024. • Public engagement via the annual Be Curious research festival.

Expected measurable benefits

It is expected that there will be benefit to health from the outputs provided to NHS Digital. England. The benefits will are hoped to occur as benefits to the health system, either directly (through improvements to data collections) or indirectly (better quality data used in analysis/research). The following headline benefits will are hoped to arise from NHS Digital’s England's usage of the tool: 1) Allow NHS Digital England to conduct integrity checking to identify occurrences of poor data quality in routinely collected data, for feedback to data providers. 2) Allow NHS Digital England to cross-reference variations in data quality issues with external influences (e.g. change [5 words unchanged] change in resource, change in service provision or structure, coding improvement initiatives). 3) Allow NHS Digital England to conduct bias checking to identify occurrences of poor data quality in routinely collected data, for feedback to data providers. 4) Support NHS Digital England in the development of new business rules for automated data profiling and feedback to data providers. 5) Improve NHS Digital’s England’s understandings of biases across time, geography and activity; NHS Digital’s England’s focus has necessarily been on data quality by provider (in order to [10 words unchanged] be addressed ), with less emphasis given to time, geography and activity. 6) Allow NHS Digital England to quantify the impact of data quality issues - e.g., whether the degree of missingness for certain conditions supports linkage match rates. 7) Improve the quality of the data used by NHS Digital England when conducting analysis. 8) Improve the quality of the data provided by NHS Digital England to third parties when conducting research/analysis. NHS Digital England manages many of the nation’s critical health and care data assets. It [21 words unchanged] Better use of health and care data will help those involved to: [5 paragraphs unchanged] Understanding the quality of data is essential in deciding whether it is fit for these uses. The benefits above will help NHS Digital England develop appropriate methods to monitor, challenge and highlight data quality issues to [9 words unchanged] the data if submitting or adjust their findings accordingly if just analysing. NHS Digital England also has a statutory responsibility, enacted in the Health and Social Care [29 words unchanged] to lead to new measurements in the report covering this statutory responsibility. The tools and methodologies developed through the QuantiCode project may also benefit other organisations in other sectors, but the data from NHS Digital England is only disseminated on the basis of there being a benefit to health. [1 paragraph unchanged] 1) Allow NHS Digital England to identify factors that affect a data quality issue called “diagnostic persistence”, which is widespread but poorly understood. 2) Support NHS Digital England in identifying new business rules for automated data profiling and feedback to data providers. 3) Improve the quality of the data used by NHS Digital England when conducting analysis. 4) Improve the robustness of data quality checks carried out by third parties when conducting research/analysis with NHS Digital England data. To put the importance of (1) in perspective, NHS Digital England has asked QuantiCode to investigate six conditions (autism, dementia, diabetes, ischemic heart disease, Parkinson’s and schizophrenia). Initial investigation has shown that, in our the 21 million episode APC dataset, 1 million episodes from 300,000 patients exhibit [12 words unchanged] or more of the conditions). By identifying factors that affect this issue, we the aim is to provide insights that allow NHS Digital England to design interventions (coding policy, validation, feedback, etc.) that improve diagnostic persistence [74 words unchanged] Understanding the patterns exhibited by those issues, makes it possible for NHS Digital England to design new business rules to automate the checking of future data (2), and exploit QuantiCode’s visualization techniques for further, continuous quality improvement. Note: The external organisations (NHS Digital England and six others; see (5a) for details) that are involved in this [68 words unchanged] techniques could be combined with ACE’s set visualization to analyse customer transactions. University of Leeds have developed a promising analysis workflow for the analysis of lost autism diagnoses. The current aim is to generalise that workflow to account for the vast majority (90+ %) of lost diagnoses across all of the above conditions that do not remit. That is hoped to yield benefits for the understanding of diagnostic persistence and, more generally, for combining visualization with machine learning for other types of complex health data analysis.

Benefits reported

[1 paragraph unchanged] Benefits to NHS Digital England and hospitals themselves: Our The ACE tool revealed unexpected gaps in sets of fields for some Admitted [40 words unchanged] the operation fields (OPERTN_nn). The latter are of particular concern to NHS Digital England because they affect the process that the Casemix Team use to calculate [15 words unchanged] by Results system for hospitals [6]. These findings from ACE allow NHS Digital England to provide feedback to the provider to rectify the problem, via the [85 words unchanged] revealing new issues. The results would be improved data quality for NHS Digital’s England’s own analysis [7], and in data and/or data quality notes provided to third parties [8]. [1 paragraph unchanged] ACE revealed two main issues with the APC data. One concerned missing [87 words unchanged] records and led the researchers to comment that they had asked NHS Digital England to derive the survival time for all patients based on the study [61 words unchanged] check their data on receipt, so that errors can be rectified (NHS Digital’s England’s time-limit is one month from receipt) [8]. [4 paragraphs unchanged] Important note: As stipulated in the project’s data agreements, the datasets from the different organisations (NHS Digital, England, LCC, Sainsbury’s, ResearchOne) are never linked and never used together in any analysis. None of the other organisations have access to the NHS Digital England data. Some health conditions (e.g., autism, dementia, diabetes, IHD, Parkinson's and schizophrenia) do not remit, so the relevant ICD code should always appear (“persist”) as one of the diagnoses in APC data, even if the patient is being treated for something else (e.g., a broken arm). However, for years the NHS has been concerned that diagnostic persistence is poor, with up to 50% of episodes missing those diagnoses for non-remitting conditions. NHS England asked University of Leeds to investigate this because the reasons are not known, and the investigations have proved challenging. Machine learning models based on decision trees, random forests and gradient boosting gave little in the way of insights because they could not handle the extreme complexity of health records data. However, a promising analysis workflow has been developed that combines set visualization techniques with on-the-fly computation, and this has been used to identify two previously unknown patterns that account for the 60% of lost autism diagnoses in this data agreement’s Admitted Patient Care dataset.

DARS-NIC-49164-R3G5K-v2.1 19 May 2021 to 30 September 2023
Title
QuantiCode: Admitted Patient Care Data
Commercial
Yes
Sublicensing
No
Datasets
1
Files released
0

Datasets: Hospital Episode Statistics Admitted Patient Care (HES APC)

What changed from DARS-NIC-49164-R3G5K-v1.8

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-49164-R3G5K-v1.8
FieldWasBecame
Start date2020-10-012021-05-19
Hospital Episode Statistics Admitted Patient Care (HES APC): common law duty of confidentialityNot statedDoes not include the flow of confidential data

Processing activities

[10 paragraphs unchanged] Microsoft Limited provide Cloud Services for the University of Leeds and are therefore listed as a data processor. They supply support to the system, but do not access data. Therefore, any access to the data held under this agreement would be considered a breach of the agreement. This includes granting of access to the database[s] containing the data.

Unchanged: Objective for processing, Expected output, Expected measurable benefits, Benefits reported.

Objective for processing

The University of Leeds are running a project called QuantiCode, to develop novel data mining & visualization tools/techniques that will transform people’s ability to analyse event sequence data such as electronic health records, social care data and retail data. These data often exist in databases that contain millions of records, each comprising hundreds of numerical and categorical variables. This project was funded by the Engineering and Physical Sciences Research Council (EPSRC; Mar 2016 – Feb 2020), and is being continued with the support of funding from the Alan Turing Institute (ATI; May 2019 – Dec 2021) and the University of Leeds (ongoing).

QuantiCode project involves seven external organisations (Bradford Teaching Hospitals NHS Foundation Trust, Consumerdata Limited, NHS Digital, Sainsbury’s Supermarket Limited, Leeds City Council, Leeds North Clinical Commissioning Group, AQL Limited), who are interested in the research, contribute through meetings, identifying analysis use cases and requirements for the tools/techniques, and receive reports and copies of the tools as outputs of the project. Three of the organisations (NHS Digital, Sainsbury’s, Leeds City Council) provided datasets, which are neither linked in any way nor used together in any analysis. None of the organisations have access to each other’s data. No record-level data from NHS Digital will be shared, or be in any way accessible, to third party organisations. Similarly, no aggregated data including small numbers (as defined in the HES Analysis Guide) may be shared with (or be in any way accessible to) third party organisations.

Under GDPR article 6(1), the lawful basis for the processing is “public task” (research is a task in the public interest), and under article 9(2) the lawful basis for processing special category data is (j) Archiving, research and statistics (with a basis in law). Like most British universities, the University of Leeds has charitable status because its primary purposes of advancing education and research are deemed to deliver a public benefit. The QuantiCode project is conducting fundamental academic research to acquire new knowledge about data mining & visualization techniques for analysing event sequence data. To help ensure that this research generalises to the scale and complexity of the real world, QuantiCode has been provided with health (from NHS Digital), social care (Leeds City Council) and retail data extracts (Sainsbury’s). As is clear from the Yielded Benefits (see 5d(iii) below), the project’s applied focus has been on health and the vast majority of the tangible benefits to date are in that domain. Those benefits include identifying an issue that affects the NHS’s Payment by Results system for hospitals, showing how the project's tools can be used to identify the source of data quality problems so hospitals are able to improve the future quality of their own data, and revealing previously unknown issues in the primary and secondary care data extracts that researchers were using to study melanoma/diabetes survival and acute myocardial infarction (heart attack) outcomes/hospitalization rates, respectively.

QuantiCode is multi-disciplinary, with academic researchers from computer science, maths, ethics, geography, and health. The NHS Digital data is only used in the part of the project that is developing novel tools/techniques for investigating data quality. Data quality is an issue that is particularly important for NHS Digital – it impacts directly on areas such as running the NHS (invoice validation, etc.), as well as some indirect benefits (research and analysis to develop policy/clinical guidelines). Specifically, QuantiCode’s objectives are to develop tools/techniques to: (1) comprehensively investigate data quality (number of missing values, percentiles, outliers, etc.), (2) make detailed investigations of patterns of missing data and, where possible, to identify factors associated with their source, and (3) investigate factors that affect diagnostic persistence (the most recent NHS Digital report shows that this affects 42% of HES data episodes for patients with conditions that do not remit https://digital.nhs.uk/data-and-information/data-tools-and-services/data-services/data-quality). By collaborating with the QuantiCode project, it is expected that there will be benefits for NHS Digital and to third parties who use NHS Digital data (e.g., see 5d (iii) Yielded Benefits, below).

To perform this research into data quality, QuantiCode requested pseudonymised admitted patient care (APC) data from hospital episode statistics (HES). A single year of HES Admitted Patient Care data was being requested in order to fulfil these objectives. The factors in (2) and (3) are currently unknown and may involve any variable in a dataset, which is why all variables (except those deemed sensitive or identifiable) were requested.

To address the GDPR principle of data minimisation the University of Leeds has only requested a single year of pseudonymised data, this the minimum amount of data needed in order to still be able to achieve the aims stated within this agreement. Data concerning individuals admitted to any English NHS hospital in a whole year is required for the purposes listed in (1), (2) and (3) to ensure that the tools/techniques scale to the volume and complexity of real NHS data. Additionally, while some data issues are common (3) others such of those in (2) can be rare, so potentially missing if only a part-year dataset, a subset of APC variables or limited geography was requested.

The University of Leeds is the sole data controller who also processes the data for the purposes described within this Agreement

Expected output

The intended outputs relating to health were:

1) Data analysis tool Version 1:This output will be a visual analytics tool, which allows users to gain an overview of missing data patterns and investigate data integrity in health datasets.

2) Research report 1: This report will describe the application of the tool to health data, and the benefits that the tool provides. The report will be submitted to a high-impact outlet such as the Journal of the American Medical Informatics Association (the pre-eminent journal for research into methods for analysing health data).

3) Data analysis tool Version 2:. This version of the visual analytics tool will allow users to investigate bias caused by data quality issues in health datasets.

4) Research report 2: This report will describe the application of the tool for bias investigations, and the benefits that the tool provides. The report will be submitted to the Journal of the American Medical Informatics Association.

Outputs 1 & 3 will not contain any data – a user will load their dataset into the tool to analyse ‘missingness’. Outputs 2 & 4 will contain only aggregate level data with small numbers suppressed in line with HES analysis guide. That data will be shown in figures that illustrate the usage of the tool.

The ultimate beneficiaries of this work will be the general public. For example, CCGs and local authorities analyse data to generate business intelligence for operations and investment, with the aim of providing us all with improved and more cost-effective services. Businesses similarly require business intelligence for operations and investment, which translates to jobs and other economic benefits. To bridge the gap between these indirect benefits from the project and popular interest in big data, the University of Leeds will conduct a range of public engagement activities which include live demonstrations at the annual Leeds Festival of Science, a short film, an on-line tutorial about the ethics surrounding data analytics, and publishing articles in the popular scientific press.

As of September 2020 the following outputs have been produced, or are expected to be produced:

Note: all outputs either did not contain any NHS Digital data or were aggregated with small numbers supressed in line with HES Analysis Guidance.

Software (see Outputs 1 & 3, above)

• A data analysis tool called ACE. Version 1 was released to the QuantiCode project partners on 20/7/18. Version 2 was released to the QuantiCode project partners on 6/11/19. This version of ACE contained substantial new functionality to provide set visualization functionality that is generic, so that sets of both missing and present data can be analysed. Some parts of the ACE were re-engineered to speed up processing and make it more scalable. As part of that, ACE was tested on a 64 GB desktop PC with datasets that contained up to 21 million records, 417 fields (i.e., sets) and 141,000 unique combinations of field (set intersections).

• A Python/Pandas library (“QCprofiling”) for computing a suite of data quality checks and visualizing the output in matrices of miniature visualizations has been developed. The checks include data type, missing values, unique values, value lengths, character pattern, percentiles, outliers and example values.

Peer reviewed academic papers (*indicates outputs that concern “Non-health benefits” outlined in 5d; all other outputs are directly related to the NHS Digital data request)

• Research reports (see Outputs 2 & 4, above): These “reports” are in fact academic papers. The first journal manuscript was submitted to the IEEE Transactions on Visualization and Computer Graphics on 24/7/2018, and unfortunately not accepted. The second was submitted to the Journal of the American Medical Informatics Association on 11/6/2020 and also not accepted. Revisions and resubmissions are planned, and in both cases further work is required (see 5a. Objective for processing, and 5b. Processing activities).

• *Adnan, M., & Ruddle, R.A. (2018). A set-based visual analytics approach to analyze retail data. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA).

• Adnan, M., Nguyen, P. H., Ruddle, R. A. & Turkay, C. (2019). Visual analytics of event data using multiple mining methods. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA).

• Ruddle, R. A., & Hall, M. S. (2019). Using miniature visualizations of descriptive statistics to investigate the quality of electronic health records. Proceedings of the 12th International Joint Conference on Biomedical Engineering Systems and Technologies - Volume 5 HEALTHINF: HEALTHINF, ISBN 978-989-758-353-7, pages 230-238.

Reports (*indicates outputs that concern “Non-health benefits”)

• Feedback to NHS Digital, correcting an error in the APC Data Dictionary (ref: NIC-309404-B7V2J; 20/6/2019)

• “QuantiCode: Benefits to Health and/or Social Care” report for NHS Digital (17/7/2019)

• *Understanding customers’ missions from the products they purchase. Technical report for Sainsbury’s. June 2019.

Presentations (*indicates outputs that concern “Non-health benefits”)

• *Briefing to Leeds City Council (LCC) at Adult Social Care Information Management and Technology (IM&T) Team Meeting (8/11/16)

• “Exploring missingness patterns in Admitted Patient Care (APC) data using ACE” to NHS Digital Head of Information Utilisation and colleagues (16/10/2017)

• Demo of ACE to NHS Digital and NHS England staff, with discussion of application to issues such as diagnostic persistence (14/11/2017)

• Demo of ACE to NHS Digital Data Quality and Casemix teams (1/12/2017)

• “Visualizing data profiles and analysis pipelines”. Presentation at Visualization for Data Science and AI. Alan Turing Institute, London, 13/9/2019.

• *Discussion of strategies for investigating data quality for LCC’s adult social care business intelligence dashboard (29/10/19)

Keynote conference presentation

• “Visualizing health data – from fundamental research to successful applications”. 13th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC), Malta, 26/2/2020

Public engagement

• Exhibition and hand-on demo of prototype of ACE at the University of Leeds Be Curious research festival (25/3/17). 847 adults and children attended.

• Release a film about “Visualizing the Quality of Data” on YouTube https://tinyurl.com/VizDataQuality (20/9/2019)

Other

• Training workshops for ACE were run in July and December 2018

The following new outputs are expected:

Note: all outputs all outputs either did not contain any NHS Digital data or will be aggregated with small numbers supressed in line with HES Analysis Guidance.

Software

• The ACE software will be made widely available for free non-commercial use, initially leveraging the existing Alan Turing Institute Health Programme (includes researchers from 13 partner universities) and the HDRUK Health Data Research Hubs (includes NHS-accredited cloud-based IT platforms). This involves publicity (Q4 2020 onward) and follow-up support of interested researchers, providing additional evaluation material that is needed for peer reviewed academic papers (see below)

• Release version 1 of the RFviz software (Q4 2021) for free use by the QuantiCode project partners and free non-commercial use by others. This software will allow a data-driven approach to be taken for the investigation of diagnostic persistence and similar data quality issues.

• New workflow software, to make the QCprofiling library more accessible to users for comprehensively investigating data quality and developing new business rules for automated data profiling (Q2 2022).

Peer reviewed academic papers

• Technical paper describing the ACE tool. For submission to a high-impact outlet such as Information Visualization (Q3 2021).

• Paper describing the application of ACE to investigating electronic health records. For submission to a high-impact outlet such as the Journal of the American Medical Informatics Association (Q1 2022).

• Paper describing the design and evaluation of RFViz to the investigation of diagnostic persistence. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q2 2022).

• Paper describing the application of the new workflow software. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q4 2022).

Reports & presentations

• Six-monthly reports/presentations to NHS Digital Data Quality team (similar to previous presentations; see above) to stimulate adoption.

Public engagement

• A demonstration of the QuantiCode software has been accepted for the AI UK showcase (23rd/24th Mar 2021; postponed from 2020 due to Covid-19 https://www.turing.ac.uk/ai-uk).

Benefits reported

To date there have been three main types of benefit. In the following, the eight original Expected Measurable Benefits (see 5d(ii)) are referenced as [1], [2], etc.

Benefits to NHS Digital and hospitals themselves: Our ACE tool revealed unexpected gaps in sets of fields for some Admitted Patient Care (APC) data records, and pinpointed the source of 90+% of the gaps as being a particular admission method at a single hospital [1]. Some of the gaps were in the diagnosis code fields (DIAG_nn) and others were in the operation fields (OPERTN_nn). The latter are of particular concern to NHS Digital because they affect the process that the Casemix Team use to calculate Healthcare Resource Groups (HRGs), which are one of the building blocks of the NHS’s Payment by Results system for hospitals [6]. These findings from ACE allow NHS Digital to provide feedback to the provider to rectify the problem, via the existing HES data quality lifecycle. ACE has subsequently revealed similar gaps in the date fields for operations (MYOPDATE_nn) and inconsistencies between the operation fields and their corresponding date fields. The speed (a few minutes) with which users were able to interactively find these unexpected missing data issues with ACE, contrasts with the years that the issues have existed yet gone undetected. It would be straightforward to implement new business rules to automatically check for these gaps [4], to complement ACE’s power in exploratory analysis for revealing new issues. The results would be improved data quality for NHS Digital’s own analysis [7], and in data and/or data quality notes provided to third parties [8].

Benefits to third party users of electronic health records: QuantiCode has concentrated its effort on assisting researchers working on two other projects – one using APC data (DARS-NIC-17649-G0X4B) to study long-term outcomes and hospitalization rates for survivors of acute myocardial infarction (heart attack), and the other using ResearchOne primary care data to study the survival from melanoma of patients with type 2 diabetes.

ACE revealed two main issues with the APC data. One concerned missing admission and episode start dates. In principle they can be estimated from each other if only one is missing, but ACE showed that that was hardly ever the case (99.8% of records that were missing one of those dates was also missing the other). However, ACE also let the researchers localize the origin to maternity records, and therefore make an informed decision to discard the records because the project was conducting heart attack research [8]. The other issue concerned survival time, which was missing in 12 million records and led the researchers to comment that they had asked NHS Digital to derive the survival time for all patients based on the study census date, regardless of their mortality status, but clearly that had not been done. This finding made the researchers realise that they had to estimate survival time for those patients who did not die prior to the census date, losing precision and precluding analyses of short-term survival patterns within 30 days of admission. In future, ACE would allow the researchers to check their data on receipt, so that errors can be rectified (NHS Digital’s time-limit is one month from receipt) [8].

Complementing ACE, the QCprofiling software revealed a number of longitudinal data quality issues [5] across nine years of APC data extracts (150 million records in total; 116 different fields). Some were the result of policy change (e.g., all values of Clinical Commissioning Group fields missing in the early extracts but not the latest, and longitudinal differences in coding depth for the operation fields) [2], others from improvements to business rules (the primary diagnosis was sometimes missing in the early extracts but never from the latest ones), some were prevalent across all of the extracts (e.g., widespread use of punctuation characters in diagnosis codes, which complicates data cleaning), one affected data linkage [6] (the use of two different (16 vs. 32 character) pseudonymized patient identifiers in one of the extracts).

The QCprofiling software also revealed data quality issues with the primary care data (90 million records; 14 tables; 41 different fields). These included validation error text appearing in a birth year field, clinical codes being padded out with ‘.’ characters (i.e., inconsistent coding precision between different data tables), and the distribution of blood pressure measurement dates leading the researchers to realise that they had not been provided with the historical data that they had expected [8]. All of these issues were new to the researchers, despite the fact that they had had the data for several months. As with the APC data, the QCprofiling software was a game-changer for investigating data quality because it allowed the researchers to see patterns that spanned a large number of records/tables/fields at the glance of an eye.

Non-health benefits: ACE was also evaluated and used for more exploratory analysis with data from two of QuantiCode’s other external partners (Leeds City Council (LCC) and Sainsbury’s), to help generalise the design of ACE. LCC were shown how ACE’s set visualization techniques could be used to investigate patterns of missing data in an anonymized adult social care service provision dataset. The main benefit lies in making explicit the patterns for discussion by analysts and stakeholders prior to use of the data for other purposes such as modelling future service demand. In some cases explanations could be offered for unexpected patterns (e.g., tacit knowledge about historical changes in which data were recorded), and in others further investigation would be required.

Sainsbury’s were shown how frequent itemset mining techniques could be combined with ACE’s set visualization to analyse customer transactions (e.g., What items are bought together? How to highlight differences and similarities between stores, or over time?). This was exploratory research, indicating a possible approach that would allow customer transactions to be analysed at a much finer level of detail than is possible with Sainsbury’s current methods.

Important note: As stipulated in the project’s data agreements, the datasets from the different organisations (NHS Digital, LCC, Sainsbury’s, ResearchOne) are never linked and never used together in any analysis. None of the other organisations have access to the NHS Digital data.

DARS-NIC-49164-R3G5K-v1.8 1 October 2020 to 30 September 2023
Title
QuantiCode: Admitted Patient Care Data
Commercial
Yes
Sublicensing
No
Datasets
1
Files released
0

Datasets: Hospital Episode Statistics Admitted Patient Care (HES APC)

What changed from DARS-NIC-49164-R3G5K-v0.4

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-49164-R3G5K-v0.4
FieldWasBecame
Start date2017-10-012020-10-01
End date2020-09-302023-09-30
Hospital Episode Statistics Admitted Patient Care (HES APC): legal basisHealth and Social Care Act 2012 – s261(2)(b)(ii)Health and Social Care Act 2012 - s261 - 'Other dissemination of information'

Objective for processing

The University of Leeds are running a project called QuantiCode. This project is funded by the Engineering and Physical Sciences Research Council (EPSRC) from March 2016 to February 2019 (details: http://gtr.rcuk.ac.uk/projects?ref=EP%2FN013980%2F1). The overall QuantiCode project is divided into 3 stages: The University of Leeds are running a project called QuantiCode, to develop novel data mining & visualization tools/techniques that will transform people’s ability to analyse event sequence data such as electronic health records, social care data and retail data. These data often exist in databases that contain millions of records, each comprising hundreds of numerical and categorical variables. This project was funded by the Engineering and Physical Sciences Research Council (EPSRC; Mar 2016 – Feb 2020), and is being continued with the support of funding from the Alan Turing Institute (ATI; May 2019 – Dec 2021) and the University of Leeds (ongoing). 1) Data fusion, covering tools for data linkage and visualizing data quality, and thought leadership in data governance. QuantiCode project involves seven external organisations (Bradford Teaching Hospitals NHS Foundation Trust, Consumerdata Limited, NHS Digital, Sainsbury’s Supermarket Limited, Leeds City Council, Leeds North Clinical Commissioning Group, AQL Limited), who are interested in the research, contribute through meetings, identifying analysis use cases and requirements for the tools/techniques, and receive reports and copies of the tools as outputs of the project. Three of the organisations (NHS Digital, Sainsbury’s, Leeds City Council) provided datasets, which are neither linked in any way nor used together in any analysis. None of the organisations have access to each other’s data. No record-level data from NHS Digital will be shared, or be in any way accessible, to third party organisations. Similarly, no aggregated data including small numbers (as defined in the HES Analysis Guide) may be shared with (or be in any way accessible to) third party organisations. 2) Analytical techniques that allow users to interactively mine longitudinal data for patterns. Under GDPR article 6(1), the lawful basis for the processing is “public task” (research is a task in the public interest), and under article 9(2) the lawful basis for processing special category data is (j) Archiving, research and statistics (with a basis in law). Like most British universities, the University of Leeds has charitable status because its primary purposes of advancing education and research are deemed to deliver a public benefit. The QuantiCode project is conducting fundamental academic research to acquire new knowledge about data mining & visualization techniques for analysing event sequence data. To help ensure that this research generalises to the scale and complexity of the real world, QuantiCode has been provided with health (from NHS Digital), social care (Leeds City Council) and retail data extracts (Sainsbury’s). As is clear from the Yielded Benefits (see 5d(iii) below), the project’s applied focus has been on health and the vast majority of the tangible benefits to date are in that domain. Those benefits include identifying an issue that affects the NHS’s Payment by Results system for hospitals, showing how the project's tools can be used to identify the source of data quality problems so hospitals are able to improve the future quality of their own data, and revealing previously unknown issues in the primary and secondary care data extracts that researchers were using to study melanoma/diabetes survival and acute myocardial infarction (heart attack) outcomes/hospitalization rates, respectively. 3) Governance-aware abstraction techniques that allow users to explore how complex data may (or may not) be simplified to reveal important patterns. QuantiCode is multi-disciplinary, with academic researchers from computer science, maths, ethics, geography, and health. The NHS Digital data is only used in the part of the project that is developing novel tools/techniques for investigating data quality. Data quality is an issue that is particularly important for NHS Digital – it impacts directly on areas such as running the NHS (invoice validation, etc.), as well as some indirect benefits (research and analysis to develop policy/clinical guidelines). Specifically, QuantiCode’s objectives are to develop tools/techniques to: (1) comprehensively investigate data quality (number of missing values, percentiles, outliers, etc.), (2) make detailed investigations of patterns of missing data and, where possible, to identify factors associated with their source, and (3) investigate factors that affect diagnostic persistence (the most recent NHS Digital report shows that this affects 42% of HES data episodes for patients with conditions that do not remit https://digital.nhs.uk/data-and-information/data-tools-and-services/data-services/data-quality). By collaborating with the QuantiCode project, it is expected that there will be benefits for NHS Digital and to third parties who use NHS Digital data (e.g., see 5d (iii) Yielded Benefits, below). The work will be evaluated by the collaborating organisations to ensure that the solutions are applicable in the real world. To perform this research into data quality, QuantiCode requested pseudonymised admitted patient care (APC) data from hospital episode statistics (HES). A single year of HES Admitted Patient Care data was being requested in order to fulfil these objectives. The factors in (2) and (3) are currently unknown and may involve any variable in a dataset, which is why all variables (except those deemed sensitive or identifiable) were requested. The QuantiCode project involves a number of collaborating organisations (University of Leeds, Bradford Teaching Hospitals NHS Foundation Trust, Consumerdata Limited, NHS Digital, Sainsbury’s Supermarket Limited, Leeds City Council, Leeds North Clinical Commissioning Group, AQ Limited), who are interested in this work and are supplying datasets and will receive reports and tools as outputs of the project. NHS Digital is one of these organisations, as NHS Digital has a number of very large datasets, and Data Quality is an issue which is particularly important – it impacts directly on areas such as running the NHS (invoice validation, etc.), as well as some indirect benefits (research and analysis to develop policy/clinical guidelines). By collaborating with the QuantiCode project, it is expected that there will be benefits for NHS Digital and benefits to healthcare more widely as a result. To address the GDPR principle of data minimisation the University of Leeds has only requested a single year of pseudonymised data, this the minimum amount of data needed in order to still be able to achieve the aims stated within this agreement. Data concerning individuals admitted to any English NHS hospital in a whole year is required for the purposes listed in (1), (2) and (3) to ensure that the tools/techniques scale to the volume and complexity of real NHS data. Additionally, while some data issues are common (3) others such of those in (2) can be rare, so potentially missing if only a part-year dataset, a subset of APC variables or limited geography was requested. Any data provided from NHS Digital to the University of Leeds for this project will not be linked in any way with the other datasets being used. No record-level data from NHS Digital will be shared, or be in any way accessible, to third party organisations (including the other collaborating organisations). Similarly, no aggregated data including small numbers (as defined in the HES Analysis Guide) may be shared with (or be in any way accessible to) third party organisations The University of Leeds is the sole data controller who also processes the data for the purposes described within this Agreement Within the bounds of the QuantiCode project, the purpose for processing healthcare data from NHS Digital is to allow different designs of visualisation and machine learning technique to be compared for their ability to meet user requirements, and allow the QuantiCode data analysis tool to be tested prior to release to NHS Digital for in-depth evaluation. A single year of HES Admitted Patient Care data is being requested in order to fulfil these objectives. The single year of hospital data is important as there are data quality process which take place at the end of the year, meaning that a sub-set of data (e.g. 6 months)will display different characteristics to the finalised (“Annual Refresh”) data produced after the end of the financial year. Without healthcare data, it is possible that the methodologies/tools cannot be applied to healthcare data, and the opportunity to improve the healthcare data will be lost. In addition, a specific purpose for processing the health data is to make a detailed analysis of patterns of “missingness” in data (the manner in which data are missing from a sample of a population). The goal is to be able to investigate patterns that involve: (a) several (3+) variables missing together, and/or (b) are dependent on the particular value of other variables (e.g., provider code and admission type). These patterns are currently unknown and such missingness may involve any variable in a dataset, which is why all variables (except those deemed sensitive or identifiable) have been requested.

Processing activities

The dataset (to be) provided by NHS Digital consists of pseudonymised, record-level data. The dataset will be transferred by the University of Leeds’ Integrated Research Campus (IRC) Data Services Team using NHS Digitals Secure Electronic File Transfer system (SEFT) The data will only be accessed by substantive employees of the University who are contributing to the project. The University of Leeds do not flow any data to NHS Digital. Under a previous iteration of this agreement NHS Digital flowed pseudonymised HES APC (2015/16) to the University of Leeds. This data was sent via the Secure Electronic File Transfer System (SEFT). There are no subsequent flows of data. The patient data from NHS Digital will not be linked with any other dataset. The data will only be accessed by substantive employees of the University who are contributing to the project, all of whom have been trained in data protection and confidentiality. The data is in a category that the University classifies as IRC-Confidential. The data will be stored on University computer systems that are NHS DSPT and ISO 27001 accredited and accessed in accordance with ISO 27001 and University policies, including the information protection policy. The Quanticode project will develop tools and methodologies for investigating data quality across a number of datasets. The aim is to test these tools and methodologies across a variety of datasets, including health data. Although the QuantiCode project as a whole will examine issues around data linkage, that work will not involve the use of data from NHS Digital data. The patient data from NHS Digital will not be linked with any other dataset, including those received from Quanticode Partners. The QuantiCode project will process datasets provided by other organisations involved in the project, and the methodologies/tools will be developed to apply across datasets as much as possible. One of the aims of this is to ensure that techniques are developed which are generically applicable, rather than each sector needing to develop their own tools. There will be no attempt to re-identify individuals from the pseudonymised data. The NHS Digital dataset will be used to help design and test a new data analysis tool, which allows users to investigate data quality in health records. The development process for the tool will involve: (a) characterising the dataset so that scalable algorithms and appropriate statistical models may be designed, (b) using the dataset as an exemplar to design and implement new interactive visualization techniques for data quality investigation, (c) running the data to refine the statistical models, and (d) testing of the data analysis tool prior to release to NHS Digital for in-depth evaluation. The Quanticode project is developing tools and methodologies for investigating data quality across a number of datasets. The aim is to test these tools and methodologies across a variety of datasets, including health data. The NHS Digital dataset will only be used where necessary for the purposes in this agreement. Live data will not be used for early stages of development/testing when it would be more appropriate to use test data. The QuantiCode project will process datasets provided by organisations involved in the project (including NHS Digital), and the methodologies/tools will be developed to apply across datasets as much as possible. One of the aims of this is to ensure that techniques are developed which are generically applicable, rather than each sector needing to develop their own tools. The datasets from the different organisations are never linked and never used together in any analysis. None of the other organisations will have access to the NHS Digital data. More specifically, the NHS Digital dataset will be used to help design and test new data analysis tools (called QCprofiling, ACE and RFviz), which allow users to investigate data quality in tabular data such as electronic health records. The following types of computation will be performed: (a) calculate descriptive statistics (number of missing values, percentiles, outliers, etc.), (b) calculate sets and set intersections, and (c) data mining (using well-established methods such as random forests, gradient boosting, entropy and information gain ). To put these in the context of achieving QuantiCode’s purpose, (a) is central to allowing users to comprehensively investigate data quality (as enabled by our QCprofiling software), and (b) is central to allowing users to go further with detailed investigations of patterns of missing data (the ACE software). Data mining techniques such as entropy and information gain (c) are essential for pinpointing the origin of patterns of missing data in ACE, and techniques such as random forests and gradient boosting are essential for investigating diagnostic persistence (the RFviz software). Output from the computations will be visualized using a wide range of techniques, as appropriate to the type and scale of data. For comprehensive data quality investigations and investigating diagnostic persistence (QCprofiling and RFviz output, respectively) those techniques span bar charts, histograms, scatterplots, box plots and character maps. ACE uses a smaller suite of techniques (bar charts, histograms and heatmaps). Real-world data is essential both for designing the above tools and testing them. That testing is important for informing design choices (e.g., what are the speed/accuracy trade-offs of data mining methods such as random forests vs. gradient boosting) and testing the scalability of the tools to reduce bottlenecks and improve performance. The NHS Digital dataset will only be used where necessary, and will only be used for the purposes described in this agreement. Live data will not be used for early stages of development/testing when it would be more appropriate to use test data. [1 paragraph unchanged]

Expected output

The intended outputs relating to health are: were: 1) Data analysis tool Version 1 (target date 28/2/18). This 1:This output will be a visual analytics tool, which allows users to gain an overview of missing data patterns and investigate data integrity in health datasets. 2) Research report 1 (target date 30/9/18). 1: This report will describe the application of the tool to health data, [25 words unchanged] Association (the pre-eminent journal for research into methods for analysing health data). 3) Data analysis tool Version 2 (target date 31/5/19). 2:. This version of the visual analytics tool will allow users to investigate bias caused by data quality issues in health datasets. 4) Research report 2 (target date 30/11/19). 2: This report will describe the application of the tool for bias investigations, [9 words unchanged] will be submitted to the Journal of the American Medical Informatics Association. Outputs 1 & 3 will not contain any data – a user will load their dataset into the tool to analyse missingness. ‘missingness’. Outputs 2 & 4 will contain only aggregate level data with small [10 words unchanged] will be shown in figures that illustrate the usage of the tool. [1 paragraph unchanged] As of September 2020 the following outputs have been produced, or are expected to be produced: Note: all outputs either did not contain any NHS Digital data or were aggregated with small numbers supressed in line with HES Analysis Guidance. Software (see Outputs 1 & 3, above) • A data analysis tool called ACE. Version 1 was released to the QuantiCode project partners on 20/7/18. Version 2 was released to the QuantiCode project partners on 6/11/19. This version of ACE contained substantial new functionality to provide set visualization functionality that is generic, so that sets of both missing and present data can be analysed. Some parts of the ACE were re-engineered to speed up processing and make it more scalable. As part of that, ACE was tested on a 64 GB desktop PC with datasets that contained up to 21 million records, 417 fields (i.e., sets) and 141,000 unique combinations of field (set intersections). • A Python/Pandas library (“QCprofiling”) for computing a suite of data quality checks and visualizing the output in matrices of miniature visualizations has been developed. The checks include data type, missing values, unique values, value lengths, character pattern, percentiles, outliers and example values. Peer reviewed academic papers (*indicates outputs that concern “Non-health benefits” outlined in 5d; all other outputs are directly related to the NHS Digital data request) • Research reports (see Outputs 2 & 4, above): These “reports” are in fact academic papers. The first journal manuscript was submitted to the IEEE Transactions on Visualization and Computer Graphics on 24/7/2018, and unfortunately not accepted. The second was submitted to the Journal of the American Medical Informatics Association on 11/6/2020 and also not accepted. Revisions and resubmissions are planned, and in both cases further work is required (see 5a. Objective for processing, and 5b. Processing activities). • *Adnan, M., & Ruddle, R.A. (2018). A set-based visual analytics approach to analyze retail data. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA). • Adnan, M., Nguyen, P. H., Ruddle, R. A. & Turkay, C. (2019). Visual analytics of event data using multiple mining methods. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA). • Ruddle, R. A., & Hall, M. S. (2019). Using miniature visualizations of descriptive statistics to investigate the quality of electronic health records. Proceedings of the 12th International Joint Conference on Biomedical Engineering Systems and Technologies - Volume 5 HEALTHINF: HEALTHINF, ISBN 978-989-758-353-7, pages 230-238. Reports (*indicates outputs that concern “Non-health benefits”) • Feedback to NHS Digital, correcting an error in the APC Data Dictionary (ref: NIC-309404-B7V2J; 20/6/2019) • “QuantiCode: Benefits to Health and/or Social Care” report for NHS Digital (17/7/2019) • *Understanding customers’ missions from the products they purchase. Technical report for Sainsbury’s. June 2019. Presentations (*indicates outputs that concern “Non-health benefits”) • *Briefing to Leeds City Council (LCC) at Adult Social Care Information Management and Technology (IM&T) Team Meeting (8/11/16) • “Exploring missingness patterns in Admitted Patient Care (APC) data using ACE” to NHS Digital Head of Information Utilisation and colleagues (16/10/2017) • Demo of ACE to NHS Digital and NHS England staff, with discussion of application to issues such as diagnostic persistence (14/11/2017) • Demo of ACE to NHS Digital Data Quality and Casemix teams (1/12/2017) • “Visualizing data profiles and analysis pipelines”. Presentation at Visualization for Data Science and AI. Alan Turing Institute, London, 13/9/2019. • *Discussion of strategies for investigating data quality for LCC’s adult social care business intelligence dashboard (29/10/19) Keynote conference presentation • “Visualizing health data – from fundamental research to successful applications”. 13th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC), Malta, 26/2/2020 Public engagement • Exhibition and hand-on demo of prototype of ACE at the University of Leeds Be Curious research festival (25/3/17). 847 adults and children attended. • Release a film about “Visualizing the Quality of Data” on YouTube https://tinyurl.com/VizDataQuality (20/9/2019) Other • Training workshops for ACE were run in July and December 2018 The following new outputs are expected: Note: all outputs all outputs either did not contain any NHS Digital data or will be aggregated with small numbers supressed in line with HES Analysis Guidance. Software • The ACE software will be made widely available for free non-commercial use, initially leveraging the existing Alan Turing Institute Health Programme (includes researchers from 13 partner universities) and the HDRUK Health Data Research Hubs (includes NHS-accredited cloud-based IT platforms). This involves publicity (Q4 2020 onward) and follow-up support of interested researchers, providing additional evaluation material that is needed for peer reviewed academic papers (see below) • Release version 1 of the RFviz software (Q4 2021) for free use by the QuantiCode project partners and free non-commercial use by others. This software will allow a data-driven approach to be taken for the investigation of diagnostic persistence and similar data quality issues. • New workflow software, to make the QCprofiling library more accessible to users for comprehensively investigating data quality and developing new business rules for automated data profiling (Q2 2022). Peer reviewed academic papers • Technical paper describing the ACE tool. For submission to a high-impact outlet such as Information Visualization (Q3 2021). • Paper describing the application of ACE to investigating electronic health records. For submission to a high-impact outlet such as the Journal of the American Medical Informatics Association (Q1 2022). • Paper describing the design and evaluation of RFViz to the investigation of diagnostic persistence. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q2 2022). • Paper describing the application of the new workflow software. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q4 2022). Reports & presentations • Six-monthly reports/presentations to NHS Digital Data Quality team (similar to previous presentations; see above) to stimulate adoption. Public engagement • A demonstration of the QuantiCode software has been accepted for the AI UK showcase (23rd/24th Mar 2021; postponed from 2020 due to Covid-19 https://www.turing.ac.uk/ai-uk).

Expected measurable benefits

[2 paragraphs unchanged] 1) Allow NHS Digital to conduct integrity checking to identify occurrences of poor data quality in routinely collected data, for feedback to data providers (target date 30/9/18). providers. 2) Allow NHS Digital to cross-reference variations in data quality issues with [8 words unchanged] priorities, change in resource, change in service provision or structure, coding improvement initiatives) (target date 30/4/19). initiatives). 3) Allow NHS Digital to conduct bias checking to identify occurrences of poor data quality in routinely collected data, for feedback to data providers (target date 30/9/19). providers. 4) Support NHS Digital in the development of new business rules for automated data profiling and feedback to data providers (target date 30/11/20). providers. [13 paragraphs unchanged] The following benefits are expected to arise during the time-period of this data sharing agreement extension: 1) Allow NHS Digital to identify factors that affect a data quality issue called “diagnostic persistence”, which is widespread but poorly understood. 2) Support NHS Digital in identifying new business rules for automated data profiling and feedback to data providers. 3) Improve the quality of the data used by NHS Digital when conducting analysis. 4) Improve the robustness of data quality checks carried out by third parties when conducting research/analysis with NHS Digital data. To put the importance of (1) in perspective, NHS Digital has asked QuantiCode to investigate six conditions (autism, dementia, diabetes, ischemic heart disease, Parkinson’s and schizophrenia). Initial investigation has shown that, in our 21 million episode APC dataset, 1 million episodes from 300,000 patients exhibit that data quality issue (diagnoses not persisting when they should for one or more of the conditions). By identifying factors that affect this issue, we aim to provide insights that allow NHS Digital to design interventions (coding policy, validation, feedback, etc.) that improve diagnostic persistence and, therefore, directly improve data quality. As a consequence, the hospitals would have more accurate information for treating their own patients, and health analysts (3) and researchers (4) would have more robust data for epidemiological studies. In a production environment, data quality checks need to be automated. QuantiCode has already yielded benefits by showing how effective interactive visualization is for discovering previously unknown issues (e.g., in diagnosis code, operation and date fields; see 5d(iii)). Understanding the patterns exhibited by those issues, makes it possible for NHS Digital to design new business rules to automate the checking of future data (2), and exploit QuantiCode’s visualization techniques for further, continuous quality improvement. Note: The external organisations (NHS Digital and six others; see (5a) for details) that are involved in this project received copies of the reports and software tools that the project produced up until 29/2/2020 (the end of the project’s original EPSRC funding). Additionally, and as described below in (iii) Yielded Benefits, Leeds City Council were shown how ACE’s set visualization techniques could be used to investigate patterns of missing data in an anonymized adult social care service provision dataset. Sainsbury’s were shown how frequent itemset mining techniques could be combined with ACE’s set visualization to analyse customer transactions.

Benefits reported

Yielded Benefits is not a requirement for new applications. To date there have been three main types of benefit. In the following, the eight original Expected Measurable Benefits (see 5d(ii)) are referenced as [1], [2], etc. Benefits to NHS Digital and hospitals themselves: Our ACE tool revealed unexpected gaps in sets of fields for some Admitted Patient Care (APC) data records, and pinpointed the source of 90+% of the gaps as being a particular admission method at a single hospital [1]. Some of the gaps were in the diagnosis code fields (DIAG_nn) and others were in the operation fields (OPERTN_nn). The latter are of particular concern to NHS Digital because they affect the process that the Casemix Team use to calculate Healthcare Resource Groups (HRGs), which are one of the building blocks of the NHS’s Payment by Results system for hospitals [6]. These findings from ACE allow NHS Digital to provide feedback to the provider to rectify the problem, via the existing HES data quality lifecycle. ACE has subsequently revealed similar gaps in the date fields for operations (MYOPDATE_nn) and inconsistencies between the operation fields and their corresponding date fields. The speed (a few minutes) with which users were able to interactively find these unexpected missing data issues with ACE, contrasts with the years that the issues have existed yet gone undetected. It would be straightforward to implement new business rules to automatically check for these gaps [4], to complement ACE’s power in exploratory analysis for revealing new issues. The results would be improved data quality for NHS Digital’s own analysis [7], and in data and/or data quality notes provided to third parties [8]. Benefits to third party users of electronic health records: QuantiCode has concentrated its effort on assisting researchers working on two other projects – one using APC data (DARS-NIC-17649-G0X4B) to study long-term outcomes and hospitalization rates for survivors of acute myocardial infarction (heart attack), and the other using ResearchOne primary care data to study the survival from melanoma of patients with type 2 diabetes. ACE revealed two main issues with the APC data. One concerned missing admission and episode start dates. In principle they can be estimated from each other if only one is missing, but ACE showed that that was hardly ever the case (99.8% of records that were missing one of those dates was also missing the other). However, ACE also let the researchers localize the origin to maternity records, and therefore make an informed decision to discard the records because the project was conducting heart attack research [8]. The other issue concerned survival time, which was missing in 12 million records and led the researchers to comment that they had asked NHS Digital to derive the survival time for all patients based on the study census date, regardless of their mortality status, but clearly that had not been done. This finding made the researchers realise that they had to estimate survival time for those patients who did not die prior to the census date, losing precision and precluding analyses of short-term survival patterns within 30 days of admission. In future, ACE would allow the researchers to check their data on receipt, so that errors can be rectified (NHS Digital’s time-limit is one month from receipt) [8]. Complementing ACE, the QCprofiling software revealed a number of longitudinal data quality issues [5] across nine years of APC data extracts (150 million records in total; 116 different fields). Some were the result of policy change (e.g., all values of Clinical Commissioning Group fields missing in the early extracts but not the latest, and longitudinal differences in coding depth for the operation fields) [2], others from improvements to business rules (the primary diagnosis was sometimes missing in the early extracts but never from the latest ones), some were prevalent across all of the extracts (e.g., widespread use of punctuation characters in diagnosis codes, which complicates data cleaning), one affected data linkage [6] (the use of two different (16 vs. 32 character) pseudonymized patient identifiers in one of the extracts). The QCprofiling software also revealed data quality issues with the primary care data (90 million records; 14 tables; 41 different fields). These included validation error text appearing in a birth year field, clinical codes being padded out with ‘.’ characters (i.e., inconsistent coding precision between different data tables), and the distribution of blood pressure measurement dates leading the researchers to realise that they had not been provided with the historical data that they had expected [8]. All of these issues were new to the researchers, despite the fact that they had had the data for several months. As with the APC data, the QCprofiling software was a game-changer for investigating data quality because it allowed the researchers to see patterns that spanned a large number of records/tables/fields at the glance of an eye. Non-health benefits: ACE was also evaluated and used for more exploratory analysis with data from two of QuantiCode’s other external partners (Leeds City Council (LCC) and Sainsbury’s), to help generalise the design of ACE. LCC were shown how ACE’s set visualization techniques could be used to investigate patterns of missing data in an anonymized adult social care service provision dataset. The main benefit lies in making explicit the patterns for discussion by analysts and stakeholders prior to use of the data for other purposes such as modelling future service demand. In some cases explanations could be offered for unexpected patterns (e.g., tacit knowledge about historical changes in which data were recorded), and in others further investigation would be required. Sainsbury’s were shown how frequent itemset mining techniques could be combined with ACE’s set visualization to analyse customer transactions (e.g., What items are bought together? How to highlight differences and similarities between stores, or over time?). This was exploratory research, indicating a possible approach that would allow customer transactions to be analysed at a much finer level of detail than is possible with Sainsbury’s current methods. Important note: As stipulated in the project’s data agreements, the datasets from the different organisations (NHS Digital, LCC, Sainsbury’s, ResearchOne) are never linked and never used together in any analysis. None of the other organisations have access to the NHS Digital data.

Objective for processing

The University of Leeds are running a project called QuantiCode, to develop novel data mining & visualization tools/techniques that will transform people’s ability to analyse event sequence data such as electronic health records, social care data and retail data. These data often exist in databases that contain millions of records, each comprising hundreds of numerical and categorical variables. This project was funded by the Engineering and Physical Sciences Research Council (EPSRC; Mar 2016 – Feb 2020), and is being continued with the support of funding from the Alan Turing Institute (ATI; May 2019 – Dec 2021) and the University of Leeds (ongoing).

QuantiCode project involves seven external organisations (Bradford Teaching Hospitals NHS Foundation Trust, Consumerdata Limited, NHS Digital, Sainsbury’s Supermarket Limited, Leeds City Council, Leeds North Clinical Commissioning Group, AQL Limited), who are interested in the research, contribute through meetings, identifying analysis use cases and requirements for the tools/techniques, and receive reports and copies of the tools as outputs of the project. Three of the organisations (NHS Digital, Sainsbury’s, Leeds City Council) provided datasets, which are neither linked in any way nor used together in any analysis. None of the organisations have access to each other’s data. No record-level data from NHS Digital will be shared, or be in any way accessible, to third party organisations. Similarly, no aggregated data including small numbers (as defined in the HES Analysis Guide) may be shared with (or be in any way accessible to) third party organisations.

Under GDPR article 6(1), the lawful basis for the processing is “public task” (research is a task in the public interest), and under article 9(2) the lawful basis for processing special category data is (j) Archiving, research and statistics (with a basis in law). Like most British universities, the University of Leeds has charitable status because its primary purposes of advancing education and research are deemed to deliver a public benefit. The QuantiCode project is conducting fundamental academic research to acquire new knowledge about data mining & visualization techniques for analysing event sequence data. To help ensure that this research generalises to the scale and complexity of the real world, QuantiCode has been provided with health (from NHS Digital), social care (Leeds City Council) and retail data extracts (Sainsbury’s). As is clear from the Yielded Benefits (see 5d(iii) below), the project’s applied focus has been on health and the vast majority of the tangible benefits to date are in that domain. Those benefits include identifying an issue that affects the NHS’s Payment by Results system for hospitals, showing how the project's tools can be used to identify the source of data quality problems so hospitals are able to improve the future quality of their own data, and revealing previously unknown issues in the primary and secondary care data extracts that researchers were using to study melanoma/diabetes survival and acute myocardial infarction (heart attack) outcomes/hospitalization rates, respectively.

QuantiCode is multi-disciplinary, with academic researchers from computer science, maths, ethics, geography, and health. The NHS Digital data is only used in the part of the project that is developing novel tools/techniques for investigating data quality. Data quality is an issue that is particularly important for NHS Digital – it impacts directly on areas such as running the NHS (invoice validation, etc.), as well as some indirect benefits (research and analysis to develop policy/clinical guidelines). Specifically, QuantiCode’s objectives are to develop tools/techniques to: (1) comprehensively investigate data quality (number of missing values, percentiles, outliers, etc.), (2) make detailed investigations of patterns of missing data and, where possible, to identify factors associated with their source, and (3) investigate factors that affect diagnostic persistence (the most recent NHS Digital report shows that this affects 42% of HES data episodes for patients with conditions that do not remit https://digital.nhs.uk/data-and-information/data-tools-and-services/data-services/data-quality). By collaborating with the QuantiCode project, it is expected that there will be benefits for NHS Digital and to third parties who use NHS Digital data (e.g., see 5d (iii) Yielded Benefits, below).

To perform this research into data quality, QuantiCode requested pseudonymised admitted patient care (APC) data from hospital episode statistics (HES). A single year of HES Admitted Patient Care data was being requested in order to fulfil these objectives. The factors in (2) and (3) are currently unknown and may involve any variable in a dataset, which is why all variables (except those deemed sensitive or identifiable) were requested.

To address the GDPR principle of data minimisation the University of Leeds has only requested a single year of pseudonymised data, this the minimum amount of data needed in order to still be able to achieve the aims stated within this agreement. Data concerning individuals admitted to any English NHS hospital in a whole year is required for the purposes listed in (1), (2) and (3) to ensure that the tools/techniques scale to the volume and complexity of real NHS data. Additionally, while some data issues are common (3) others such of those in (2) can be rare, so potentially missing if only a part-year dataset, a subset of APC variables or limited geography was requested.

The University of Leeds is the sole data controller who also processes the data for the purposes described within this Agreement

Expected output

The intended outputs relating to health were:

1) Data analysis tool Version 1:This output will be a visual analytics tool, which allows users to gain an overview of missing data patterns and investigate data integrity in health datasets.

2) Research report 1: This report will describe the application of the tool to health data, and the benefits that the tool provides. The report will be submitted to a high-impact outlet such as the Journal of the American Medical Informatics Association (the pre-eminent journal for research into methods for analysing health data).

3) Data analysis tool Version 2:. This version of the visual analytics tool will allow users to investigate bias caused by data quality issues in health datasets.

4) Research report 2: This report will describe the application of the tool for bias investigations, and the benefits that the tool provides. The report will be submitted to the Journal of the American Medical Informatics Association.

Outputs 1 & 3 will not contain any data – a user will load their dataset into the tool to analyse ‘missingness’. Outputs 2 & 4 will contain only aggregate level data with small numbers suppressed in line with HES analysis guide. That data will be shown in figures that illustrate the usage of the tool.

The ultimate beneficiaries of this work will be the general public. For example, CCGs and local authorities analyse data to generate business intelligence for operations and investment, with the aim of providing us all with improved and more cost-effective services. Businesses similarly require business intelligence for operations and investment, which translates to jobs and other economic benefits. To bridge the gap between these indirect benefits from the project and popular interest in big data, the University of Leeds will conduct a range of public engagement activities which include live demonstrations at the annual Leeds Festival of Science, a short film, an on-line tutorial about the ethics surrounding data analytics, and publishing articles in the popular scientific press.

As of September 2020 the following outputs have been produced, or are expected to be produced:

Note: all outputs either did not contain any NHS Digital data or were aggregated with small numbers supressed in line with HES Analysis Guidance.

Software (see Outputs 1 & 3, above)

• A data analysis tool called ACE. Version 1 was released to the QuantiCode project partners on 20/7/18. Version 2 was released to the QuantiCode project partners on 6/11/19. This version of ACE contained substantial new functionality to provide set visualization functionality that is generic, so that sets of both missing and present data can be analysed. Some parts of the ACE were re-engineered to speed up processing and make it more scalable. As part of that, ACE was tested on a 64 GB desktop PC with datasets that contained up to 21 million records, 417 fields (i.e., sets) and 141,000 unique combinations of field (set intersections).

• A Python/Pandas library (“QCprofiling”) for computing a suite of data quality checks and visualizing the output in matrices of miniature visualizations has been developed. The checks include data type, missing values, unique values, value lengths, character pattern, percentiles, outliers and example values.

Peer reviewed academic papers (*indicates outputs that concern “Non-health benefits” outlined in 5d; all other outputs are directly related to the NHS Digital data request)

• Research reports (see Outputs 2 & 4, above): These “reports” are in fact academic papers. The first journal manuscript was submitted to the IEEE Transactions on Visualization and Computer Graphics on 24/7/2018, and unfortunately not accepted. The second was submitted to the Journal of the American Medical Informatics Association on 11/6/2020 and also not accepted. Revisions and resubmissions are planned, and in both cases further work is required (see 5a. Objective for processing, and 5b. Processing activities).

• *Adnan, M., & Ruddle, R.A. (2018). A set-based visual analytics approach to analyze retail data. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA).

• Adnan, M., Nguyen, P. H., Ruddle, R. A. & Turkay, C. (2019). Visual analytics of event data using multiple mining methods. Proceedings of the EuroVis Workshop on Visual Analytics (EuroVA).

• Ruddle, R. A., & Hall, M. S. (2019). Using miniature visualizations of descriptive statistics to investigate the quality of electronic health records. Proceedings of the 12th International Joint Conference on Biomedical Engineering Systems and Technologies - Volume 5 HEALTHINF: HEALTHINF, ISBN 978-989-758-353-7, pages 230-238.

Reports (*indicates outputs that concern “Non-health benefits”)

• Feedback to NHS Digital, correcting an error in the APC Data Dictionary (ref: NIC-309404-B7V2J; 20/6/2019)

• “QuantiCode: Benefits to Health and/or Social Care” report for NHS Digital (17/7/2019)

• *Understanding customers’ missions from the products they purchase. Technical report for Sainsbury’s. June 2019.

Presentations (*indicates outputs that concern “Non-health benefits”)

• *Briefing to Leeds City Council (LCC) at Adult Social Care Information Management and Technology (IM&T) Team Meeting (8/11/16)

• “Exploring missingness patterns in Admitted Patient Care (APC) data using ACE” to NHS Digital Head of Information Utilisation and colleagues (16/10/2017)

• Demo of ACE to NHS Digital and NHS England staff, with discussion of application to issues such as diagnostic persistence (14/11/2017)

• Demo of ACE to NHS Digital Data Quality and Casemix teams (1/12/2017)

• “Visualizing data profiles and analysis pipelines”. Presentation at Visualization for Data Science and AI. Alan Turing Institute, London, 13/9/2019.

• *Discussion of strategies for investigating data quality for LCC’s adult social care business intelligence dashboard (29/10/19)

Keynote conference presentation

• “Visualizing health data – from fundamental research to successful applications”. 13th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC), Malta, 26/2/2020

Public engagement

• Exhibition and hand-on demo of prototype of ACE at the University of Leeds Be Curious research festival (25/3/17). 847 adults and children attended.

• Release a film about “Visualizing the Quality of Data” on YouTube https://tinyurl.com/VizDataQuality (20/9/2019)

Other

• Training workshops for ACE were run in July and December 2018

The following new outputs are expected:

Note: all outputs all outputs either did not contain any NHS Digital data or will be aggregated with small numbers supressed in line with HES Analysis Guidance.

Software

• The ACE software will be made widely available for free non-commercial use, initially leveraging the existing Alan Turing Institute Health Programme (includes researchers from 13 partner universities) and the HDRUK Health Data Research Hubs (includes NHS-accredited cloud-based IT platforms). This involves publicity (Q4 2020 onward) and follow-up support of interested researchers, providing additional evaluation material that is needed for peer reviewed academic papers (see below)

• Release version 1 of the RFviz software (Q4 2021) for free use by the QuantiCode project partners and free non-commercial use by others. This software will allow a data-driven approach to be taken for the investigation of diagnostic persistence and similar data quality issues.

• New workflow software, to make the QCprofiling library more accessible to users for comprehensively investigating data quality and developing new business rules for automated data profiling (Q2 2022).

Peer reviewed academic papers

• Technical paper describing the ACE tool. For submission to a high-impact outlet such as Information Visualization (Q3 2021).

• Paper describing the application of ACE to investigating electronic health records. For submission to a high-impact outlet such as the Journal of the American Medical Informatics Association (Q1 2022).

• Paper describing the design and evaluation of RFViz to the investigation of diagnostic persistence. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q2 2022).

• Paper describing the application of the new workflow software. For submission to a high-impact outlet such as the IEEE Transactions on Visualization and Computer Graphics or the Journal of the American Medical Informatics Association (Q4 2022).

Reports & presentations

• Six-monthly reports/presentations to NHS Digital Data Quality team (similar to previous presentations; see above) to stimulate adoption.

Public engagement

• A demonstration of the QuantiCode software has been accepted for the AI UK showcase (23rd/24th Mar 2021; postponed from 2020 due to Covid-19 https://www.turing.ac.uk/ai-uk).

Benefits reported

To date there have been three main types of benefit. In the following, the eight original Expected Measurable Benefits (see 5d(ii)) are referenced as [1], [2], etc.

Benefits to NHS Digital and hospitals themselves: Our ACE tool revealed unexpected gaps in sets of fields for some Admitted Patient Care (APC) data records, and pinpointed the source of 90+% of the gaps as being a particular admission method at a single hospital [1]. Some of the gaps were in the diagnosis code fields (DIAG_nn) and others were in the operation fields (OPERTN_nn). The latter are of particular concern to NHS Digital because they affect the process that the Casemix Team use to calculate Healthcare Resource Groups (HRGs), which are one of the building blocks of the NHS’s Payment by Results system for hospitals [6]. These findings from ACE allow NHS Digital to provide feedback to the provider to rectify the problem, via the existing HES data quality lifecycle. ACE has subsequently revealed similar gaps in the date fields for operations (MYOPDATE_nn) and inconsistencies between the operation fields and their corresponding date fields. The speed (a few minutes) with which users were able to interactively find these unexpected missing data issues with ACE, contrasts with the years that the issues have existed yet gone undetected. It would be straightforward to implement new business rules to automatically check for these gaps [4], to complement ACE’s power in exploratory analysis for revealing new issues. The results would be improved data quality for NHS Digital’s own analysis [7], and in data and/or data quality notes provided to third parties [8].

Benefits to third party users of electronic health records: QuantiCode has concentrated its effort on assisting researchers working on two other projects – one using APC data (DARS-NIC-17649-G0X4B) to study long-term outcomes and hospitalization rates for survivors of acute myocardial infarction (heart attack), and the other using ResearchOne primary care data to study the survival from melanoma of patients with type 2 diabetes.

ACE revealed two main issues with the APC data. One concerned missing admission and episode start dates. In principle they can be estimated from each other if only one is missing, but ACE showed that that was hardly ever the case (99.8% of records that were missing one of those dates was also missing the other). However, ACE also let the researchers localize the origin to maternity records, and therefore make an informed decision to discard the records because the project was conducting heart attack research [8]. The other issue concerned survival time, which was missing in 12 million records and led the researchers to comment that they had asked NHS Digital to derive the survival time for all patients based on the study census date, regardless of their mortality status, but clearly that had not been done. This finding made the researchers realise that they had to estimate survival time for those patients who did not die prior to the census date, losing precision and precluding analyses of short-term survival patterns within 30 days of admission. In future, ACE would allow the researchers to check their data on receipt, so that errors can be rectified (NHS Digital’s time-limit is one month from receipt) [8].

Complementing ACE, the QCprofiling software revealed a number of longitudinal data quality issues [5] across nine years of APC data extracts (150 million records in total; 116 different fields). Some were the result of policy change (e.g., all values of Clinical Commissioning Group fields missing in the early extracts but not the latest, and longitudinal differences in coding depth for the operation fields) [2], others from improvements to business rules (the primary diagnosis was sometimes missing in the early extracts but never from the latest ones), some were prevalent across all of the extracts (e.g., widespread use of punctuation characters in diagnosis codes, which complicates data cleaning), one affected data linkage [6] (the use of two different (16 vs. 32 character) pseudonymized patient identifiers in one of the extracts).

The QCprofiling software also revealed data quality issues with the primary care data (90 million records; 14 tables; 41 different fields). These included validation error text appearing in a birth year field, clinical codes being padded out with ‘.’ characters (i.e., inconsistent coding precision between different data tables), and the distribution of blood pressure measurement dates leading the researchers to realise that they had not been provided with the historical data that they had expected [8]. All of these issues were new to the researchers, despite the fact that they had had the data for several months. As with the APC data, the QCprofiling software was a game-changer for investigating data quality because it allowed the researchers to see patterns that spanned a large number of records/tables/fields at the glance of an eye.

Non-health benefits: ACE was also evaluated and used for more exploratory analysis with data from two of QuantiCode’s other external partners (Leeds City Council (LCC) and Sainsbury’s), to help generalise the design of ACE. LCC were shown how ACE’s set visualization techniques could be used to investigate patterns of missing data in an anonymized adult social care service provision dataset. The main benefit lies in making explicit the patterns for discussion by analysts and stakeholders prior to use of the data for other purposes such as modelling future service demand. In some cases explanations could be offered for unexpected patterns (e.g., tacit knowledge about historical changes in which data were recorded), and in others further investigation would be required.

Sainsbury’s were shown how frequent itemset mining techniques could be combined with ACE’s set visualization to analyse customer transactions (e.g., What items are bought together? How to highlight differences and similarities between stores, or over time?). This was exploratory research, indicating a possible approach that would allow customer transactions to be analysed at a much finer level of detail than is possible with Sainsbury’s current methods.

Important note: As stipulated in the project’s data agreements, the datasets from the different organisations (NHS Digital, LCC, Sainsbury’s, ResearchOne) are never linked and never used together in any analysis. None of the other organisations have access to the NHS Digital data.

DARS-NIC-49164-R3G5K-v0.4 1 October 2017 to 30 September 2020
Title
QuantiCode: Admitted Patient Care Data
Commercial
Yes
Sublicensing
No
Datasets
1
Files released
1

Datasets: Hospital Episode Statistics Admitted Patient Care (HES APC)

Objective for processing

The University of Leeds are running a project called QuantiCode. This project is funded by the Engineering and Physical Sciences Research Council (EPSRC) from March 2016 to February 2019 (details: http://gtr.rcuk.ac.uk/projects?ref=EP%2FN013980%2F1). The overall QuantiCode project is divided into 3 stages:

1) Data fusion, covering tools for data linkage and visualizing data quality, and thought leadership in data governance.

2) Analytical techniques that allow users to interactively mine longitudinal data for patterns.

3) Governance-aware abstraction techniques that allow users to explore how complex data may (or may not) be simplified to reveal important patterns.

The work will be evaluated by the collaborating organisations to ensure that the solutions are applicable in the real world.

The QuantiCode project involves a number of collaborating organisations (University of Leeds, Bradford Teaching Hospitals NHS Foundation Trust, Consumerdata Limited, NHS Digital, Sainsbury’s Supermarket Limited, Leeds City Council, Leeds North Clinical Commissioning Group, AQ Limited), who are interested in this work and are supplying datasets and will receive reports and tools as outputs of the project. NHS Digital is one of these organisations, as NHS Digital has a number of very large datasets, and Data Quality is an issue which is particularly important – it impacts directly on areas such as running the NHS (invoice validation, etc.), as well as some indirect benefits (research and analysis to develop policy/clinical guidelines). By collaborating with the QuantiCode project, it is expected that there will be benefits for NHS Digital and benefits to healthcare more widely as a result.

Any data provided from NHS Digital to the University of Leeds for this project will not be linked in any way with the other datasets being used. No record-level data from NHS Digital will be shared, or be in any way accessible, to third party organisations (including the other collaborating organisations). Similarly, no aggregated data including small numbers (as defined in the HES Analysis Guide) may be shared with (or be in any way accessible to) third party organisations

Within the bounds of the QuantiCode project, the purpose for processing healthcare data from NHS Digital is to allow different designs of visualisation and machine learning technique to be compared for their ability to meet user requirements, and allow the QuantiCode data analysis tool to be tested prior to release to NHS Digital for in-depth evaluation. A single year of HES Admitted Patient Care data is being requested in order to fulfil these objectives. The single year of hospital data is important as there are data quality process which take place at the end of the year, meaning that a sub-set of data (e.g. 6 months)will display different characteristics to the finalised (“Annual Refresh”) data produced after the end of the financial year. Without healthcare data, it is possible that the methodologies/tools cannot be applied to healthcare data, and the opportunity to improve the healthcare data will be lost.

In addition, a specific purpose for processing the health data is to make a detailed analysis of patterns of “missingness” in data (the manner in which data are missing from a sample of a population). The goal is to be able to investigate patterns that involve: (a) several (3+) variables missing together, and/or (b) are dependent on the particular value of other variables (e.g., provider code and admission type). These patterns are currently unknown and such missingness may involve any variable in a dataset, which is why all variables (except those deemed sensitive or identifiable) have been requested.

Expected output

The intended outputs relating to health are:

1) Data analysis tool Version 1 (target date 28/2/18). This output will be a visual analytics tool, which allows users to gain an overview of missing data patterns and investigate data integrity in health datasets.

2) Research report 1 (target date 30/9/18). This report will describe the application of the tool to health data, and the benefits that the tool provides. The report will be submitted to a high-impact outlet such as the Journal of the American Medical Informatics Association (the pre-eminent journal for research into methods for analysing health data).

3) Data analysis tool Version 2 (target date 31/5/19). This version of the visual analytics tool will allow users to investigate bias caused by data quality issues in health datasets.

4) Research report 2 (target date 30/11/19). This report will describe the application of the tool for bias investigations, and the benefits that the tool provides. The report will be submitted to the Journal of the American Medical Informatics Association.

Outputs 1 & 3 will not contain any data – a user will load their dataset into the tool to analyse missingness. Outputs 2 & 4 will contain only aggregate level data with small numbers suppressed in line with HES analysis guide. That data will be shown in figures that illustrate the usage of the tool.

The ultimate beneficiaries of this work will be the general public. For example, CCGs and local authorities analyse data to generate business intelligence for operations and investment, with the aim of providing us all with improved and more cost-effective services. Businesses similarly require business intelligence for operations and investment, which translates to jobs and other economic benefits. To bridge the gap between these indirect benefits from the project and popular interest in big data, the University of Leeds will conduct a range of public engagement activities which include live demonstrations at the annual Leeds Festival of Science, a short film, an on-line tutorial about the ethics surrounding data analytics, and publishing articles in the popular scientific press.

Benefits reported

Yielded Benefits is not a requirement for new applications.

Register history

When this agreement appeared in, or was edited in, each monthly edition of the register. Built by comparing every edition this site holds, the earliest of which is July 2021.

Cite this page

NHS England (2026) Data Uses Register, September 2026 edition, agreement DARS-NIC-49164-R3G5K, “QuantiCode: Admitted Patient Care Data”. Read via NHS Data Access Explorer (unofficial), https://healthdatauses.uk/agreements/dars-nic-49164-r3g5k/ (accessed [date]).

This address stays the same, but the page is rebuilt with each monthly edition, so the citation names the edition it shows. Every edition's data is kept in the facts store.

Source: datausesregister_september2026.xlsx, September 2026 edition of the NHS England Data Uses Register. Search that workbook for DARS-NIC-49164-R3G5K to see the original rows.