Creating synthetic data for health research
University College London (UCL) · Academic
Expired The latest version ended on 28 March 2024. The September 2026 register still lists the agreement, but its term has passed.
- Reference
- DARS-NIC-419453-G3G1G
- Latest version
- v0.4
- Term of latest version
- 29 March 2021 to 28 March 2024
- Start date
- 29 March 2021
- Data controller
- Sole Data Controller
- Commercial purposes
- No
- Sublicensing
- No
- Files released to date
- 0
Why the data was released
Objective for processing
The study aims to evaluate methods for creating ‘synthetic’ datasets for health research. The idea is to create artificial datasets that ‘look’ like the original data source (preserving the structure and statistical properties of the data and relationships between variables) but that do not contain information on any real individuals, and therefore pose no confidentiality risks. If such datasets can be created, synthetic data could be used by researchers to understand the structure of the data, develop data cleaning protocols, codes and algorithms, and test out methods. Final analyses could then be conducted once approvals are in place (in secure settings), or alternatively, by the data providers themselves (so that researchers would not need any access to confidential information).
The concept of synthetic data is not new, but its use is increasing. For example, synthetic versions of general practice data (from the Clinical Practice Research Datalink) have recently been generated, including for COVID-19 research (https://www.cprd.com/content/synthetic-data). Health Data Research UK have also recently prioritised work in this area (for more information, see https://www.hdruk.ac.uk/synthetic-data/).
To achieve this aim, the research team at University College London (UCL) will compare a range of methods for synthetic data generation for generating synthetic versions of Hospital Episode Statistics (HES) in particular Admitted Patient Care (APC).
Legal basis
The legal basis for processing personal data for this purpose data at UCL falls under Article 6(1)(e) of the General Data Protection Regulations (GDPR), i.e. “a task carried out in the public interest”. It also falls under Article 9(2)(j), “processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes”.
There is clear public interest for this application, as it could lead to a significant streamlining of research using electronic health data. The potential for using electronic health data for timely health research has recently been highlighted by COVID-19. Strict governance restrictions to protect confidentiality mean that data, when released, are highly anonymised and/or only accessed in secure research settings. This is for good reason - there are some who argue that individual-level data can never be truly anonymous. However, methods to anonymise data, such as by removing exact event dates or categorising variables, can mean that the resulting data are not sufficiently granular for research purposes. Synthetic datasets could provide an alternative resource for health researchers.
Findings from the study will help data providers decide whether providing synthetic versions of electronic health data could help address the increasing pressure to deliver timely outputs in the context of increasing numbers (and complexity) of data access applications. Sensitive or potentially identifiable datasets such as Hospital Episode Statistics have great potential for economic and social impact, leading to better informed policy decisions and effective public services. Widening the use of these data through synthetic data would therefore lead to increased efficiency in health research, ultimately benefiting the public.
The researchers will access pseudonymised data only (as only pseudonymised data has been released under NIC-393510).
How the data requested will achieve the aim
The HES data requested will achieve the aim of this study by providing an ‘original’ data source that can be used as a basis for evaluating a range of synthetic data generators. For example, the UCL team will select a core set of variables to be synthesized including patient characteristics (index of multiple deprivation score (IMD), ethnic group, sex), birth outcomes, number of admissions, and high-dimensional fields such as diagnosis fields.
Exploring whether it is possible to accurately generate ‘lookalike’ variables in a synthetic dataset will help inform researchers and data providers on the value of synthetic data. Information from HES will also allow the researchers to determine for which types of data the methods are effective. For example, they will be able to determine whether the methods can be used to generate synthetic versions of early HES data (from 1997, likely to be less complete) and more recent data (from 2019, more complete).
The data will be used to evaluate three synthetic data generators, Synthpop, Simulacrum and Jomo. Synthpop and Jomo are implemented in open source software (R packages). Simulacrum was developed by Health Data Insight (a social enterprise overseen by the Office of the Regulator of Community Interest Companies). UCL will evaluate these generators, in terms of how well they can create synthetic datasets, by the following:
- Assess general utility by visualising marginal distributions of key variables and by estimating the standardised propensity mean square error (pMSE). The pMSE is a measure of poor discrimination between the original and the synthetic data (a positive feature in this context) and is derived from a logistic regression model for the propensity of a record to be from the original dataset. Coefficients of the propensity model will be inspected to identify ways in which the synthetic and original data diverge.
- Assess specific utility by estimating coefficients from selected models of interest using the synthetic and the real data, and then deriving standardised differences and percentage bias for the coefficients of interest. UCL will assess the extent to which inferences based on the synthetic data are robust, by estimating the overlap of confidence intervals for coefficients derived from the original and synthetic datasets using the interval overlap measure. Results will be averaged across multiple versions of the synthetic data.
Relevant background to the request
This is a methodological research study funded by the Economic and Social research Council (ESRC).
Relationship between proposed project and associated work
UCL are proposing to re-use an existing extract of HES APC data (held by members of the wider research team under a separate DSA; NIC-393510). UCL are requesting a re-purposing of this DSA to allow them to access specifically the years 1997/98 and 2019/20, which will enable them to establish whether methods can be used to generate synthetic versions of both early HES APC data and more recent, more complete data.
The purpose of the request
The research team are requesting that health data captured in HES are used to evaluate whether it is possible to generate synthetic versions of health data that can be used for health research. The purpose of this request is to answer a set of research questions about the feasibility and usability of synthetic data, aiming to generate evidence on the usefulness of synthetic data for data providers and health researchers.
For this purpose, the research team are requesting pseudonymised HES APC data for 1997/98 and 2019/20. National data are required in order to capture the variation in data quality across different providers and to evaluate whether synthetic data methods can handle large amounts of data accurately. There are no alternative ways of achieving the purpose of this application. The research team will use the minimum data required in order to answer the research questions.
Organisations involved
Data controller: UCL
Data processor: UCL
Although the study involves a co-applicant at the London School of Hygiene and Tropical Medicine (LSHTM), they will only contribute advice (particularly on the use of the Jomo package). They will not access any HES data. LSHTM is not considered as a Data Controller or Data Processor.
There are no funders or commissioners directly involved in the project.
No party involved in the application will receive any form of commercial benefit from the use of data.
Processing activities
The research team at UCL will extract a core set of variables including patient characteristics (IMD, ethnic group, sex) birth outcomes, number of admissions and high dimensional fields such as diagnosis fields, from an existing HES APC extract held under a separate Data Sharing Agreement (NIC-393510). The intention is to use this original data set and replace all of the values with synthetic ones, causing minimal distortion of the statistical information contained in the original data set. This would result in a new dataset, in which every value for every variable will be synthetic, i.e. the synthetic data will not contain any records that correspond to a real person. The new extract will be transferred to a new location in the UCL Data Safe Haven. All analysis will take place within the UCL Data Safe Haven.
The data flows are summarised as follows:
1. HES APC data with no identifiers will be extracted for 1997/98 and 2019/20 from NIC-393510-J1Q6T
2. The new extract will be transferred to a new location on the UCL Data Safe Haven.
3. All analysis will take place in the UCL Data Safe Haven.
There will be no attempts to identify individuals. Risk of re-identification will be mitigated by checking all outputs for small cell sizes. No potentially disclosive outputs will be shared or published. Data processing will only be carried out by substantive employees of UCL who have been appropriately trained in data protection and confidentiality.
Expected output
The main output will consist of a set of guidelines and evidence on the use of synthetic data, including a comparison of approaches. These guidelines will be developed alongside data providers and other researchers as part of an engagement phase of the study. UCL will disseminate the guidelines using existing networks, e.g. including with colleagues at the Office for National Statistics (ONS), the Department for Education (DfE), NHS Digital and Public Health England (PHE), as well as Administrative Dara Research UK (ADR UK) and Health Data Research UK. Findings will be used by data providers to inform ongoing research into the use of synthetic data. Findings will also be published as peer review publications in high quality journals (e.g. PLoS One, submitting within 3 years of data access). The researchers will also work with members of the public to co-produce a range of outputs suitable for communicating results to members of the public interested in the use of electronic health data for research, e.g. through training events or fact sheets.
All journal articles will be published with open access, to ensure the wide dissemination of the study’s results to data providers, healthcare professionals, governance bodies, and other researchers. Results of the study will also be made available in both clinical and methodological research forums: abstracts will be submitted to the following conferences within 2 years of data access: International Population Data Linkage Network, Public Health Science.
As the main output will be guidelines on the use of synthetic data, and will not inform any decisions about individuals, the researchers do not expect that an a EQIA will be required. However, the guidelines will include as assessment of how well synthetic data can preserve information in the real data about people with protected characteristics (e,g., to ensure that ethnic groups are represented in the same way within the synthetic and real data).
Data will not be used for sales or marketing purposes.
Expected measurable benefits
The research will benefit the provision of health care and the promotion of health, by informing policy on whether synthetic versions of electronic healthcare datasets can be shared with researchers, in order to streamline the research process. This will have direct relevance to all health research using healthcare datasets such as HES. The research is in the public interest, because the public have vocalised opinions about the need for timely access to high quality healthcare data, especially in light of COVID-19. The results of this study will provide evidence on whether synthetic data can be used to speed up the data access applications, data management, and data analysis stages of research. The study will directly benefit the Health and Social Care sector by providing data providers, researchers, governance bodies and policy makers with detailed and up-to-date evidence to aid decision making about the use of synthetic data to support a wide range of health research. Our guidelines on the appropriate use of synthetic data will help facilitate timely access to administrative datasets, improve the efficiency of research and streamline approval processes in the context of increasing demands on data providers. Our work will help minimise access to identifiable personal data, by allowing researchers to develop methods using anonymised, synthetic data, with final models being implemented by data providers or within secure settings.
The study team will engage with data providers, researchers and the public throughout the study. UCL will offer workshops on synthetic data and the results of our study with NHS Digital, DfE and ONS. This will allow the study team and data providers to establish views on the resource implications associated with synthetic datasets, and the likely efficiency benefits of being able to provide synthetic datasets to researchers.
Benefits reported so far
Yielded Benefits is not a requirement for new applications.
Datasets on the latest version
Legal basis for provision: Health and Social Care Act 2012 - s261 - 'Other dissemination of information'
| Dataset | Type of data | Sensitivity | Frequency | Confidential data |
|---|---|---|---|---|
| Hospital Episode Statistics Admitted Patient Care (HES APC) | Anonymised - ICO Code Compliant | Non-Sensitive | One-Off | Does not include the flow of confidential data |
Files released
Files released counts only files released externally by DARS. Access granted in NHS England's own systems, such as its Secure Data Environment, is not included.
No files recorded as released under this agreement.
Version history
The register lists each renewal of this agreement as a separate row. This site has 1 version.
DARS-NIC-419453-G3G1G-v0.4 29 March 2021 to 28 March 2024
- Title
- Creating synthetic data for health research
- Commercial
- No
- Sublicensing
- No
- Datasets
- 1
- Files released
- 0
Datasets: Hospital Episode Statistics Admitted Patient Care (HES APC)
Register history
When this agreement appeared in, or was edited in, each monthly edition of the register. Built by comparing every edition this site holds, the earliest of which is July 2021.
-
July 2021 —
already listed in the earliest edition this site holds, so it may be older. 1 version: DARS-NIC-419453-G3G1G-v0.4
Cite this page
NHS England (2026) Data Uses Register, September 2026 edition, agreement DARS-NIC-419453-G3G1G, “Creating synthetic data for health research”. Read via NHS Data Access Explorer (unofficial), https://healthdatauses.uk/agreements/dars-nic-419453-g3g1g/ (accessed [date]).
This address stays the same, but the page is rebuilt with each monthly edition, so the citation names the edition it shows. Every edition's data is kept in the facts store.
Source: datausesregister_september2026.xlsx, September 2026 edition of the NHS England Data Uses Register. Search that workbook for DARS-NIC-419453-G3G1G to see the original rows.