Unofficial. This site is an experimental reformatting of data published by NHS England. It is not endorsed by NHS England. Always check the official Data Uses Register before relying on anything here.

Genomics England: Use of data within the National Genomics Research Library (NGRL)

Genomics England · Research

In term In term in the September 2026 edition: the latest version runs to 2 April 2027.

Reference
DARS-NIC-12784-R8W7V
Current version
v15.2
Term of current version
3 April 2024 to 2 April 2027
Start date
Before 1 February 2019
Data controller
Sole Data Controller
Commercial purposes
Yes
Sublicensing
Yes
Files released to date
2,853

Why the data was released

Objective for processing

Genomics England requires access to NHS England Data for use in their National Genomics Research Library (NGRL), which operates as a Trusted Research Environment (TRE).

National Genomic Research Library (NGRL)

The NGRL is a secure national resource of genomic, health and sample data managed by Genomics England, which builds on the research environment created by Genomics England for the 100,000 Genomes Project that completed recruitment in 2018 and involved the sequencing of approximately 100,000 genomes. It contains cohorts of patients/participants recruited via programmes set up for the NHS Genomic Medicine Service (GMS), the 100,000 Genomes Project, and other studies/programmes where patients’/participants’ genomes have been sequenced, and provides a national standardised genomic research resource. Being able to compare all patient data in one place provides researchers with an opportunity to better understand diseases, develop new treatments and can lead to new discoveries.

The NGRL contains NHS England Data linked at record-level with data on the 100,000 Genomes Project cohort. Genomics England will also store NHS England Data linked at record-level with data on other Genomics England programmes.

This Data Sharing Agreement (DSA) covers Data provided for the 100,000 Genomes Projects cohort and the NHS Genomic Medicine Service (NHS GMS) cohorts.

NHS Genomic Medicine Service (NHS GMS)

Following the successfully delivery of the 100,000 Genomes Project, Genomics England begun work to deliver the NHS GMS. This service will offer patients the dual opportunity of routine clinical care alongside a choice to participate in research spanning across all genomic tests within the NHS, starting with whole genomes. Development of the NHS GMS builds on the evidence generated by the 100,000 Genomes Project, but also extends to other genomic testing other than Whole Genome Sequencing (WGS).

The NHS GMS is led and commissioned by NHS England and will consist of:

> NHS Genomic Laboratory Hubs (GLHs) that will work as part of a National Genomic testing service. The provisions in this service will be determined by a national genomic test directory that outlines the testing strategies and technology to be employed for rare and inherited disease, cancer and other defined conditions/ applications.

> Clinical Genetic Services

> Cancer services using genomic analysis to guide treatment

Within the GMS, two particular cohorts are defined:

(a) The Rare Diseases cohort

The NGRL rare diseases cohort will consist of families with rare diseases, based on family structures appropriate to the provision of the GMS. This will harness the strength of current UK rare disease programmes and will advance current understanding of rare disease mechanisms. It may also impact on common diseases that share similar phenotypes. It will also offer opportunities for biomarker, clinical, and interventional studies through industrial partnerships. Genomics England will actively support the implementation of the UK Rare Disease Strategy. Building on work already undertaken by the 100,000 Genomes Project, this will facilitate the generation of a national data resource of all genomic data, with a focus on WGS but including all genomic testing.

The goals of processing data of patients with rare disease are:

• To increase discovery of pathogenic variants (the gene variant responsible for causing disease) for rare disease.

• To add value with additional biological insights that build confidence in commonly accepted pathogenic variants.

• To enhance the clinical interpretation of WGS in rare disease.

• To develop a programme of functional pathways for genomic tests other than whole genome sequencing, specifically, transcriptomics (the study of all the ribonucleic acid (RNA) molecules within a cell, otherwise known as the transcriptome), epigenetics (the study of how cells control gene activity without changing the DNA sequence), micro RNAs and biomarkers.

• To return findings to the NHS for feedback to patients.

• To create a unique dataset for rare diseases that may enable therapeutic innovation.

(b) The Cancer cohort

The NGRL will continue to learn from and collaborate with other projects who are producing an inventory of genomic, transcriptomic and epigenomic changes in a wide range of different tumour types. Researchers from these projects access the data via an approved sublicensing agreement with Genomics England.

The goals of processing data for patients with cancer are to:

• Use WGS to identify novel driver mutations for cancer and to understand its evolutionary genetic architecture through primary and secondary malignant disease (by multiple biopsy and WGS).

• Partner stratified healthcare programmes and outcome studies with patients from the NHS in England, to enable understanding of WGS benefits in defining predictors of therapeutic response to cancer therapies.

• To use other genomic testing approaches to offer additional biological insights into cancer.

• To utilise WGS to identify new pathways for cancer therapies and improved diagnostic characterisation.

Other forms of collaboration include direct partnerships to develop tools to improve systems and services. For example, partnering with Lifebit to take advantage of their technical genomic data tooling.

This DSA also permits the reuse of Data from the GMS Rare Diseases or Cancer cohorts and/or from the 100K Genomes cohort alongside the Generations cohort data obtained under DARS-NIC-733503-V0X9Q.

The following NHS England Data will be accessed:

> Hospital Episode Statistics (HES) - these Datasets provide the core clinical data for participants and are vital to the provision of a detailed longitudinal medical history for participants. Specifically, the following HES Datasets are required:

- Admitted Patient Care (APC)

- Accident & Emergency (A&E)

- Critical Care (CC)

- Outpatients (OP)

> Emergency Care Data Set (ECDS) – necessary to understand which patients attend emergency departments and what treatment they receive in order to assess if there are associations with genetic markers.

> Diagnostic Imaging Dataset (DID) – necessary to provide invaluable, detailed information to build on participants’ phenotypes (observable characteristics), e.g., tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

> Mental Health – necessary because the GMS includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are expected to be some of the most prevalent conditions within the Project. For the 100,000 genomes project approximately 10% of project participants have a mental health record. Mental health Data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

> Civil Registration Mortality – necessary because deaths Data is essential for performing survival analyses; this is crucial information for research in combination with other medical history. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

> Cancer Registration – necessary because it provides a clinical background for the patient's cancer journey which may help provide scientific and medical insights when compared against the patient's genome.

> MRIS Events Mortality and Demographic Data - still required for open-ended research, specifically for researchers which have already been using this Data and need to go back to it. Losing this Dataset may cause disruption to their research.

NHS England data is matched to the consented cohorts in the NGRL and therefore provides a more comprehensive medical history, and going forward, a more comprehensive patient journey. NHS England is the richest source of the data required.

The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records, such as blood pressure, cholesterol, glucose, and pharmacogenomics (a field of research that studies how a person's genes affect how he or she responds to medications.), adds significant value. The richness of the NGRL datasets will allow Genomics England to move beyond the primary phenotype of the rare disease, cancer or infectious disease that led to the patient’s initial WGS in the context of other continuous traits, diseases and response to therapy including harm.

The level of the Data will be:

> Identifiable

For the following Datasets: HES APC; HES OP; HES A&E; ECDS; Cancer Registration Data; Civil Registrations of Death, many indirect identifiable data items are required because they provide valuable data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHS England. To ensure patients will not be identified the researchers and their projects are examined before access is granted to the Data and an agreement to not re-identify patients is signed. Further Genomics England monitors all Data leaving the NGRL and will not allow patient records to be exported.

The Data will be minimised as follows:

> Limited to cohorts who consented to participate in (1) the 100,000 genomes project who also consented to longitudinal research (~80,000 participants) or (2) the Genomics Medicines Service (expected ~100,000 new additions per year).

Genomics England will request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.

Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above.

The lawful basis for processing personal data under the UK GDPR is:

> Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.

Genomics England has determined the processing is necessary for its legitimate interests in carrying out medical research on the causes, diagnosis and treatment of rare diseases and cancers.

The lawful basis for processing special category data under the UK GDPR is:

> Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

It is necessary for Genomics England to process special category participant data for carrying out medical research on the causes, diagnosis and treatment of rare diseases and cancers, which is expected to benefit patients.

The funding is provided by the Department of Health and Social Care. The funding is specifically for the projects described. Funding is in place until March 2025, with the intention to renew this funding periodically.

The funder will have no ability to suppress or otherwise limit the publication of findings.

Lifebit provides IT support to Genomics England.

Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.

Representatives from patient and public bodies have an important role to play in Genomics England commercial initiatives. These representatives ensure transparency is upheld, and the interest of those whose data is being used is always being respected.

In the early stages of the Library, Genomics England undertook a range of work to ensure that potential participant’s views were included in the formulation of the ethical policies submitted for research ethics approval and in the development of patient information. The views of different groups of potential participants (those affected by cancer, rare disease, and those from BAME communities) in relation to ethical issues raised by the 100,000 Genomes Project were sought and findings were published on the Genomics England website (See all reports under ‘patient and public involvement - https://www.genomicsengland.co.uk/library-and-resources/ and the Genomics England Engagement Strategy). Genomics England will continue to engage with these stakeholders. Further to this, each of the 13 currently recruiting NHS Genomic Medicine Centres had dedicated Patient and Public Involvement leads (PPI) who are responsible for engaging with and involving local potential participant groups from diverse backgrounds. It is expected that the future NHS GMS will continue these local PPI activities to shape and inform the service.

A Participant Panel has also been established. This 30-strong group has provided invaluable advice on a range of topics, for instance, in shaping how analysis is monitored, how results are returned, and how advice and support should be framed. Participant Panel members have either donated samples to the Library themselves or are carers of participants. They take part in a wide variety of consultative groups, such as the Genomics England Ethics Advisory Committee but most importantly are guardians of the dataset, with representatives on the Access Review Committee. Participants play an important part in every decision made about access to data.

SUB-LICENCING:

Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.

It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.

The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.

All requests for data access will be subject to the following considerations:

• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).

• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).

• Provision of a signed Genomics England data access agreement to the Access Review Committee.

• Prioritisation of access according to resource availability.

• Facilitation of high-quality health research

Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:

• Diagnostic discovery

• Pre-clinical research

• Clinical Trials Referral

• Real World Evidence / Market Access

Approval process for Commercial organisations for access to the NGRL:

Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.

The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.

Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.

Approval process for GeCIP users (academic) of the NGRL:

• Researcher visits Genomics England website to enrol as a GECIP member

• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)

• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.

• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.

• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.

• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.

All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:

• Identification of diagnoses originally missed by the standardised pipeline

• Feedback of new diagnoses to patients

• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone

Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.

Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.

AIRLOCK POLICY:

The following rules are applied to all airlock requests:

1. All relevant details of the summary results to be transferred must be provided with every request.

2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.

3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.

4. Summary results requested for transfer are assessed using the following criteria:

a. whether the request aligns with the users ARC approval in full;

b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;

c. any data security implications;

d. any disclosure risks;

e. the technical feasibility and associated cost of the request;

f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;

g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.

The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:

• Technical Lead

• User Community Representative

• Bioinformatics Director

• Caldicott Guardian

• Chief Scientist representative

The Data will be processed worldwide.

Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.

GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:

• UK academic research institutions (e.g., universities, research institutions etc.)

• NHS trusts or authorities

• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project

• Foreign universities and research institutions that carry out significant research activity

• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))

• Foreign healthcare organisations (private or public) that undertake significant research activity

To be eligible for data access as a GeCIP member, applicants must meet these requirements:

• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.

• Their host institution has verified that they are affiliated with that institution.

• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).

• The GeCIP domain lead has approved the application.

• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.

• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.

All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).

The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.

GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.

All personnel accessing the Data have been appropriately trained in data protection and confidentiality.

The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:

> National Cancer Registration and Analysis Service (NCRAS)

and uncurated NCRAS data

> Secure Anonymised Information Linkage (SAIL) data; Welsh data

> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)

> NHS Trusts data

The Data will not be linked with any other data.

The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.

To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.

Data accessed under sub licence will not be re-identified.

Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.

Data shared through the Airlock process is aggregate data only and is therefore not personal data so does not require a legal basis under the UK GDPR.

A release register detailing any sub licences and onward sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse

Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract.

In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users.

NHS England will require the ability to audit the sub licensee.

Processing activities

Genomics England will transfer data to NHS England. The data will consist of identifying details (specifically study ID, NHS Number, Date of Birth, Surname, Forename, Gender, Postcode and Other Given Name) which are required for the cohort to be linked with NHS England Data. This is the minimum requirement of identifiers required for linkage to guarantee complete matching.

NHS England Data will provide the relevant records from the HES APC, HES CC, HES OP, HES A&E, ECDS, MHSDS, MHMDS, Cancer Registration Data, Civil Registrations of Death, PROM and DIDs datasets to Genomics England’s Amazon Web Services (AWS) cloud storage. The Data will:

> Contain directly identifying Data items including but not limited to: Names, Postcode, Cause of Deaths, Place of Birth, Cancer Registration Number, which are required to provide maximum insight, and therefore maximum value to the researchers accessing the Data.

The NHS England Data is pseudonymised within the AWS cloud and is then loaded into the NGRL. Raw, identifiable files are kept in a secure location on AWS.

The Data will not be transferred to any other location.

The Data will be stored on the NGRL and the AWS Cloud at Genomics England.

Genomics England stores NGRL data on the Cloud provided by Amazon Web Services (AWS).

The Data will be accessed by authorised personnel via remote access.

The Controller(s) must confirm and provide evidence upon audit by NHS England that access via any remote device complies with the data security obligations within this DSA and the Data Sharing Framework Contract.

For remote access:

- Remote access will only be from secure locations situated within the territory of use (as further restricted elsewhere within the DSA if so done) stated within this DSA;

- Access controls granting users the minimum level of access required are in place;

- Remote access is only via secure connections (e.g., VPNs or secure protocols) to protect data;

- Multifactor authentication (MFA) is required for remote access;

- Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for remote access;

- All remote access is undertaken within the scope of the organisation’s DSPT (or other security arrangements as per this DSA) and complies with the organisation’s remote access policy.

The above applies in addition to any condition set out elsewhere within the DSA (e.g. who may carry out processing, and for what purpose).

Data is physically stored in England.

Remote access is permitted from the following specified countries: UK, EEA Countries, United States, Canada, Australia, Qatar, Republic of Korea, Japan, Switzerland, Brazil, India, New Zealand, Argentina

Should any country on the permitted list above become a high risk country through the duration of this DSA, the Recipient will cease disseminating data to researchers/organisations based in that country and request that data already disseminated be destroyed.

Should the Recipient wish to share data with any countries not listed above, it will require an update to this DSA.

Should Genomics England wish to facilitate remote access from a country that is not listed above, prior written agreement from NHS England must be obtained.

Genomics England upholds the following safeguards and controls:

1. Compliance with National Cyber Security Centre (NCSC) guidance, leading to the implementation of geo-blocking measures for IP addresses originating from Iran, Russia, North Korea and Belarus

2. Collaboration with the NCSC and other security partners to identify and block potentially risky IP addresses, irrespective of their country of origin.

3. Implementation of two email authentication methods, namely Domain-based Message Authentication Reporting and Conformance (DMARC) and Sender Policy Framework (SPF), to detect and respond to spoofing and spam. This is crucial, as these activities often target our firewalls from international IP addresses.

4. Introduction of additional assurance activities related to international access within the Office 365 estate.

5. Conducting due diligence on companies associated with BGI Genomics.

6. Responsibilities of the Access Review Committee (ARC) include the thorough review of applications and applicants.

7. Continuous improvement of Information Governance training and cybersecurity awareness at Genomics England.

Access to confidential patient identifiable Data is restricted to an extremely limited number of employees of Genomics England, accessible on AWS.

Substantive employees of Genomics England and researchers who are a member of the GeCIP and Discovery Forum will process the Data for the purposes described above.

Expected output

Researchers will have their own dissemination and communication strategies, however a full list of scientific publications and conferences/posters will be made available on the Genomics England website on an ongoing basis. The expected outputs of the processing will be:

> Submissions to peer reviewed journals

> Presentations at conferences

> Posters

> Creation of a database of all genomic data, including all genomic and omics tests.

The outputs will not contain NHS England Data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the Dataset(s) from which the information was derived.

The outputs will be communicated to relevant recipients through the following dissemination channels:

> Journals

> Posters

> Website: A list of publications is kept up-to-date on the Genomics website: https://www.genomicsengland.co.uk/research/publications?

> Presentations at appropriate conferences

> Upload of findings onto the ‘Discovery Forum’: Genomics England works with industry partners through the Discovery Forum. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Additionally, it allows the NGRL users to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed. These partners act as a critical friend and have already made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's NGRL.

Expected measurable benefits

Gene discovery in the NGRL will create significant opportunities for scientific innovation through routine service, the focus on residual unmet need, and emphasis upon national and international collaborations. The library is expected to enable genomically-driven reclassification of rare diseases leading to opportunities to recall patients for deeper phenotyping through Rare Diseases Translational Research Collaboration (RD-TRC). RD-TRC has been setup by the National Institute for Health and Care Research and its aim is to provide research infrastructure that harnesses the strength of the NHS to support discoveries and translational research on rare diseases. These data are expected to pave the way for functional characterisation of findings, thereby adding further value to datasets, improving diagnostic utility and possibly identifying new targets and therapies.

The use of the Data could help to achieve the following benefits:

> Through the international coalition of research intellects known as the Genomics England Clinical Interpretation Partnership (GeCIP) and the Discovery Forum, the framework for Genomics England to work with Industry:

• Create a mechanism for research to continually improve the accuracy and reliability of information fed back to patients

• Add to knowledge of the genetic basis of disease

• Increase opportunities for clinical trials

• Build the evidence base to accelerate the introduction of new technologies into healthcare

> Stimulate and enhance UK industry and investment

> Provide access to this unique research data resource to industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices

> Attract inward investment from life science companies, with an aim of increasing opportunities of access to medicines that would otherwise be unavailable to UK patients

> Result in new scientific insights and discoveries

> Information linked to continually updated with long-term patient health and personal information to aid analysis by researchers.

> Increase public knowledge and support for genomic medicine by delivering an ethical and transparent programme, retaining patient and public trust and confidence. This is aided by work with a range of partners to increase knowledge of genomics.

Specific example: NHS Genomic Medicine Service (NHS GMS):

Accelerate the uptake of genomic medicine in the NHS and other healthcare systems:

> Work with NHS England and other partners to deliver a scalable NHS GMS including all genomic tests and an informatics platform to enable these services to be made widely available for NHS patients.

> Work with Devolved Nation administrations, where appropriate, to achieve the same aims of the NHS GMS.

Within the GMS, expected benefits for specific cohorts:

(a) The Rare Diseases cohort:

> Increase discovery of pathogenic variants for rare disease

> Add value with additional biological insights that build confidence in putative pathogenic variants

> Enhance the clinical interpretation of WGS in rare disease

> Develop a programme of functional pathways for other genomic tests (i.e. not whole genome sequencing), specifically transcriptomics, epigenetics, micro RNAs and biomarkers

> Return findings to the NHS for feedback to patients

> Create a unique dataset for rare diseases that may enable therapeutic innovation

(b) The Cancer cohort:

> Use WGS to identify novel driver mutations for cancer and to understand its evolutionary genetic architecture through primary and secondary malignant disease

> Partner stratified healthcare programmes and outcome studies with patients from the NHS in England, to enable understanding of WGS benefits in defining predictors of therapeutic response to cancer therapies

> To use approaches using other genomic tests to offer additional biological insights into cancer

> To utilise WGS to identify new pathways for cancer therapies and improved diagnostic characterisation.

The expected patient benefit is to provide clinical diagnosis, and in time, new or more effective treatments for NHS patients. The discovery of new causes of disease, the offer of tailored therapies to create the best outcomes, and the priming new or more effective treatments for NHS patients, are other expected patient benefits.

Benefits reported so far

A number of benefits have already been yielded from the use of the NHS England Data held under this Agreement:

• Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The project found that using WGS led to a new diagnosis for 25% of the participants. Of these new diagnoses, 14% found variations in regions of the genome that would be missed by other methods, including other types of non-whole genomic tests.

• New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. This has enhanced genomic healthcare research by creating the largest genomic healthcare data resource in the world, which in turn will uncover answers for participants both now and in the future through genomic-level analysis of conditions.

• Accelerating the uptake of genomic medicine in the NHS: working with NHS England and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. This acceleration significantly contributed to delivering the Genomics Medicine Service (GMS) for the NHS, which makes whole genome sequencing part of routine healthcare.

• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. After involving participants in all stages of the pioneering 100,000 Genomes Project and putting a trusted system in place contributed to a major dialogue led by Ipsos MORI and commissioned by Genomics England and co-funded by UK Research and Innovation’s Sciencewise programme in 2019 (post 100,000 genomes project completion) found the public are enthusiastic and optimistic about the potential for genomic medicine.

Datasets on the current version

Legal basis for provision: Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital

Datasets approved under DARS-NIC-12784-R8W7V-v15.2
DatasetType of dataSensitivity FrequencyConfidential data
Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set Identifiable Non-Sensitive Ongoing Consent (Reasonable Expectation)
Cancer Registration Data Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
Civil Registrations of Death Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
Diagnostic Imaging Data Set (DID) Identifiable Non-Sensitive Ongoing Consent (Reasonable Expectation)
Emergency Care Data Set (ECDS) Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
HES-ID to MPS-ID HES Accident and Emergency Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
HES-ID to MPS-ID HES Admitted Patient Care Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
HES-ID to MPS-ID HES Outpatients Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Hospital Episode Statistics Accident and Emergency (HES A and E) Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
Hospital Episode Statistics Admitted Patient Care (HES APC) Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
Hospital Episode Statistics Critical Care (HES Critical Care) Identifiable Non-Sensitive Ongoing Consent (Reasonable Expectation)
Hospital Episode Statistics Outpatients (HES OP) Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
Mental Health and Learning Disabilities Data Set (MHLDDS) Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
Mental Health Minimum Data Set (MHMDS) Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
Mental Health Services Data Set (MHSDS) Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
MRIS - Cause of Death Report Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
MRIS - Cohort Event Notification Report Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
MRIS - Flagging Current Status Report Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
MRIS - List Cleaning Report Identifiable Sensitive Ongoing Consent (Reasonable Expectation)
MRIS - Members and Postings Report Identifiable Sensitive Ongoing Consent (Reasonable Expectation)

Files released

Files released counts only files released externally by DARS. Access granted in NHS England's own systems, such as its Secure Data Environment, is not included.

This agreement permits sublicensing: the applicant may pass data on to others. Anything passed on is not recorded in this register.

Patient opt-outs were not applied to any of the 2,853 files released under this agreement, across every version. About opt-outs

Files released against version 15.2 of this agreement, summarised by dataset.

Files released under DARS-NIC-12784-R8W7V-v15.2
DatasetFilesFirst releasedLast releasedOpt-outs applied
Mental Health Services Data Set (MHSDS)401 July 2024June 2026No
Mental Health Minimum Data Set (MHMDS)24 July 2024June 2026No
Mental Health and Learning Disabilities Data Set (MHLDDS)18 July 2024June 2026No
Cancer Registration Data10 April 2024July 2026No
Civil Registrations of Death10 April 2024July 2026No
Emergency Care Data Set (ECDS)9 July 2024June 2026No
Hospital Episode Statistics Admitted Patient Care (HES APC)9 July 2024June 2026No
Hospital Episode Statistics Critical Care (HES Critical Care)9 July 2024June 2026No
Hospital Episode Statistics Outpatients (HES OP)9 July 2024June 2026No
Hospital Episode Statistics Accident and Emergency (HES A and E)7 July 2024June 2026No
Diagnostic Imaging Data Set (DID)4 November 2025July 2026No

Version history

The register lists each renewal of this agreement as a separate row. This site has 11 versions — earlier versions existed before this site's records begin.

DARS-NIC-12784-R8W7V-v15.2 3 April 2024 to 2 April 2027
Title
Genomics England: Use of data within the National Genomics Research Library (NGRL)
Commercial
Yes
Sublicensing
Yes
Datasets
20
Files released
510

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report

What changed from DARS-NIC-12784-R8W7V-v14.5

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v14.5
FieldWasBecame
Start date2023-11-242024-04-03
End date2026-11-232027-04-02

Objective for processing

Amendments included in version (v14) are not permitted under this Data Sharing Agreement. V14 permits ongoing processing for the purposes described in v13. V14 is a three year agreement, however data will only to be released for the next six months. This is a pragmatic approach to enable Genomics England to continue accessing data until proposed amendments in v14 are approved. Summary of processing permitted under v13: > Processing was permitted for the 100,000 genomes project which was designed to create a new Genomic Medicine Service (GMS) for the NHS. Genomics England also continued to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project. > The following datasets were accessed: Hospital Episode Statistics: Outpatients (HES OP), HES Admitted Patient Care (APC), HES Outpatients (OP), HES Critical Care (CC), HES Accident & Emergency (A&E), Diagnostic Imaging Dataset (DID), Mental Health Minimum Data Set (MHMDS), Mental Health and Learning Disabilities Data Set (MHLDDS), Mental Health Services Data Set (MHSDS), Patient Reported Outcome Measures (PROMs), Civil Registrations of Death and Demographics. Demographics and PROMs data is no longer required and has been destroyed. > All Genomics research analysis was carried out within the Genomics England Trusted Research Environment (TRE).> Genomics England were permitted to share pseudonymised data under sublicense, including commercially through its Discovery Forum via their TRE. > The territory of use was worldwide. > NCRAS and uncurated NCRAS are also available within the TRE to link to the NHS England datasets held under this Agreement using the Study_ID. > Genomics England is the controller on this Agreement. > AWS and Lifebit are processors on this Agreement. They provide the technical infrastructure to host the research environment that data is loaded into, and researchers’ access. > The following special condition was included: “The data controller must remove all those participants from the cohort who have been recruited using the Consultee process, before submitting the cohort to NHS England for data.” SUMMARY OF CHANGES BETWEEN THE V13 AND V14 DSA. THE CHANGES BELOW WILL NOT BE PERMITTED UNDER THIS V14 DSA BUT WILL BE INCLUDED IN A FUTURE AMENDMENT: > Section 5 has been re-written to put the focus on the data being used broadly within the National Genomics Research Library (NGRL), formerly known as the Genomics TRE. > Detail of the GMS has been added as an example use of the NGRL data. > The Generations study has been added as an example use of the NGRL data, and with it, identifiers for this new cohort of newborns and their mothers is to be sent to NHS England for linkage. > Additional consent materials for the NGRL and Generations study have been provided to NHS England for review and inclusion. This is to widen the cohort from only participants of the 100,000 genomes project to also include participants recruited through the GMS or via the Generations study. > Additional datasets (MSDS and CSDS) have been requested for the Generations cohort only. (to be disseminated under a sister DSA (DARS-NIC-733503-V0X9Q) *** [2 paragraphs unchanged] The NGRL is a secure national resource of genomic, health and sample [46 words unchanged] Genomic Medicine Service (GMS), the 100,000 Genomes Project, and other studies/programmes where patients/participants patients’/participants’ genomes have been sequenced, and provides a national standardised genomic research resource. [16 words unchanged] better understand diseases, develop new treatments and can lead to new discoveries. Only research using cohorts covered by this Agreement for the purposes specified can process the NHS England Data. The following are examples of the aims of the use of data within the NGRL provided by Genomics England: The NGRL contains NHS England Data linked at record-level with data on the 100,000 Genomes Project cohort. Genomics England will also store NHS England Data linked at record-level with data on other Genomics England programmes. (1) This Data Sharing Agreement (DSA) covers Data provided for the 100,000 Genomes Projects cohort and the NHS Genomic Medicine Service (NHS GMS) cohorts. NHS Genomic Medicine Service (NHS GMS) [23 paragraphs unchanged] (2) Generations Study This DSA also permits the reuse of Data from the GMS Rare Diseases or Cancer cohorts and/or from the 100K Genomes cohort alongside the Generations cohort data obtained under DARS-NIC-733503-V0X9Q. Every day, nine babies in the UK are born with a rare genetic condition that could be treated, prevented, or even cured if only it had been diagnosed when those babies were newborns. The Generation Study is aiming to find out if this situation can be improved by recruiting 100,000 newborn babies at NHS Trusts throughout England and conducting WGS to screen for over 200 rare genetic conditions that are treatable in early childhood. It is hoped that babies affected by these conditions will be identified more quickly, treated earlier and therefore have improved clinical outcomes. The 3 research questions that Genomics England hope to answer are: 1. Can we better diagnose, and therefore care for, children with rare diseases? 2. Can we give researchers opportunities to improve their understanding of rare disease, to develop new treatments, and diagnoses, and better understand how our genes affect our health? 3. Should we, and if so how should we, use a baby’s genome throughout their lifetime as a resource they and their doctors can use if, for example, they become ill as they get older? Additional processing for the Generation Study involves but is not limited to: > Identifying discrepancies between WGS interpretation results and Bloodspot data results from the CSDS Dataset. This will be conducted on AWS cloud via an automated system developed by Genomics England. Discrepancies will be reported in the cases interpretation portal to be made available to the treating clinician. > Evaluation of the cost effectiveness, health related outcomes, demographics monitoring and impact on the NHS. Further details of the outputs are detailed below. [12 paragraphs unchanged] > Community Services Data Set (CSDS) – necessary to evaluate the cost effectiveness of the Generation Study, to estimate the impact of WGS in newborns, and identify discrepancies in diseases identified between WGS interpretation results and the CSDS Dataset, which will be reported in the interpretation portal to be made available to the treating clinician. NHS England data is matched to the consented cohorts in the NGRL and therefore provides a more comprehensive medical history, and going forward, a more comprehensive patient journey. NHS England is the richest source of the data required. >Maternity Services Data Set (MSDS) – necessary for discovery research purposes. [3 paragraphs unchanged] For the following Datasets: HES APC; HES OP; HES A&E; ECDS; Cancer Registration Data; Civil Registrations of Death, many indirect identifiable Data data items have been requested are required because they provide valuable Data data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. NHS England. To ensure patients won’t will not be identified the researchers and their projects are examined before access is granted to the Data, Data and an agreement to not re-identify patients is signed. Further Genomics England monitors all Data leaving the NGRL and will not allow patient records to be exported. [1 paragraph unchanged] > Limited to a study cohort identified by Genomics England, comprising of patients recruited via cohorts who consented to participate in (1) the 100,000 genomes project who also consented to longitudinal research (~80,000 participants), participants) or (2) the Genomics Medicines Service (expected ~100,000 new additions per year) or (3) the Generations Project, which follows newborn babies (expected ~50,000 new additions per year - recruitment is due to begin in Dec 2023 and will continue until 100,000 participants have been recruited in 2025). year). > Genomics England will request full history of patient Data to provide maximum [21 words unchanged] no other alternative or less intrusive ways of achieving the purpose described. Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data. NHS England has commissioned Genomics England to undertake the work. NHS England does not specify what Data are required to deliver the work nor how the Data shall be processed to achieve that purpose. Such decisions are taken by Genomics England. [96 paragraphs unchanged] A release register detailing any sub licences and onward sharin sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract. In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users. NHS England will require the ability to audit the sub licensee.

Processing activities

[1 paragraph unchanged] NHS England Data will provide the relevant records from the HES APC, HES CC, HES OP, HES A&E, ECDS, MHSDS, MHMDS, Cancer Registration Data, Civil Registrations of Death, PROMs, DIDs, CSDS PROM and MSDS Datasets DIDs datasets to Genomics England’s Amazon Web Services (AWS) cloud storage. The CSDS, MSDS, HES APC, HES OP, HES CC, HES A&E, ECDS, Civil Registrations of Deaths and DIDs Data for the Generations Study cohort will flow via DARS-NIC-733503-V0X9Q-v0. The Data will: [12 paragraphs unchanged] - Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for the remote access; [2 paragraphs unchanged] The Data will be processed worldwide. Data is physically stored in England. Remote access is permitted from the following specified countries: UK, EEA Countries, United States, Canada, Australia, Qatar, Republic of Korea, Japan, Switzerland, Brazil, India, New Zealand, Argentina Should any country on the permitted list above become a high risk country through the duration of this DSA, the Recipient will cease disseminating data to researchers/organisations based in that country and request that data already disseminated be destroyed. Should the Recipient wish to share data with any countries not listed above, it will require an update to this DSA. Should Genomics England wish to facilitate remote access from a country that is not listed above, prior written agreement from NHS England must be obtained. Genomics England upholds the following safeguards and controls: 1. Compliance with National Cyber Security Centre (NCSC) guidance, leading to the implementation of geo-blocking measures for IP addresses originating from Iran, Russia, North Korea and Belarus 2. Collaboration with the NCSC and other security partners to identify and block potentially risky IP addresses, irrespective of their country of origin. 3. Implementation of two email authentication methods, namely Domain-based Message Authentication Reporting and Conformance (DMARC) and Sender Policy Framework (SPF), to detect and respond to spoofing and spam. This is crucial, as these activities often target our firewalls from international IP addresses. 4. Introduction of additional assurance activities related to international access within the Office 365 estate. 5. Conducting due diligence on companies associated with BGI Genomics. 6. Responsibilities of the Access Review Committee (ARC) include the thorough review of applications and applicants. 7. Continuous improvement of Information Governance training and cybersecurity awareness at Genomics England. [2 paragraphs unchanged]

Expected measurable benefits

[31 paragraphs unchanged] Specific Example: Generations Study: Within the Generations Study, Genomics England intend to use NHSE Data to provide evidence to answer the questions set out above for evaluation purposes. Specifically, HES and CSDS will be used for the following: • Cost effectiveness To approximate the true costs associated with additional / fewer =healthcare encounters of WGS in newborns, to include A&E attendances, outpatient appointments, admissions, allied health professional appointments, procedures, medication and treatment for all participants that screen positive through the Generation Study. • Health related outcomes To estimate the impact of WGS in newborns on the following, for all participants that screen positive through the Generation Study: o Diagnostic Odyssey to include i) time from first clinical contact to diagnosis, ii) age at diagnosis, iii) frequency and duration of health encounters during the diagnostic period. o Health encounters such as A&E attendances, outpatient appointments, admissions, allied health professional appointments over a defined time period. o Interventions for example procedures, medication and treatment over a defined time period. o Mortality to include i) age at death, and ii) cause of death. • Demographics monitoring To ensure enrolled participants are representative of the wider English population based on a variety of demographic variables. • Impacts on the NHS To identify if there has been any impact on the uptake of existing NHS newborn screening amongst participants of the Generation Study Genomics England also hope to link NHSE Data to genomic data for discovery research purposes. HES, CSDS and MSDS are expected to prove useful for researchers seeking opportunities to improve their understanding of rare disease, to develop new treatments, and diagnoses, and better understand how genes affect health.

Benefits reported

[3 paragraphs unchanged] • Accelerating the uptake of genomic medicine in the NHS: working with NHSE NHS England and other partners to deliver a scale-able WGS and informatics platform to [59 words unchanged] for the NHS, which makes whole genome sequencing part of routine healthcare. [2 paragraphs unchanged]

Unchanged: Expected output.

DARS-NIC-12784-R8W7V-v14.5 24 November 2023 to 23 November 2026
Title
Genomics England: Use of data within the National Genomics Research Library (NGRL)
Commercial
Yes
Sublicensing
Yes
Datasets
20
Files released
135

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report

What changed from DARS-NIC-12784-R8W7V-v13.4

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v13.4
FieldWasBecame
TitleGenomics England (MR1418) - Renewal Request for tranche of data across multiple data sets.Genomics England: Use of data within the National Genomics Research Library (NGRL)
Start date2023-08-012023-11-24
End date2023-10-312026-11-23
Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set: type of dataAnonymised - ICO Code CompliantIdentifiable
HES-ID to MPS-ID HES Accident and Emergency: type of dataAnonymised - ICO Code CompliantIdentifiable
HES-ID to MPS-ID HES Admitted Patient Care: type of dataAnonymised - ICO Code CompliantIdentifiable
HES-ID to MPS-ID HES Outpatients: type of dataAnonymised - ICO Code CompliantIdentifiable

Datasets: − Demographics; − Patient Reported Outcome Measures (Linkable to HES)

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics. Amendments included in version (v14) are not permitted under this Data Sharing Agreement. V14 permits ongoing processing for the purposes described in v13. V14 is a three year agreement, however data will only to be released for the next six months. This is a pragmatic approach to enable Genomics England to continue accessing data until proposed amendments in v14 are approved. Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records. Summary of processing permitted under v13: Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem. > Processing was permitted for the 100,000 genomes project which was designed to create a new Genomic Medicine Service (GMS) for the NHS. Genomics England also continued to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project. The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, through the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure. > The following datasets were accessed: Hospital Episode Statistics: Outpatients (HES OP), HES Admitted Patient Care (APC), HES Outpatients (OP), HES Critical Care (CC), HES Accident & Emergency (A&E), Diagnostic Imaging Dataset (DID), Mental Health Minimum Data Set (MHMDS), Mental Health and Learning Disabilities Data Set (MHLDDS), Mental Health Services Data Set (MHSDS), Patient Reported Outcome Measures (PROMs), Civil Registrations of Death and Demographics. Demographics and PROMs data is no longer required and has been destroyed. Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to participants: for many they are receiving a diagnosis for the first time. > All Genomics research analysis was carried out within the Genomics England Trusted Research Environment (TRE).> Genomics England were permitted to share pseudonymised data under sublicense, including commercially through its Discovery Forum via their TRE. The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness. > The territory of use was worldwide. The richness of the high-quality data sets is crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy. > NCRAS and uncurated NCRAS are also available within the TRE to link to the NHS England datasets held under this Agreement using the Study_ID. The 100,000 Genomes Project completed recruitment of rare disease participants and cancer patients in early 2019. > Genomics England is the controller on this Agreement. Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project. > AWS and Lifebit are processors on this Agreement. They provide the technical infrastructure to host the research environment that data is loaded into, and researchers’ access. To achieve these goals Genomics wish to continue to access the following datasets: > The following special condition was included: “The data controller must remove all those participants from the cohort who have been recruited using the Consultee process, before submitting the cohort to NHS England for data.” Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants. SUMMARY OF CHANGES BETWEEN THE V13 AND V14 DSA. Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations. THE CHANGES BELOW WILL NOT BE PERMITTED UNDER THIS V14 DSA BUT WILL BE INCLUDED IN A FUTURE AMENDMENT: Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants. > Section 5 has been re-written to put the focus on the data being used broadly within the National Genomics Research Library (NGRL), formerly known as the Genomics TRE. Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records. > Detail of the GMS has been added as an example use of the NGRL data. Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts. > The Generations study has been added as an example use of the NGRL data, and with it, identifiers for this new cohort of newborns and their mothers is to be sent to NHS England for linkage. Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that Genomics' data applications are correct, do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate. > Additional consent materials for the NGRL and Generations study have been provided to NHS England for review and inclusion. This is to widen the cohort from only participants of the 100,000 genomes project to also include participants recruited through the GMS or via the Generations study. Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum. > Additional datasets (MSDS and CSDS) have been requested for the Generations cohort only. (to be disseminated under a sister DSA (DARS-NIC-733503-V0X9Q) As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their pseudonymised genome and health data. *** The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set. Genomics England requires access to NHS England Data for use in their National Genomics Research Library (NGRL), which operates as a Trusted Research Environment (TRE). The lawful basis for processing participant data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of participants. National Genomic Research Library (NGRL) The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers. The NGRL is a secure national resource of genomic, health and sample data managed by Genomics England, which builds on the research environment created by Genomics England for the 100,000 Genomes Project that completed recruitment in 2018 and involved the sequencing of approximately 100,000 genomes. It contains cohorts of patients/participants recruited via programmes set up for the NHS Genomic Medicine Service (GMS), the 100,000 Genomes Project, and other studies/programmes where patients/participants genomes have been sequenced, and provides a national standardised genomic research resource. Being able to compare all patient data in one place provides researchers with an opportunity to better understand diseases, develop new treatments and can lead to new discoveries. Only research using cohorts covered by this Agreement for the purposes specified can process the NHS England Data. The processing is necessary to support and enable Genomics England’s legitimate interests in: The following are examples of the aims of the use of data within the NGRL provided by Genomics England: · Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment) (1) NHS Genomic Medicine Service (NHS GMS) · Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer. Following the successfully delivery of the 100,000 Genomes Project, Genomics England begun work to deliver the NHS GMS. This service will offer patients the dual opportunity of routine clinical care alongside a choice to participate in research spanning across all genomic tests within the NHS, starting with whole genomes. Development of the NHS GMS builds on the evidence generated by the 100,000 Genomes Project, but also extends to other genomic testing other than Whole Genome Sequencing (WGS). Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees. The NHS GMS is led and commissioned by NHS England and will consist of: The beneficiaries are: > NHS Genomic Laboratory Hubs (GLHs) that will work as part of a National Genomic testing service. The provisions in this service will be determined by a national genomic test directory that outlines the testing strategies and technology to be employed for rare and inherited disease, cancer and other defined conditions/ applications. · Participants through the work we do will influence their care; > Clinical Genetic Services · researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data; > Cancer services using genomic analysis to guide treatment · and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK. Within the GMS, two particular cohorts are defined: (a) The Rare Diseases cohort The NGRL rare diseases cohort will consist of families with rare diseases, based on family structures appropriate to the provision of the GMS. This will harness the strength of current UK rare disease programmes and will advance current understanding of rare disease mechanisms. It may also impact on common diseases that share similar phenotypes. It will also offer opportunities for biomarker, clinical, and interventional studies through industrial partnerships. Genomics England will actively support the implementation of the UK Rare Disease Strategy. Building on work already undertaken by the 100,000 Genomes Project, this will facilitate the generation of a national data resource of all genomic data, with a focus on WGS but including all genomic testing. The goals of processing data of patients with rare disease are: • To increase discovery of pathogenic variants (the gene variant responsible for causing disease) for rare disease. • To add value with additional biological insights that build confidence in commonly accepted pathogenic variants. • To enhance the clinical interpretation of WGS in rare disease. • To develop a programme of functional pathways for genomic tests other than whole genome sequencing, specifically, transcriptomics (the study of all the ribonucleic acid (RNA) molecules within a cell, otherwise known as the transcriptome), epigenetics (the study of how cells control gene activity without changing the DNA sequence), micro RNAs and biomarkers. • To return findings to the NHS for feedback to patients. • To create a unique dataset for rare diseases that may enable therapeutic innovation. (b) The Cancer cohort The NGRL will continue to learn from and collaborate with other projects who are producing an inventory of genomic, transcriptomic and epigenomic changes in a wide range of different tumour types. Researchers from these projects access the data via an approved sublicensing agreement with Genomics England. The goals of processing data for patients with cancer are to: • Use WGS to identify novel driver mutations for cancer and to understand its evolutionary genetic architecture through primary and secondary malignant disease (by multiple biopsy and WGS). • Partner stratified healthcare programmes and outcome studies with patients from the NHS in England, to enable understanding of WGS benefits in defining predictors of therapeutic response to cancer therapies. • To use other genomic testing approaches to offer additional biological insights into cancer. • To utilise WGS to identify new pathways for cancer therapies and improved diagnostic characterisation. Other forms of collaboration include direct partnerships to develop tools to improve systems and services. For example, partnering with Lifebit to take advantage of their technical genomic data tooling. (2) Generations Study Every day, nine babies in the UK are born with a rare genetic condition that could be treated, prevented, or even cured if only it had been diagnosed when those babies were newborns. The Generation Study is aiming to find out if this situation can be improved by recruiting 100,000 newborn babies at NHS Trusts throughout England and conducting WGS to screen for over 200 rare genetic conditions that are treatable in early childhood. It is hoped that babies affected by these conditions will be identified more quickly, treated earlier and therefore have improved clinical outcomes. The 3 research questions that Genomics England hope to answer are: 1. Can we better diagnose, and therefore care for, children with rare diseases? 2. Can we give researchers opportunities to improve their understanding of rare disease, to develop new treatments, and diagnoses, and better understand how our genes affect our health? 3. Should we, and if so how should we, use a baby’s genome throughout their lifetime as a resource they and their doctors can use if, for example, they become ill as they get older? Additional processing for the Generation Study involves but is not limited to: > Identifying discrepancies between WGS interpretation results and Bloodspot data results from the CSDS Dataset. This will be conducted on AWS cloud via an automated system developed by Genomics England. Discrepancies will be reported in the cases interpretation portal to be made available to the treating clinician. > Evaluation of the cost effectiveness, health related outcomes, demographics monitoring and impact on the NHS. Further details of the outputs are detailed below. The following NHS England Data will be accessed: > Hospital Episode Statistics (HES) - these Datasets provide the core clinical data for participants and are vital to the provision of a detailed longitudinal medical history for participants. Specifically, the following HES Datasets are required: - Admitted Patient Care (APC) - Accident & Emergency (A&E) - Critical Care (CC) - Outpatients (OP) > Emergency Care Data Set (ECDS) – necessary to understand which patients attend emergency departments and what treatment they receive in order to assess if there are associations with genetic markers. > Diagnostic Imaging Dataset (DID) – necessary to provide invaluable, detailed information to build on participants’ phenotypes (observable characteristics), e.g., tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations. > Mental Health – necessary because the GMS includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are expected to be some of the most prevalent conditions within the Project. For the 100,000 genomes project approximately 10% of project participants have a mental health record. Mental health Data are therefore vital in ensuring that a complete and relevant medical history is available for all participants. > Civil Registration Mortality – necessary because deaths Data is essential for performing survival analyses; this is crucial information for research in combination with other medical history. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts. > Cancer Registration – necessary because it provides a clinical background for the patient's cancer journey which may help provide scientific and medical insights when compared against the patient's genome. > MRIS Events Mortality and Demographic Data - still required for open-ended research, specifically for researchers which have already been using this Data and need to go back to it. Losing this Dataset may cause disruption to their research. > Community Services Data Set (CSDS) – necessary to evaluate the cost effectiveness of the Generation Study, to estimate the impact of WGS in newborns, and identify discrepancies in diseases identified between WGS interpretation results and the CSDS Dataset, which will be reported in the interpretation portal to be made available to the treating clinician. >Maternity Services Data Set (MSDS) – necessary for discovery research purposes. The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records, such as blood pressure, cholesterol, glucose, and pharmacogenomics (a field of research that studies how a person's genes affect how he or she responds to medications.), adds significant value. The richness of the NGRL datasets will allow Genomics England to move beyond the primary phenotype of the rare disease, cancer or infectious disease that led to the patient’s initial WGS in the context of other continuous traits, diseases and response to therapy including harm. The level of the Data will be: > Identifiable For the following Datasets: HES APC; HES OP; HES A&E; ECDS; Cancer Registration Data; Civil Registrations of Death, many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. To ensure patients won’t be identified the researchers and their projects are examined before access is granted to the Data, and an agreement to not re-identify patients is signed. Further Genomics England monitors all Data leaving the NGRL and will not allow patient records to be exported. The Data will be minimised as follows: > Limited to a study cohort identified by Genomics England, comprising of patients recruited via (1) the 100,000 genomes project who also consented to longitudinal research (~80,000 participants), (2) the Genomics Medicines Service (expected ~100,000 new additions per year) or (3) the Generations Project, which follows newborn babies (expected ~50,000 new additions per year - recruitment is due to begin in Dec 2023 and will continue until 100,000 participants have been recruited in 2025). > Genomics England will request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described. Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data. NHS England has commissioned Genomics England to undertake the work. NHS England does not specify what Data are required to deliver the work nor how the Data shall be processed to achieve that purpose. Such decisions are taken by Genomics England. The lawful basis for processing personal data under the UK GDPR is: > Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party. Genomics England has determined the processing is necessary for its legitimate interests in carrying out medical research on the causes, diagnosis and treatment of rare diseases and cancers. The lawful basis for processing special category data under the UK GDPR is: > Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject. It is necessary for Genomics England to process special category participant data for carrying out medical research on the causes, diagnosis and treatment of rare diseases and cancers, which is expected to benefit patients. The funding is provided by the Department of Health and Social Care. The funding is specifically for the projects described. Funding is in place until March 2025, with the intention to renew this funding periodically. The funder will have no ability to suppress or otherwise limit the publication of findings. Lifebit provides IT support to Genomics England. Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England. Representatives from patient and public bodies have an important role to play in Genomics England commercial initiatives. These representatives ensure transparency is upheld, and the interest of those whose data is being used is always being respected. In the early stages of the Library, Genomics England undertook a range of work to ensure that potential participant’s views were included in the formulation of the ethical policies submitted for research ethics approval and in the development of patient information. The views of different groups of potential participants (those affected by cancer, rare disease, and those from BAME communities) in relation to ethical issues raised by the 100,000 Genomes Project were sought and findings were published on the Genomics England website (See all reports under ‘patient and public involvement - https://www.genomicsengland.co.uk/library-and-resources/ and the Genomics England Engagement Strategy). Genomics England will continue to engage with these stakeholders. Further to this, each of the 13 currently recruiting NHS Genomic Medicine Centres had dedicated Patient and Public Involvement leads (PPI) who are responsible for engaging with and involving local potential participant groups from diverse backgrounds. It is expected that the future NHS GMS will continue these local PPI activities to shape and inform the service. A Participant Panel has also been established. This 30-strong group has provided invaluable advice on a range of topics, for instance, in shaping how analysis is monitored, how results are returned, and how advice and support should be framed. Participant Panel members have either donated samples to the Library themselves or are carers of participants. They take part in a wide variety of consultative groups, such as the Genomics England Ethics Advisory Committee but most importantly are guardians of the dataset, with representatives on the Access Review Committee. Participants play an important part in every decision made about access to data. SUB-LICENCING: Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted. It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics. The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not. All requests for data access will be subject to the following considerations: • Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval). • Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG). • Provision of a signed Genomics England data access agreement to the Access Review Committee. • Prioritisation of access according to resource availability. • Facilitation of high-quality health research Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into: • Diagnostic discovery • Pre-clinical research • Clinical Trials Referral • Real World Evidence / Market Access Approval process for Commercial organisations for access to the NGRL: Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal. The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups. Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits. Approval process for GeCIP users (academic) of the NGRL: • Researcher visits Genomics England website to enrol as a GECIP member • Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project) • Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access. • On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC. • Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable. • If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated. All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes: • Identification of diagnoses originally missed by the standardised pipeline • Feedback of new diagnoses to patients • Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access. Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England. AIRLOCK POLICY: The following rules are applied to all airlock requests: 1. All relevant details of the summary results to be transferred must be provided with every request. 2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection. 3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer. 4. Summary results requested for transfer are assessed using the following criteria: a. whether the request aligns with the users ARC approval in full; b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL; c. any data security implications; d. any disclosure risks; e. the technical feasibility and associated cost of the request; f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared; g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals. The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of: • Technical Lead • User Community Representative • Bioinformatics Director • Caldicott Guardian • Chief Scientist representative The Data will be processed worldwide. Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following: • UK academic research institutions (e.g., universities, research institutions etc.) • NHS trusts or authorities • UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project • Foreign universities and research institutions that carry out significant research activity • UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE)) • Foreign healthcare organisations (private or public) that undertake significant research activity To be eligible for data access as a GeCIP member, applicants must meet these requirements: • Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy. • Their host institution has verified that they are affiliated with that institution. • The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below). • The GeCIP domain lead has approved the application. • Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure. • Data access will not then be granted until a researcher has successfully passed mandatory information governance training. All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)). The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel. GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work. All personnel accessing the Data have been appropriately trained in data protection and confidentiality. The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data: > National Cancer Registration and Analysis Service (NCRAS) and uncurated NCRAS data > Secure Anonymised Information Linkage (SAIL) data; Welsh data > Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum) > NHS Trusts data The Data will not be linked with any other data. The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset. To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate. Data accessed under sub licence will not be re-identified. Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL. Data shared through the Airlock process is aggregate data only and is therefore not personal data so does not require a legal basis under the UK GDPR. A release register detailing any sub licences and onward sharin

Processing activities

All organisations party to this agreement must comply with the Data Sharing Framework Contract requirements, including those regarding the use (and purposes of that use) by "Personnel" (as defined within the Data Sharing Framework Contract i.e.: employees, agents and contractors of the Data Recipient who may have access to that data). Genomics England will transfer data to NHS England. The data will consist of identifying details (specifically study ID, NHS Number, Date of Birth, Surname, Forename, Gender, Postcode and Other Given Name) which are required for the cohort to be linked with NHS England Data. This is the minimum requirement of identifiers required for linkage to guarantee complete matching. PROCESSING ACTIVITIES NHS England Data will provide the relevant records from the HES APC, HES CC, HES OP, HES A&E, ECDS, MHSDS, MHMDS, Cancer Registration Data, Civil Registrations of Death, PROMs, DIDs, CSDS and MSDS Datasets to Genomics England’s Amazon Web Services (AWS) cloud storage. The CSDS, MSDS, HES APC, HES OP, HES CC, HES A&E, ECDS, Civil Registrations of Deaths and DIDs Data for the Generations Study cohort will flow via DARS-NIC-733503-V0X9Q-v0. The Data will: The first stage of processing focuses on quality verification. This ensures that the data set is complete, accurate and complies with the NHS data dictionary or relevant specification. Participant identifiers in the dataset are verified against Genomics England’s participant details and any updates required to identifiable data fields, e.g., dates of birth, are highlighted. Finally, the data set is reviewed against recent participant withdrawals so that any withdrawals notified after the data application was made can be removed from the data sets. > Contain directly identifying Data items including but not limited to: Names, Postcode, Cause of Deaths, Place of Birth, Cancer Registration Number, which are required to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Following this, the data are pseudonymised, as all subsequent processing can be performed without direct identifiers. Genomics England has compiled lists of identifiable and sensitive fields for each data set in line with details provided by NHS England and following internal review of data sets. Pseudonymisation is a key facet of the Genomics England resource. Pseudonymised data are uploaded to a secure trusted research environment (TRE) hosted by Genomics England on a quarterly basis, where they are linked to participant genomes and primary clinical data. The NHS England Data is pseudonymised within the AWS cloud and is then loaded into the NGRL. Raw, identifiable files are kept in a secure location on AWS. The second stage of processing involves the selection of a pseudonymised cohort of participants that fulfil a specific research request. Researchers are members of a Genomics England Clinical Interpretation Partnership (GECiP) or the Discovery Forum. Research requests are assessed to ensure that they are included in the approved use purposes set out in the Genomics England Protocol and fall within the scope of the relevant GECiP or the Discovery Forum. Researchers declare any data they wish to bring into the TRE and any tools they wish to use for analysis. The Data will not be transferred to any other location. The third stage of processing is the analysis of the pseudonymised data sets within the TRE. Researchers perform all the analysis and processing within the environment: they do not extract pseudonymised data. Results data are placed in a secure folder for anonymisation verification before extraction. The Data will be stored on the NGRL and the AWS Cloud at Genomics England. Genomics England provided NHS England with a cohort for linkage. Data updates are disseminated on a regular, ongoing basis to enable analysis of the latest available data. Genomics England stores NGRL data on the Cloud provided by Amazon Web Services (AWS). THE TRUSTED RESEARCH ENVIRONMENT (TRE) The Data will be accessed by authorised personnel via remote access. All research analysis on the Genomics England data-set will only be carried out via a secure analysis environment hosted within the Genomics England data centre – the Genomics England TRE. Analytical tools and applications are available within the TRE. No sequencing or clinical data are made available for download, users cannot copy or paste out of the TRE, and there is limited internet access within it (i.e. whitelisted sites). Movement of files into and out of the TRE is governed via an ‘Airlock’ Policy. The Controller(s) must confirm and provide evidence upon audit by NHS England that access via any remote device complies with the data security obligations within this DSA and the Data Sharing Framework Contract. Academic researchers access the TRE by applying to be a member of a GeCIP domain. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following: For remote access: • UK academic research institutions (e.g. universities, research institutions etc.) - Remote access will only be from secure locations situated within the territory of use (as further restricted elsewhere within the DSA if so done) stated within this DSA; • NHS trusts or authorities - Access controls granting users the minimum level of access required are in place; • UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project - Remote access is only via secure connections (e.g., VPNs or secure protocols) to protect Data; • Foreign universities and research institutions that carry out significant research activity - Multifactor authentication (MFA) is required for remote access; • UK and foreign governmental departments that carry out significant research activity (e.g. MRC, NIH, PHE) - Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for the remote access; • Foreign healthcare organisations (private or public) that undertake significant research activity - All remote access is undertaken within the scope of the organisation’s DSPT (or other security arrangements as per this DSA) and complies with the organisation’s remote access policy. Membership is not open to those who are self-employed or employed by: The above applies in addition to any condition set out elsewhere within the DSA (e.g. who may carry out processing, and for what purpose). • private UK healthcare institutions The Data will be processed worldwide. • commercial companies. Access to confidential patient identifiable Data is restricted to an extremely limited number of employees of Genomics England, accessible on AWS. To be eligible for data access as a GeCIP member, applicants must meet these requirements: Substantive employees of Genomics England and researchers who are a member of the GeCIP and Discovery Forum will process the Data for the purposes described above. • Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy. • Their host institution has verified that they are affiliated with that institution. • The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below). • The GeCIP domain lead has approved the application. Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure. Data access will not then be granted until a researcher has successfully passed mandatory information governance training. NCRAS and uncurated NCRAS are also available within the TRE to link to the NHS England datasets held under this Agreement using the Study_ID. No other clinical non-NHS England data is linked to data held in the TRE under this Agreement. COMMERCIAL RESEARCHER ACCESS TO THE TRE Genomics England operates a membership-based forum – the Discovery Forum – which is open to a range of companies world-wide and allows access to the TRE. It provides a platform for collaboration between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Each Discovery Forum member signs a Data Access Agreement with Genomics England. This states the research purposes which the company is authorised to carry out, and stipulates the number of genomes sequences that can be accessed. It covers the Company’s behaviour and working practices: in particular it binds users to Genomics England’s Airlock Policy, Information Governance, IT Security and Data Protection Polices. Companies need to nominate named individuals to be their Researchers who must complete information governance training before accessing data. Once the Data Access Agreement is in place, each research project undertaken by the Company within the TRE must receive prior Access Review Committee (ARC) approval. Discovery Forum members access the TRE in a similar manner to GeCIP Researchers: all research is carried out within the TRE, and any movement of results out of the environment occurs only through the Airlock Process. THE ACCESS REVIEW COMMITTEE The ARC provides an independent examination of requests for data access, with regards to the acceptable uses of the Genomics England dataset which are outlined in the 100,000 Genomes Project Protocol and Data Access and Acceptable Uses Policy. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel which is made up of participants and parents/carers involved in the 100,000 genomes project. THE AIRLOCK PROCESS The Genomics England TRE has been developed with the intention that all data analysis is carried out within it and that the only data to leave it are summary results. An Airlock process has been established which enables material (data, files, tools etc) to be moved in or out of the TRE in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access. Removal of summary results therefore requires an Airlock request. The following rules are applied to all Airlock requests: 1. All relevant details of the summary results to be transferred must be provided with every request. 2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection. 3. All summary results transferred will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer. 4. Summary results requested for transfer are assessed using the following criteria: • whether the request aligns with the users ARC approval in full • whether the request can clearly be demonstrated to be aligned with a registered project in the TRE • any data security implications • any disclosure risks • the technical feasibility and associated cost of the request • when importing data, its scientific value to the community of researchers within the TRE, and when and how it will be shared • when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals. The Airlock process is governed by the Airlock Policy, which defines the process and governance of the Airlock process. A set of Airlock Policy Guidelines presents the rules-of-thumb/principles that will be referenced by both the researcher (during preparation of analysis results) and the output checker (during output-checking). Analysed results are inspected to ensure they cannot be used to disclose the identity of the participants. Checking of summary results by the Airlock Review Team is governed by a set of principles that guide individual decisions. By using a principles-based approach where each case is assessed individually the security of the dataset is maintained by exporting only appropriate data. Review of transfer requests resulting in public-sharing/publication of data will be checked more stringently. Any approved Airlock export can only be used for the specific use detailed in the original request. The TRE contains pseudonymised longitudinal data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be used and what can be exported. Where an export contains pseudonymised longitudinal data, Genomics England will always apply the requirements placed on them as conditions of having access to the data. All Airlock requests go through a robust approval process and the Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. Where a precedent has been set and a clear set of principles and rules are in place for types of research, the Airlock Manager can approve the request. For more complicated requests or where no precedent has been set these will go to the Airlock Committee and be reviewed by the Airlock Review Team. The Airlock Review Team is a delegation of the Genomics England Chief Scientist responsible for oversight of all airlock requests in accordance with the Airlock Policy and the groups Terms of Reference. It comprises: • Technical Lead • User Community Representative • Bioinformatics Director • Caldicott Guardian • Chief Scientist representative SUB LICENSING Genomics England has developed the TRE to allow registered third parties to access pseudonymised versions of the data that it holds, for the purposes of approved research. The TRE contains External Data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be accessed. Genomics England will always apply the requirements placed on them as holders of External Data to users of the TRE as a condition of having access to the data. The data is NOT for onward sharing outside of the TRE. DATA CONTROLLER Genomics England in its provision of whole genome sequencing are applying to NHS England for secondary clinical data to link to the genomic data. With regards to data provided by NHS England, Genomics England are the sole data controller. Genomics England provides NHS England with linking data in order to receive longitudinal data sets. These data sets are delivered to Genomics England by NHS England on a quarterly basis having been approved by the Independent Group Advising on the Release of Data (IGARD). Genomics England identifies the linking data and agrees with NHS England the scope of the longitudinal data being provided. Genomics England determines the method of pseudonymisation and storage within the TRE and secures this data for use by approved researchers only. Genomics England determines who these researchers are. Genomics England is the Data Controller for longitudinal data sets processed in the Genomics England TRE. Researchers in academic, educational or commercial organisations Access to pseudonymised data in the TRE which will include longitudinal data sets (HES etc) provided by NHS England. Access to the TRE only allowed under access agreement. The individual researchers are Data Controllers when carrying out research within the research environment. Research which is approved using the data will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/ DATA PROCESSORS AWS and Lifebit are additional data processors on this Agreement. They provide the technical infrastructure to host the research environment that data is loaded into, and researchers access. o Only summary level data can be removed from the environment. o Approved researchers will only be able to access Lifebit’s Platform as a service (PaaS) Cloud Operating System (OS) through a virtual desktop. o Secondary data will be ingested into CloudOS. o CloudOS will be hosted within GEL’s London AWS environment - All data is encrypted in transit and at rest. o CloudOS controls access to the secondary data. o A security and Data Protection Impact Assessment (DPIA) will be conducted prior to loading live Lifebit: Lifebit has been selected as platform partner to deliver the Research Environment after reviewing several proposals. The UK-based Subject Matter Expert (SME) offered a proven and innovative technology solution offering a blend of robustness and ease of use. Lifebit CloudOS provides a secure and collaborative workspace to enable researchers to easily perform genomic data analysis. The platform will deliver an intuitive, integrated and collaborative user experience and enables fast, effective research outcomes across a wide range of academic and biotech/pharma researchers with varying levels of technical competency. DATA MINIMISATION Genomics England’s TRE aligns to the current NHS guidance and is currently cited as an example of best practice. The detail to support this is below. In the HDR UK TRE Principles and Best Practices paper from December 2021, the Genomics England model is explicitly described (page 17) and this document has a foreword authored by the Director of Data Policy, NHSX and Director of Tech Policy, NHSX. https://www.hdruk.ac.uk/news/new-principles-published-to-improve-public-confidence-in-access-and-use-of-data-for-health-research-through-trusted-research-environments. Conceptually, in terms of the 5 safes framework, the safes work together to protect the privacy of the individual. The TRE model makes 4 of the safes: safe-setting, safe-people, safe-projects and safe-outputs very strong, meaning that it’s possible to allow the criteria for safe-data to be relaxed (i.e., de-identification only) while maintaining the overall level of privacy protection. The TRE model effectively moves data minimisation to the output stage - ’safe outputs’, i.e., the airlock, where the airlock managers and the airlock committee do review requests to ensure summary data is minimised to only what is necessary to demonstrate externally a particular research result. The data released by the Airlock project is aggregated summary data. Given these measures, there has even been a discussion about data within TREs being regarded as "functionally anonymous” because of these safeguards, which would put them outside the constraints of GDPR. The risk is further minimised as all the researchers are under contractual obligation not to re-identify individuals from the data. By allowing researchers access to all the data it allows "hypothesis free research" and means they may discover things they didn't suspect. The full dataset is “adequate and relevant” for all researchers to access and the TRE framework provides the required limitation. Research in the RE is around discovering genome phenome relationships and given the complexity of the genome and how little we understand about it, it’s impossible to predict what phenotype terms may be involved, so researchers need to be able to explore all of them, rather than repeatedly requesting different datasets. Each individual participant has agreed and is aware that their data is in the TRE and accessed by researchers. Access to the data is necessary for this purpose to ensure the research can be carried out and the most value extracted from the data. Genomics England do have a wide range of controls outlined above to make sure the risk is minimised. Both Article 89 and Article 25(1) of the UK GDPR support this approach. This will be limited to the selected cohort and additions and deletions will be updated regularly. There will be no data linkage undertaken with NHS England data provided under this agreement that is not already noted in the agreement. PROMS terms and conditions will be adhered to. PROMS data is only available for non-commercial purposes, such as academic research, or in connection with delivering services to the NHS. A Data Protection Impact Assessment was conducted in November 2021 prior to the transfer of the data to the new environment and is being continually assessed and updated by the data protection team.

Expected output

Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the TRE to support this ongoing research. Each research group who accesses this data Researchers will have their own defined output strategy dissemination and expectations which they will aim to deliver against. A communication strategies, however a full list of scientific publications and conferences/posters can will be found made available on the Genomics England website. Recent specific examples website on an ongoing basis. The expected outputs of journal outputs are as follows: the processing will be: - A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. Alexander J.M. Blakes, Htoo Wai, Ian Davies, et al. Genome Medicine Jul 2022 DOI: 10.1186/s13073-022-01087-x. This demonstrated the clinical value of examining non-canonical splicing variants in individuals with unsolved rare diseases by identifying 35 new likely diagnoses from 258 probands with an unsolved rare disease. > Submissions to peer reviewed journals - Prognosis and oncogenomic profiling of patients with tropomyosin receptor kinase fusion cancer in the 100,000 genomes project. John Bridgewater, Xiaolong Jiao, Mounika Parimi, et al. Cancer Treatment and Research Communications Aug 2022 DOI: 10.1016/j.ctarc.2022.100623. This study supports the hypothesis that Neurotrophic tyrosine receptor kinase (NTRK) gene fusions are primary oncogenic drivers and the co-occurrence of NTRK gene fusions with other oncogenic alterations is rare. Overall survival was not statistically significant between NTRK and non-NTRK tumour patients. > Presentations at conferences - Participant experiences of genome sequencing for rare diseases in the 100,000 Genomes Project: a mixed methods study. Michelle Peter, Jennifer Hammond, Saskia C. Sanderson, et al. European Journal of Human Genetics Mar 2022 DOI: https://doi.org/10.1038/s41431-022-01065-2. This study highlights how expectation-setting is critical when offering genome sequencing, and that post-test counselling is important regardless of the GS result received, with parents perhaps needing additional emotional support. > Posters > Creation of a database of all genomic data, including all genomic and omics tests. The outputs will not contain NHS England Data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the Dataset(s) from which the information was derived. The outputs will be communicated to relevant recipients through the following dissemination channels: > Journals > Posters > Website: A list of publications is kept up-to-date on the Genomics website: https://www.genomicsengland.co.uk/research/publications? > Presentations at appropriate conferences > Upload of findings onto the ‘Discovery Forum’: Genomics England works with industry partners through the Discovery Forum. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Additionally, it allows the NGRL users to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed. These partners act as a critical friend and have already made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's NGRL.

Expected measurable benefits

The aims of the 100,000 genomes project are: Gene discovery in the NGRL will create significant opportunities for scientific innovation through routine service, the focus on residual unmet need, and emphasis upon national and international collaborations. The library is expected to enable genomically-driven reclassification of rare diseases leading to opportunities to recall patients for deeper phenotyping through Rare Diseases Translational Research Collaboration (RD-TRC). RD-TRC has been setup by the National Institute for Health and Care Research and its aim is to provide research infrastructure that harnesses the strength of the NHS to support discoveries and translational research on rare diseases. These data are expected to pave the way for functional characterisation of findings, thereby adding further value to datasets, improving diagnostic utility and possibly identifying new targets and therapies. • Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The use of the Data could help to achieve the following benefits: • New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. > Through the international coalition of research intellects known as the Genomics England Clinical Interpretation Partnership (GeCIP) and the Discovery Forum, the framework for Genomics England to work with Industry: • Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. • Create a mechanism for research to continually improve the accuracy and reliability of information fed back to patients • Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. • Add to knowledge of the genetic basis of disease • Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. • Increase opportunities for clinical trials • Build the evidence base to accelerate the introduction of new technologies into healthcare > Stimulate and enhance UK industry and investment > Provide access to this unique research data resource to industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices > Attract inward investment from life science companies, with an aim of increasing opportunities of access to medicines that would otherwise be unavailable to UK patients > Result in new scientific insights and discoveries > Information linked to continually updated with long-term patient health and personal information to aid analysis by researchers. > Increase public knowledge and support for genomic medicine by delivering an ethical and transparent programme, retaining patient and public trust and confidence. This is aided by work with a range of partners to increase knowledge of genomics. Specific example: NHS Genomic Medicine Service (NHS GMS): Accelerate the uptake of genomic medicine in the NHS and other healthcare systems: > Work with NHS England and other partners to deliver a scalable NHS GMS including all genomic tests and an informatics platform to enable these services to be made widely available for NHS patients. > Work with Devolved Nation administrations, where appropriate, to achieve the same aims of the NHS GMS. Within the GMS, expected benefits for specific cohorts: (a) The Rare Diseases cohort: > Increase discovery of pathogenic variants for rare disease > Add value with additional biological insights that build confidence in putative pathogenic variants > Enhance the clinical interpretation of WGS in rare disease > Develop a programme of functional pathways for other genomic tests (i.e. not whole genome sequencing), specifically transcriptomics, epigenetics, micro RNAs and biomarkers > Return findings to the NHS for feedback to patients > Create a unique dataset for rare diseases that may enable therapeutic innovation (b) The Cancer cohort: > Use WGS to identify novel driver mutations for cancer and to understand its evolutionary genetic architecture through primary and secondary malignant disease > Partner stratified healthcare programmes and outcome studies with patients from the NHS in England, to enable understanding of WGS benefits in defining predictors of therapeutic response to cancer therapies > To use approaches using other genomic tests to offer additional biological insights into cancer > To utilise WGS to identify new pathways for cancer therapies and improved diagnostic characterisation. The expected patient benefit is to provide clinical diagnosis, and in time, new or more effective treatments for NHS patients. The discovery of new causes of disease, the offer of tailored therapies to create the best outcomes, and the priming new or more effective treatments for NHS patients, are other expected patient benefits. Specific Example: Generations Study: Within the Generations Study, Genomics England intend to use NHSE Data to provide evidence to answer the questions set out above for evaluation purposes. Specifically, HES and CSDS will be used for the following: • Cost effectiveness To approximate the true costs associated with additional / fewer =healthcare encounters of WGS in newborns, to include A&E attendances, outpatient appointments, admissions, allied health professional appointments, procedures, medication and treatment for all participants that screen positive through the Generation Study. • Health related outcomes To estimate the impact of WGS in newborns on the following, for all participants that screen positive through the Generation Study: o Diagnostic Odyssey to include i) time from first clinical contact to diagnosis, ii) age at diagnosis, iii) frequency and duration of health encounters during the diagnostic period. o Health encounters such as A&E attendances, outpatient appointments, admissions, allied health professional appointments over a defined time period. o Interventions for example procedures, medication and treatment over a defined time period. o Mortality to include i) age at death, and ii) cause of death. • Demographics monitoring To ensure enrolled participants are representative of the wider English population based on a variety of demographic variables. • Impacts on the NHS To identify if there has been any impact on the uptake of existing NHS newborn screening amongst participants of the Generation Study Genomics England also hope to link NHSE Data to genomic data for discovery research purposes. HES, CSDS and MSDS are expected to prove useful for researchers seeking opportunities to improve their understanding of rare disease, to develop new treatments, and diagnoses, and better understand how genes affect health.

Benefits reported

A number of benefits have already been yielded from the use of the NHS England Data held under this Agreement: [5 paragraphs unchanged]

Objective for processing

Amendments included in version (v14) are not permitted under this Data Sharing Agreement. V14 permits ongoing processing for the purposes described in v13. V14 is a three year agreement, however data will only to be released for the next six months. This is a pragmatic approach to enable Genomics England to continue accessing data until proposed amendments in v14 are approved.

Summary of processing permitted under v13:

> Processing was permitted for the 100,000 genomes project which was designed to create a new Genomic Medicine Service (GMS) for the NHS. Genomics England also continued to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

> The following datasets were accessed: Hospital Episode Statistics: Outpatients (HES OP), HES Admitted Patient Care (APC), HES Outpatients (OP), HES Critical Care (CC), HES Accident & Emergency (A&E), Diagnostic Imaging Dataset (DID), Mental Health Minimum Data Set (MHMDS), Mental Health and Learning Disabilities Data Set (MHLDDS), Mental Health Services Data Set (MHSDS), Patient Reported Outcome Measures (PROMs), Civil Registrations of Death and Demographics. Demographics and PROMs data is no longer required and has been destroyed.

> All Genomics research analysis was carried out within the Genomics England Trusted Research Environment (TRE).> Genomics England were permitted to share pseudonymised data under sublicense, including commercially through its Discovery Forum via their TRE.

> The territory of use was worldwide.

> NCRAS and uncurated NCRAS are also available within the TRE to link to the NHS England datasets held under this Agreement using the Study_ID.

> Genomics England is the controller on this Agreement.

> AWS and Lifebit are processors on this Agreement. They provide the technical infrastructure to host the research environment that data is loaded into, and researchers’ access.

> The following special condition was included: “The data controller must remove all those participants from the cohort who have been recruited using the Consultee process, before submitting the cohort to NHS England for data.”

SUMMARY OF CHANGES BETWEEN THE V13 AND V14 DSA.

THE CHANGES BELOW WILL NOT BE PERMITTED UNDER THIS V14 DSA BUT WILL BE INCLUDED IN A FUTURE AMENDMENT:

> Section 5 has been re-written to put the focus on the data being used broadly within the National Genomics Research Library (NGRL), formerly known as the Genomics TRE.

> Detail of the GMS has been added as an example use of the NGRL data.

> The Generations study has been added as an example use of the NGRL data, and with it, identifiers for this new cohort of newborns and their mothers is to be sent to NHS England for linkage.

> Additional consent materials for the NGRL and Generations study have been provided to NHS England for review and inclusion. This is to widen the cohort from only participants of the 100,000 genomes project to also include participants recruited through the GMS or via the Generations study.

> Additional datasets (MSDS and CSDS) have been requested for the Generations cohort only. (to be disseminated under a sister DSA (DARS-NIC-733503-V0X9Q)

***

Genomics England requires access to NHS England Data for use in their National Genomics Research Library (NGRL), which operates as a Trusted Research Environment (TRE).

National Genomic Research Library (NGRL)

The NGRL is a secure national resource of genomic, health and sample data managed by Genomics England, which builds on the research environment created by Genomics England for the 100,000 Genomes Project that completed recruitment in 2018 and involved the sequencing of approximately 100,000 genomes. It contains cohorts of patients/participants recruited via programmes set up for the NHS Genomic Medicine Service (GMS), the 100,000 Genomes Project, and other studies/programmes where patients/participants genomes have been sequenced, and provides a national standardised genomic research resource. Being able to compare all patient data in one place provides researchers with an opportunity to better understand diseases, develop new treatments and can lead to new discoveries. Only research using cohorts covered by this Agreement for the purposes specified can process the NHS England Data.

The following are examples of the aims of the use of data within the NGRL provided by Genomics England:

(1) NHS Genomic Medicine Service (NHS GMS)

Following the successfully delivery of the 100,000 Genomes Project, Genomics England begun work to deliver the NHS GMS. This service will offer patients the dual opportunity of routine clinical care alongside a choice to participate in research spanning across all genomic tests within the NHS, starting with whole genomes. Development of the NHS GMS builds on the evidence generated by the 100,000 Genomes Project, but also extends to other genomic testing other than Whole Genome Sequencing (WGS).

The NHS GMS is led and commissioned by NHS England and will consist of:

> NHS Genomic Laboratory Hubs (GLHs) that will work as part of a National Genomic testing service. The provisions in this service will be determined by a national genomic test directory that outlines the testing strategies and technology to be employed for rare and inherited disease, cancer and other defined conditions/ applications.

> Clinical Genetic Services

> Cancer services using genomic analysis to guide treatment

Within the GMS, two particular cohorts are defined:

(a) The Rare Diseases cohort

The NGRL rare diseases cohort will consist of families with rare diseases, based on family structures appropriate to the provision of the GMS. This will harness the strength of current UK rare disease programmes and will advance current understanding of rare disease mechanisms. It may also impact on common diseases that share similar phenotypes. It will also offer opportunities for biomarker, clinical, and interventional studies through industrial partnerships. Genomics England will actively support the implementation of the UK Rare Disease Strategy. Building on work already undertaken by the 100,000 Genomes Project, this will facilitate the generation of a national data resource of all genomic data, with a focus on WGS but including all genomic testing.

The goals of processing data of patients with rare disease are:

• To increase discovery of pathogenic variants (the gene variant responsible for causing disease) for rare disease.

• To add value with additional biological insights that build confidence in commonly accepted pathogenic variants.

• To enhance the clinical interpretation of WGS in rare disease.

• To develop a programme of functional pathways for genomic tests other than whole genome sequencing, specifically, transcriptomics (the study of all the ribonucleic acid (RNA) molecules within a cell, otherwise known as the transcriptome), epigenetics (the study of how cells control gene activity without changing the DNA sequence), micro RNAs and biomarkers.

• To return findings to the NHS for feedback to patients.

• To create a unique dataset for rare diseases that may enable therapeutic innovation.

(b) The Cancer cohort

The NGRL will continue to learn from and collaborate with other projects who are producing an inventory of genomic, transcriptomic and epigenomic changes in a wide range of different tumour types. Researchers from these projects access the data via an approved sublicensing agreement with Genomics England.

The goals of processing data for patients with cancer are to:

• Use WGS to identify novel driver mutations for cancer and to understand its evolutionary genetic architecture through primary and secondary malignant disease (by multiple biopsy and WGS).

• Partner stratified healthcare programmes and outcome studies with patients from the NHS in England, to enable understanding of WGS benefits in defining predictors of therapeutic response to cancer therapies.

• To use other genomic testing approaches to offer additional biological insights into cancer.

• To utilise WGS to identify new pathways for cancer therapies and improved diagnostic characterisation.

Other forms of collaboration include direct partnerships to develop tools to improve systems and services. For example, partnering with Lifebit to take advantage of their technical genomic data tooling.

(2) Generations Study

Every day, nine babies in the UK are born with a rare genetic condition that could be treated, prevented, or even cured if only it had been diagnosed when those babies were newborns. The Generation Study is aiming to find out if this situation can be improved by recruiting 100,000 newborn babies at NHS Trusts throughout England and conducting WGS to screen for over 200 rare genetic conditions that are treatable in early childhood. It is hoped that babies affected by these conditions will be identified more quickly, treated earlier and therefore have improved clinical outcomes.

The 3 research questions that Genomics England hope to answer are:

1. Can we better diagnose, and therefore care for, children with rare diseases?

2. Can we give researchers opportunities to improve their understanding of rare disease, to develop new treatments, and diagnoses, and better understand how our genes affect our health?

3. Should we, and if so how should we, use a baby’s genome throughout their lifetime as a resource they and their doctors can use if, for example, they become ill as they get older?

Additional processing for the Generation Study involves but is not limited to:

> Identifying discrepancies between WGS interpretation results and Bloodspot data results from the CSDS Dataset. This will be conducted on AWS cloud via an automated system developed by Genomics England. Discrepancies will be reported in the cases interpretation portal to be made available to the treating clinician.

> Evaluation of the cost effectiveness, health related outcomes, demographics monitoring and impact on the NHS. Further details of the outputs are detailed below.

The following NHS England Data will be accessed:

> Hospital Episode Statistics (HES) - these Datasets provide the core clinical data for participants and are vital to the provision of a detailed longitudinal medical history for participants. Specifically, the following HES Datasets are required:

- Admitted Patient Care (APC)

- Accident & Emergency (A&E)

- Critical Care (CC)

- Outpatients (OP)

> Emergency Care Data Set (ECDS) – necessary to understand which patients attend emergency departments and what treatment they receive in order to assess if there are associations with genetic markers.

> Diagnostic Imaging Dataset (DID) – necessary to provide invaluable, detailed information to build on participants’ phenotypes (observable characteristics), e.g., tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

> Mental Health – necessary because the GMS includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are expected to be some of the most prevalent conditions within the Project. For the 100,000 genomes project approximately 10% of project participants have a mental health record. Mental health Data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

> Civil Registration Mortality – necessary because deaths Data is essential for performing survival analyses; this is crucial information for research in combination with other medical history. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

> Cancer Registration – necessary because it provides a clinical background for the patient's cancer journey which may help provide scientific and medical insights when compared against the patient's genome.

> MRIS Events Mortality and Demographic Data - still required for open-ended research, specifically for researchers which have already been using this Data and need to go back to it. Losing this Dataset may cause disruption to their research.

> Community Services Data Set (CSDS) – necessary to evaluate the cost effectiveness of the Generation Study, to estimate the impact of WGS in newborns, and identify discrepancies in diseases identified between WGS interpretation results and the CSDS Dataset, which will be reported in the interpretation portal to be made available to the treating clinician.

>Maternity Services Data Set (MSDS) – necessary for discovery research purposes.

The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records, such as blood pressure, cholesterol, glucose, and pharmacogenomics (a field of research that studies how a person's genes affect how he or she responds to medications.), adds significant value. The richness of the NGRL datasets will allow Genomics England to move beyond the primary phenotype of the rare disease, cancer or infectious disease that led to the patient’s initial WGS in the context of other continuous traits, diseases and response to therapy including harm.

The level of the Data will be:

> Identifiable

For the following Datasets: HES APC; HES OP; HES A&E; ECDS; Cancer Registration Data; Civil Registrations of Death, many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. To ensure patients won’t be identified the researchers and their projects are examined before access is granted to the Data, and an agreement to not re-identify patients is signed. Further Genomics England monitors all Data leaving the NGRL and will not allow patient records to be exported.

The Data will be minimised as follows:

> Limited to a study cohort identified by Genomics England, comprising of patients recruited via (1) the 100,000 genomes project who also consented to longitudinal research (~80,000 participants), (2) the Genomics Medicines Service (expected ~100,000 new additions per year) or (3) the Generations Project, which follows newborn babies (expected ~50,000 new additions per year - recruitment is due to begin in Dec 2023 and will continue until 100,000 participants have been recruited in 2025).

> Genomics England will request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.

Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data.

NHS England has commissioned Genomics England to undertake the work. NHS England does not specify what Data are required to deliver the work nor how the Data shall be processed to achieve that purpose. Such decisions are taken by Genomics England.

The lawful basis for processing personal data under the UK GDPR is:

> Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.

Genomics England has determined the processing is necessary for its legitimate interests in carrying out medical research on the causes, diagnosis and treatment of rare diseases and cancers.

The lawful basis for processing special category data under the UK GDPR is:

> Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

It is necessary for Genomics England to process special category participant data for carrying out medical research on the causes, diagnosis and treatment of rare diseases and cancers, which is expected to benefit patients.

The funding is provided by the Department of Health and Social Care. The funding is specifically for the projects described. Funding is in place until March 2025, with the intention to renew this funding periodically.

The funder will have no ability to suppress or otherwise limit the publication of findings.

Lifebit provides IT support to Genomics England.

Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.

Representatives from patient and public bodies have an important role to play in Genomics England commercial initiatives. These representatives ensure transparency is upheld, and the interest of those whose data is being used is always being respected.

In the early stages of the Library, Genomics England undertook a range of work to ensure that potential participant’s views were included in the formulation of the ethical policies submitted for research ethics approval and in the development of patient information. The views of different groups of potential participants (those affected by cancer, rare disease, and those from BAME communities) in relation to ethical issues raised by the 100,000 Genomes Project were sought and findings were published on the Genomics England website (See all reports under ‘patient and public involvement - https://www.genomicsengland.co.uk/library-and-resources/ and the Genomics England Engagement Strategy). Genomics England will continue to engage with these stakeholders. Further to this, each of the 13 currently recruiting NHS Genomic Medicine Centres had dedicated Patient and Public Involvement leads (PPI) who are responsible for engaging with and involving local potential participant groups from diverse backgrounds. It is expected that the future NHS GMS will continue these local PPI activities to shape and inform the service.

A Participant Panel has also been established. This 30-strong group has provided invaluable advice on a range of topics, for instance, in shaping how analysis is monitored, how results are returned, and how advice and support should be framed. Participant Panel members have either donated samples to the Library themselves or are carers of participants. They take part in a wide variety of consultative groups, such as the Genomics England Ethics Advisory Committee but most importantly are guardians of the dataset, with representatives on the Access Review Committee. Participants play an important part in every decision made about access to data.

SUB-LICENCING:

Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.

It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.

The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.

All requests for data access will be subject to the following considerations:

• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).

• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).

• Provision of a signed Genomics England data access agreement to the Access Review Committee.

• Prioritisation of access according to resource availability.

• Facilitation of high-quality health research

Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:

• Diagnostic discovery

• Pre-clinical research

• Clinical Trials Referral

• Real World Evidence / Market Access

Approval process for Commercial organisations for access to the NGRL:

Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.

The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.

Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.

Approval process for GeCIP users (academic) of the NGRL:

• Researcher visits Genomics England website to enrol as a GECIP member

• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)

• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.

• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.

• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.

• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.

All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:

• Identification of diagnoses originally missed by the standardised pipeline

• Feedback of new diagnoses to patients

• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone

Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.

Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.

AIRLOCK POLICY:

The following rules are applied to all airlock requests:

1. All relevant details of the summary results to be transferred must be provided with every request.

2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.

3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.

4. Summary results requested for transfer are assessed using the following criteria:

a. whether the request aligns with the users ARC approval in full;

b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;

c. any data security implications;

d. any disclosure risks;

e. the technical feasibility and associated cost of the request;

f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;

g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.

The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:

• Technical Lead

• User Community Representative

• Bioinformatics Director

• Caldicott Guardian

• Chief Scientist representative

The Data will be processed worldwide.

Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.

GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:

• UK academic research institutions (e.g., universities, research institutions etc.)

• NHS trusts or authorities

• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project

• Foreign universities and research institutions that carry out significant research activity

• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))

• Foreign healthcare organisations (private or public) that undertake significant research activity

To be eligible for data access as a GeCIP member, applicants must meet these requirements:

• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.

• Their host institution has verified that they are affiliated with that institution.

• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).

• The GeCIP domain lead has approved the application.

• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.

• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.

All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).

The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.

GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.

All personnel accessing the Data have been appropriately trained in data protection and confidentiality.

The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:

> National Cancer Registration and Analysis Service (NCRAS)

and uncurated NCRAS data

> Secure Anonymised Information Linkage (SAIL) data; Welsh data

> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)

> NHS Trusts data

The Data will not be linked with any other data.

The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.

To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.

Data accessed under sub licence will not be re-identified.

Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.

Data shared through the Airlock process is aggregate data only and is therefore not personal data so does not require a legal basis under the UK GDPR.

A release register detailing any sub licences and onward sharin

Expected output

Researchers will have their own dissemination and communication strategies, however a full list of scientific publications and conferences/posters will be made available on the Genomics England website on an ongoing basis. The expected outputs of the processing will be:

> Submissions to peer reviewed journals

> Presentations at conferences

> Posters

> Creation of a database of all genomic data, including all genomic and omics tests.

The outputs will not contain NHS England Data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the Dataset(s) from which the information was derived.

The outputs will be communicated to relevant recipients through the following dissemination channels:

> Journals

> Posters

> Website: A list of publications is kept up-to-date on the Genomics website: https://www.genomicsengland.co.uk/research/publications?

> Presentations at appropriate conferences

> Upload of findings onto the ‘Discovery Forum’: Genomics England works with industry partners through the Discovery Forum. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Additionally, it allows the NGRL users to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed. These partners act as a critical friend and have already made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's NGRL.

Benefits reported

A number of benefits have already been yielded from the use of the NHS England Data held under this Agreement:

• Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The project found that using WGS led to a new diagnosis for 25% of the participants. Of these new diagnoses, 14% found variations in regions of the genome that would be missed by other methods, including other types of non-whole genomic tests.

• New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. This has enhanced genomic healthcare research by creating the largest genomic healthcare data resource in the world, which in turn will uncover answers for participants both now and in the future through genomic-level analysis of conditions.

• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. This acceleration significantly contributed to delivering the Genomics Medicine Service (GMS) for the NHS, which makes whole genome sequencing part of routine healthcare.

• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. After involving participants in all stages of the pioneering 100,000 Genomes Project and putting a trusted system in place contributed to a major dialogue led by Ipsos MORI and commissioned by Genomics England and co-funded by UK Research and Innovation’s Sciencewise programme in 2019 (post 100,000 genomes project completion) found the public are enthusiastic and optimistic about the potential for genomic medicine.

DARS-NIC-12784-R8W7V-v13.4 1 August 2023 to 31 October 2023
Title
Genomics England (MR1418) - Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
22
Files released
8

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v12.5

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v12.5
FieldWasBecame
Start date2023-02-012023-08-01
End date2023-07-312023-10-31
Demographics: legal basisHealth and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS DigitalHealth and Social Care Act 2012 – s261(2)(c)

Datasets: − Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset

Unchanged: Objective for processing, Processing activities, Expected output, Expected measurable benefits, Benefits reported.

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics.

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, through the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to participants: for many they are receiving a diagnosis for the first time.

The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness.

The richness of the high-quality data sets is crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project completed recruitment of rare disease participants and cancer patients in early 2019.

Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to access the following datasets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that Genomics' data applications are correct, do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their pseudonymised genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing participant data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of participants.

The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment)

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

Expected output

Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the TRE to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. A full list of scientific publications and conferences/posters can be found on the Genomics England website. Recent specific examples of journal outputs are as follows:

- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. Alexander J.M. Blakes, Htoo Wai, Ian Davies, et al. Genome Medicine Jul 2022 DOI: 10.1186/s13073-022-01087-x. This demonstrated the clinical value of examining non-canonical splicing variants in individuals with unsolved rare diseases by identifying 35 new likely diagnoses from 258 probands with an unsolved rare disease.

- Prognosis and oncogenomic profiling of patients with tropomyosin receptor kinase fusion cancer in the 100,000 genomes project. John Bridgewater, Xiaolong Jiao, Mounika Parimi, et al. Cancer Treatment and Research Communications Aug 2022 DOI: 10.1016/j.ctarc.2022.100623. This study supports the hypothesis that Neurotrophic tyrosine receptor kinase (NTRK) gene fusions are primary oncogenic drivers and the co-occurrence of NTRK gene fusions with other oncogenic alterations is rare. Overall survival was not statistically significant between NTRK and non-NTRK tumour patients.

- Participant experiences of genome sequencing for rare diseases in the 100,000 Genomes Project: a mixed methods study. Michelle Peter, Jennifer Hammond, Saskia C. Sanderson, et al. European Journal of Human Genetics Mar 2022 DOI: https://doi.org/10.1038/s41431-022-01065-2. This study highlights how expectation-setting is critical when offering genome sequencing, and that post-test counselling is important regardless of the GS result received, with parents perhaps needing additional emotional support.

Benefits reported

• Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The project found that using WGS led to a new diagnosis for 25% of the participants. Of these new diagnoses, 14% found variations in regions of the genome that would be missed by other methods, including other types of non-whole genomic tests.

• New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. This has enhanced genomic healthcare research by creating the largest genomic healthcare data resource in the world, which in turn will uncover answers for participants both now and in the future through genomic-level analysis of conditions.

• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. This acceleration significantly contributed to delivering the Genomics Medicine Service (GMS) for the NHS, which makes whole genome sequencing part of routine healthcare.

• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. After involving participants in all stages of the pioneering 100,000 Genomes Project and putting a trusted system in place contributed to a major dialogue led by Ipsos MORI and commissioned by Genomics England and co-funded by UK Research and Innovation’s Sciencewise programme in 2019 (post 100,000 genomes project completion) found the public are enthusiastic and optimistic about the potential for genomic medicine.

DARS-NIC-12784-R8W7V-v12.5 1 February 2023 to 31 July 2023
Title
Genomics England (MR1418) - Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
23
Files released
185

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v11.2

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v11.2
FieldWasBecame
Start date2022-10-242023-02-01
End date2023-10-122023-07-31

Objective for processing

[9 paragraphs unchanged] To achieve these goals Genomics wish to continue to retain access the following datasets: [19 paragraphs unchanged] This Agreement also covers the use of the data for other purposes by Genomics England, subject to a valid Data Sharing Agreement being in place. The data held under this Agreement must not be used for other purposes until this has been agreed by NHS Digital under another Data Sharing Agreement.

Processing activities

[3 paragraphs unchanged] Following this, the data are pseudonymised, as all subsequent processing can be [13 words unchanged] fields for each data set in line with details provided by NHS Digital England and following internal review of data sets. Pseudonymisation is a key facet [23 words unchanged] basis, where they are linked to participant genomes and primary clinical data. [2 paragraphs unchanged] Genomics England provided NHS Digital with a cohort for linkage. No additional data will flow under this Agreement (v11). Genomics England provided NHS England with a cohort for linkage. Data updates are disseminated on a regular, ongoing basis to enable analysis of the latest available data. [18 paragraphs unchanged] NCRAS and uncurated NCRAS are also available within the TRE to link to the NHS England datasets held under this Agreement using the Study_ID. No other clinical non-NHS England data is linked to data held in the TRE under this Agreement. [32 paragraphs unchanged] Genomics England in its provision of whole genome sequencing are applying to NHS Digital England for secondary clinical data to link to the genomic data. With regards to data provided by NHS Digital, England, Genomics England are the sole data controller. Genomics England provides NHS Digital England with linking data in order to receive longitudinal data sets. These data sets are delivered to Genomics England by NHS Digital England on a quarterly basis having been approved by the Independent Group Advising on the Release of Data (IGARD). Genomics England identifies the linking data and agrees with NHS Digital England the scope of the longitudinal data being provided. Genomics England determines the [13 words unchanged] use by approved researchers only. Genomics England determines who these researchers are. [2 paragraphs unchanged] Access to pseudonymised data in the TRE which will include longitudinal data sets (HES etc) provided by NHS Digital. England. Access to the TRE only allowed under access agreement. [21 paragraphs unchanged] There will be no data linkage undertaken with NHS Digital England data provided under this agreement that is not already noted in the agreement. [3 paragraphs unchanged]

Unchanged: Expected output, Expected measurable benefits, Benefits reported.

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics.

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, through the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to participants: for many they are receiving a diagnosis for the first time.

The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness.

The richness of the high-quality data sets is crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project completed recruitment of rare disease participants and cancer patients in early 2019.

Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to access the following datasets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that Genomics' data applications are correct, do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their pseudonymised genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing participant data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of participants.

The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment)

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

Expected output

Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the TRE to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. A full list of scientific publications and conferences/posters can be found on the Genomics England website. Recent specific examples of journal outputs are as follows:

- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. Alexander J.M. Blakes, Htoo Wai, Ian Davies, et al. Genome Medicine Jul 2022 DOI: 10.1186/s13073-022-01087-x. This demonstrated the clinical value of examining non-canonical splicing variants in individuals with unsolved rare diseases by identifying 35 new likely diagnoses from 258 probands with an unsolved rare disease.

- Prognosis and oncogenomic profiling of patients with tropomyosin receptor kinase fusion cancer in the 100,000 genomes project. John Bridgewater, Xiaolong Jiao, Mounika Parimi, et al. Cancer Treatment and Research Communications Aug 2022 DOI: 10.1016/j.ctarc.2022.100623. This study supports the hypothesis that Neurotrophic tyrosine receptor kinase (NTRK) gene fusions are primary oncogenic drivers and the co-occurrence of NTRK gene fusions with other oncogenic alterations is rare. Overall survival was not statistically significant between NTRK and non-NTRK tumour patients.

- Participant experiences of genome sequencing for rare diseases in the 100,000 Genomes Project: a mixed methods study. Michelle Peter, Jennifer Hammond, Saskia C. Sanderson, et al. European Journal of Human Genetics Mar 2022 DOI: https://doi.org/10.1038/s41431-022-01065-2. This study highlights how expectation-setting is critical when offering genome sequencing, and that post-test counselling is important regardless of the GS result received, with parents perhaps needing additional emotional support.

Benefits reported

• Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The project found that using WGS led to a new diagnosis for 25% of the participants. Of these new diagnoses, 14% found variations in regions of the genome that would be missed by other methods, including other types of non-whole genomic tests.

• New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. This has enhanced genomic healthcare research by creating the largest genomic healthcare data resource in the world, which in turn will uncover answers for participants both now and in the future through genomic-level analysis of conditions.

• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. This acceleration significantly contributed to delivering the Genomics Medicine Service (GMS) for the NHS, which makes whole genome sequencing part of routine healthcare.

• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. After involving participants in all stages of the pioneering 100,000 Genomes Project and putting a trusted system in place contributed to a major dialogue led by Ipsos MORI and commissioned by Genomics England and co-funded by UK Research and Innovation’s Sciencewise programme in 2019 (post 100,000 genomes project completion) found the public are enthusiastic and optimistic about the potential for genomic medicine.

DARS-NIC-12784-R8W7V-v11.2 24 October 2022 to 12 October 2023
Title
Genomics England (MR1418) - Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
23
Files released
0

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v10.2

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v10.2
FieldWasBecame
TitleGenomics England (MR1418) -Renewal Request for tranche of data across multiple data sets.Genomics England (MR1418) - Renewal Request for tranche of data across multiple data sets.
Start date2022-07-132022-10-24
End date2022-10-122023-10-12

Objective for processing

[7 paragraphs unchanged] The 100,000 Genomes Project has completed recruitment of rare disease participants and cancer patients in early 2019. [1 paragraph unchanged] To achieve these goals Genomics wish to continue to collect retain the following datasets: [19 paragraphs unchanged] To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service (GMS). Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS. This Agreement also covers the use of the data for other purposes by Genomics England, subject to a valid Data Sharing Agreement being in place. The data held under this Agreement must not be used for other purposes until this has been agreed by NHS Digital under another Data Sharing Agreement.

Processing activities

[2 paragraphs unchanged] The first stage of processing focuses on quality verification. This ensures that [24 words unchanged] Genomics England’s participant details and any updates required to identifiable data fields, e.g. e.g., dates of birth, are highlighted. Finally Finally, the data set is reviewed against recent participant withdrawals so that any withdrawals notified after the data application was made can be removed from the data sets. Following this this, the data are pseudonymised, as all subsequent processing can be performed without [59 words unchanged] basis, where they are linked to participant genomes and primary clinical data. The second stage of processing involves the selection of a pseudonymised cohort [32 words unchanged] included in the approved use purposes set out in the Genomics England Protocol, Protocol and fall within the scope of the relevant GECiP or the Discovery [9 words unchanged] into the TRE and any tools they wish to use for analysis. [1 paragraph unchanged] Genomics England provide NHS Digital with a cohort for linkage and they receive data from NHS Digital on a monthly basis. Every quarter Genomics provide an updated cohort to NHS Digital who provide the historical data for the extra cohort members. The cohort is already flagged with NHS Digital so Genomics will only receive the historical data for the extra cohort members each quarter. Genomics England provided NHS Digital with a cohort for linkage. No additional data will flow under this Agreement (v11). [66 paragraphs unchanged] Microsoft Limited: Genomics England use Cloud based processing and therefore will use Microsoft Azure Cloud as a processing location. Microsoft UK provide Cloud Services for Genomics England and are therefore listed as a data processor. They supply support to the system, but do not access data. Therefore, any access to the data held under this Agreement would be considered a breach of the Agreement. This includes granting of access to the database containing the data. [14 paragraphs unchanged]

Expected output

The project was established to sequence 100,000 genomes from around 85,000 NHS patients affected by a rare disease, or cancer. Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the TRE to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. A full list of scientific publications and conferences/posters can be found on the Genomics England website. Recent specific examples of journal outputs are as follows: The Project would also create a new genomic medicine service for the NHS – transforming the way people are cared for and bringing advanced diagnosis and personalised treatments to all those who need them. - A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. Alexander J.M. Blakes, Htoo Wai, Ian Davies, et al. Genome Medicine Jul 2022 DOI: 10.1186/s13073-022-01087-x. This demonstrated the clinical value of examining non-canonical splicing variants in individuals with unsolved rare diseases by identifying 35 new likely diagnoses from 258 probands with an unsolved rare disease. Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem. - Prognosis and oncogenomic profiling of patients with tropomyosin receptor kinase fusion cancer in the 100,000 genomes project. John Bridgewater, Xiaolong Jiao, Mounika Parimi, et al. Cancer Treatment and Research Communications Aug 2022 DOI: 10.1016/j.ctarc.2022.100623. This study supports the hypothesis that Neurotrophic tyrosine receptor kinase (NTRK) gene fusions are primary oncogenic drivers and the co-occurrence of NTRK gene fusions with other oncogenic alterations is rare. Overall survival was not statistically significant between NTRK and non-NTRK tumour patients. Recruitment of participants to the 100,000 Genomes Project was completed in 2018, with the 100,000th sequence achieved in December 2018. Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Trusted Research Environment (TRE). - Participant experiences of genome sequencing for rare diseases in the 100,000 Genomes Project: a mixed methods study. Michelle Peter, Jennifer Hammond, Saskia C. Sanderson, et al. European Journal of Human Genetics Mar 2022 DOI: https://doi.org/10.1038/s41431-022-01065-2. This study highlights how expectation-setting is critical when offering genome sequencing, and that post-test counselling is important regardless of the GS result received, with parents perhaps needing additional emotional support. Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the TRE on the dates shown above.

Unchanged: Expected measurable benefits, Benefits reported.

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics.

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, through the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to participants: for many they are receiving a diagnosis for the first time.

The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness.

The richness of the high-quality data sets is crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project completed recruitment of rare disease participants and cancer patients in early 2019.

Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to retain the following datasets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that Genomics' data applications are correct, do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their pseudonymised genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing participant data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of participants.

The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment)

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

This Agreement also covers the use of the data for other purposes by Genomics England, subject to a valid Data Sharing Agreement being in place. The data held under this Agreement must not be used for other purposes until this has been agreed by NHS Digital under another Data Sharing Agreement.

Expected output

Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the TRE to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. A full list of scientific publications and conferences/posters can be found on the Genomics England website. Recent specific examples of journal outputs are as follows:

- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. Alexander J.M. Blakes, Htoo Wai, Ian Davies, et al. Genome Medicine Jul 2022 DOI: 10.1186/s13073-022-01087-x. This demonstrated the clinical value of examining non-canonical splicing variants in individuals with unsolved rare diseases by identifying 35 new likely diagnoses from 258 probands with an unsolved rare disease.

- Prognosis and oncogenomic profiling of patients with tropomyosin receptor kinase fusion cancer in the 100,000 genomes project. John Bridgewater, Xiaolong Jiao, Mounika Parimi, et al. Cancer Treatment and Research Communications Aug 2022 DOI: 10.1016/j.ctarc.2022.100623. This study supports the hypothesis that Neurotrophic tyrosine receptor kinase (NTRK) gene fusions are primary oncogenic drivers and the co-occurrence of NTRK gene fusions with other oncogenic alterations is rare. Overall survival was not statistically significant between NTRK and non-NTRK tumour patients.

- Participant experiences of genome sequencing for rare diseases in the 100,000 Genomes Project: a mixed methods study. Michelle Peter, Jennifer Hammond, Saskia C. Sanderson, et al. European Journal of Human Genetics Mar 2022 DOI: https://doi.org/10.1038/s41431-022-01065-2. This study highlights how expectation-setting is critical when offering genome sequencing, and that post-test counselling is important regardless of the GS result received, with parents perhaps needing additional emotional support.

Benefits reported

• Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The project found that using WGS led to a new diagnosis for 25% of the participants. Of these new diagnoses, 14% found variations in regions of the genome that would be missed by other methods, including other types of non-whole genomic tests.

• New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. This has enhanced genomic healthcare research by creating the largest genomic healthcare data resource in the world, which in turn will uncover answers for participants both now and in the future through genomic-level analysis of conditions.

• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. This acceleration significantly contributed to delivering the Genomics Medicine Service (GMS) for the NHS, which makes whole genome sequencing part of routine healthcare.

• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. After involving participants in all stages of the pioneering 100,000 Genomes Project and putting a trusted system in place contributed to a major dialogue led by Ipsos MORI and commissioned by Genomics England and co-funded by UK Research and Innovation’s Sciencewise programme in 2019 (post 100,000 genomes project completion) found the public are enthusiastic and optimistic about the potential for genomic medicine.

DARS-NIC-12784-R8W7V-v10.2 13 July 2022 to 12 October 2022
Title
Genomics England (MR1418) -Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
23
Files released
305

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v9.6

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v9.6
FieldWasBecame
Start date2022-04-132022-07-13
End date2022-07-122022-10-12

Processing activities

[1 paragraph unchanged] PROCESSING ACTIVITIES [1 paragraph unchanged] Following this the data are pseudonymised, as all subsequent processing can be [38 words unchanged] of the Genomics England resource. Pseudonymised data are uploaded to a secure trusted research environment (TRE) hosted by Genomics England on a quarterly basis, where they are linked to participant genomes and primary clinical data. The second stage of processing involves the selection of a pseudonymised cohort [56 words unchanged] Discovery Forum. Researchers declare any data they wish to bring into the research environment TRE and any tools they wish to use for analysis. The third stage of processing is the analysis of the pseudonymised data sets within the research environment. TRE. Researchers perform all the analysis and processing within the environment: they do [5 words unchanged] data are placed in a secure folder for anonymisation verification before extraction. [1 paragraph unchanged] The Research Environment: THE TRUSTED RESEARCH ENVIRONMENT (TRE) All research analysis on the Genomics England data-set will only be carried [5 words unchanged] environment hosted within the Genomics England data centre – the Genomics England Research Environment. TRE. Analytical tools and applications are available within the Research Environment. TRE. No sequencing or clinical data are made available for download, users cannot copy or paste out of the Research Environment, TRE, and there is limited internet access within it (i.e. whitelisted sites). Movement of files into and out of the Research Environment TRE is governed via an ‘Airlock’ Policy. Academic researcher access to the Research Environment Academic researchers access the TRE by applying to be a member of a GeCIP domain. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following: Academic researchers access the Research Environment by applying to be a member of a GeCIP domain. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following: [15 paragraphs unchanged] COMMERCIAL RESEARCHER ACCESS TO THE RESEARCH ENVIRONMENT. TRE Genomics England operates a membership-based forum – the Discovery Forum – which is open to a range of companies world-wide and allows access to the Research Environment. TRE. It provides a platform for collaboration between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Each Discovery Forum member signs a Data Access Agreement with Genomics England. [74 words unchanged] is in place, each research project undertaken by the Company within the Research Environment TRE must receive prior Access Review Committee (ARC) approval. Discovery Forum members access the Research Environment TRE in a similar manner to GeCIP Researchers: all research is carried out within the Research Environment, TRE, and any movement of results out of the environment occurs only through the Airlock Process. [3 paragraphs unchanged] The Genomics England Research Environment TRE has been developed with the intention that all data analysis is carried out within it and that the only data to leave it are analytical summary results. An Airlock process has been established which enables material (data, files, tools etc) to be moved in or out of the Research Environment TRE in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access. Removal of summary results therefore requires an Airlock request. [1 paragraph unchanged] 1. All relevant details of the files summary results to be transferred must be provided with every request. 2. All files summary results transferred may must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any files summary results rejected along with the reason for the rejection. 3. All files summary results transferred will be checked for viruses and malware and those failing this [9 words unchanged] the requestors to resolve such issues before re-submitting the file for transfer. 4. Files Summary results requested for transfer are assessed using the following criteria: • whether the request aligns with the user’s users ARC approval in full • whether the request can clearly be demonstrated to be aligned with a registered project in the Research Environment TRE [3 paragraphs unchanged] • when importing data, its scientific value to the community of researchers within the Research Environment, TRE, and when and how it will be shared [2 paragraphs unchanged] Analysed results are inspected to ensure they cannot be used to disclose the identity of the participant. participants. Checking of statistical output summary results by the Airlock Review Team is governed by a generalizable set of principles that guide individual decisions and ensure flexible evaluation of the Genomics England dataset. decisions. By using a principles-based approach where each case is assessed individually the security of the dataset is maintained by exporting only ‘safe’ appropriate data. Review of transfer requests resulting in public-sharing/publication of data will be [7 words unchanged] can only be used for the specific use detailed in the original export. request. The TRE contains pseudonymised longitudinal data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be used and what can be exported. Where an export contains pseudonymised longitudinal data, Genomics England will always apply the requirements placed on them as conditions of having access to the data. The Research Environment contains External Data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be used and what can be exported. Where an export contains External Data, Genomics England will always apply the requirements placed on them as conditions of having access to the data. In some cases, particularly concerning the export of record-level data, these will be more conservative than those applied to 100,000 Genomes Data alone. All Airlock requests go through a robust approval process and the Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. Where a precedent has been set and a clear set of principles and rules are in place for types of research, the Airlock Manager can approve the request. For more complicated requests or where no precedent has been set these will go to the Airlock Committee and be reviewed by the Airlock Review Team. The Airlock Review Team is a delegation of the Genomics England Chief Scientist responsible for oversight of all airlock requests in accordance with the Airlock Policy and the group’s groups Terms of Reference. It comprises: • Senior Information Risk Office (SIRO) [6 paragraphs unchanged] Genomics England has developed the Research Environment TRE to allow registered third parties to access pseudonymised versions of the data that it holds, for the purposes of approved research. The Research Environment TRE contains External Data (for example Hospital Episodes Statistics [HES]) which is subject [30 words unchanged] placed on them as holders of External Data to users of the Research Environment TRE as a condition of having access to the data. The data is NOT for onward sharing outside of the Research Environment. TRE. Data control: DATA CONTROLLER [1 paragraph unchanged] Genomics England provides NHS Digital with linking data in order to receive [50 words unchanged] provided. Genomics England determines the method of pseudonymisation and storage within the Research Environment TRE and secures this data for use by approved researchers only. Genomics England determines who these researchers are. Genomics England is the Data Controller for longitudinal data sets processed in the Genomics England Research Library. TRE. [1 paragraph unchanged] Access to pseudonymised data in the Research Environment TRE which will include longitudinal data sets (HES etc) provided by NHS Digital. Access to the Research Environment TRE only allowed under access agreement. [1 paragraph unchanged] Data Processors: DATA PROCESSORS Genomics England have ceased using UKCloud (UKC) as a data processor after migrating all data to Amazon Web Services (AWS). The data in UKC is planned to be destroyed in March 2022. A detailed plan has been created which will ensure the deletion runbook will be followed, which is aligned to Genomics England standard operating procedures. [9 paragraphs unchanged] Data Minimisation: Microsoft Limited: Genomics England use Cloud based processing and therefore will use Microsoft Azure Cloud as a processing location. Microsoft UK provide Cloud Services for Genomics England and are therefore listed as a data processor. They supply support to the system, but do not access data. Therefore, any access to the data held under this Agreement would be considered a breach of the Agreement. This includes granting of access to the database containing the data. DATA MINIMISATION Genomics England’s TRE aligns to the current NHS guidance and is currently cited as an example of best practice. The detail to support this is below. In the HDR UK TRE Principles and Best Practices paper from December 2021, the Genomics England model is explicitly described (page 17) and this document has a foreword authored by the Director of Data Policy, NHSX and Director of Tech Policy, NHSX. https://www.hdruk.ac.uk/news/new-principles-published-to-improve-public-confidence-in-access-and-use-of-data-for-health-research-through-trusted-research-environments. Conceptually, in terms of the 5 safes framework, the safes work together to protect the privacy of the individual. The TRE model makes 4 of the safes: safe-setting, safe-people, safe-projects and safe-outputs very strong, meaning that it’s possible to allow the criteria for safe-data to be relaxed (i.e., de-identification only) while maintaining the overall level of privacy protection. The TRE model effectively moves data minimisation to the output stage - ’safe outputs’, i.e., the airlock, where the airlock managers and the airlock committee do review requests to ensure summary data is minimised to only what is necessary to demonstrate externally a particular research result. The data released by the Airlock project is aggregated summary data. Given these measures, there has even been a discussion about data within TREs being regarded as "functionally anonymous” because of these safeguards, which would put them outside the constraints of GDPR. The risk is further minimised as all the researchers are under contractual obligation not to re-identify individuals from the data. By allowing researchers access to all the data it allows "hypothesis free research" and means they may discover things they didn't suspect. The full dataset is “adequate and relevant” for all researchers to access and the TRE framework provides the required limitation. Research in the RE is around discovering genome phenome relationships and given the complexity of the genome and how little we understand about it, it’s impossible to predict what phenotype terms may be involved, so researchers need to be able to explore all of them, rather than repeatedly requesting different datasets. Each individual participant has agreed and is aware that their data is in the TRE and accessed by researchers. Access to the data is necessary for this purpose to ensure the research can be carried out and the most value extracted from the data. Genomics England do have a wide range of controls outlined above to make sure the risk is minimised. Both Article 89 and Article 25(1) of the UK GDPR support this approach. [3 paragraphs unchanged] PROMS data is only available for non-commercial purposes, such as academic research, or in connection with delivering services to the NHS NHS. A Data Protection Impact Assessment was conducted in November 2021 prior to the transfer of the data to the new environment and is being continually assessed and updated by the data protection team.

Expected output

[3 paragraphs unchanged] Recruitment of participants to the 100,000 Genomes Project was completed in 2018, [15 words unchanged] life-long clinical data from the participants and making these available in the Trusted Research Environment. Environment (TRE). Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the Research Environment TRE on the dates shown above.

Benefits reported

The clinical benefit the 100,000 project brought to the participants acted as evidence to support the NHS to commit to whole genome sequencing as a standard of care for rare disease and cancer clinical indications: https://www.england.nhs.uk/genomics/nhs-genomic-med-service/ • Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The project found that using WGS led to a new diagnosis for 25% of the participants. Of these new diagnoses, 14% found variations in regions of the genome that would be missed by other methods, including other types of non-whole genomic tests. The yielded research benefits are being published on the Genomics England website’s publication page: https://www.genomicsengland.co.uk/about-gecip/publications/ and conference research outputs page: https://www.genomicsengland.co.uk/about-gecip/publications/further-research-outputs/ • New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. This has enhanced genomic healthcare research by creating the largest genomic healthcare data resource in the world, which in turn will uncover answers for participants both now and in the future through genomic-level analysis of conditions. Further Genomics England has built upon its commitment to lead on Government’s technology and innovation agenda by forging partnership with industry. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners have a variety of outcomes measures and not all focus on publication for confirmation. We do however invite industry partners to present their work at conferences and keep up-to-date with their progress through our Strategic partnership directors. • Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. This acceleration significantly contributed to delivering the Genomics Medicine Service (GMS) for the NHS, which makes whole genome sequencing part of routine healthcare. • Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. • Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. After involving participants in all stages of the pioneering 100,000 Genomes Project and putting a trusted system in place contributed to a major dialogue led by Ipsos MORI and commissioned by Genomics England and co-funded by UK Research and Innovation’s Sciencewise programme in 2019 (post 100,000 genomes project completion) found the public are enthusiastic and optimistic about the potential for genomic medicine.

Unchanged: Objective for processing, Expected measurable benefits.

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics.

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, through the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to participants: for many they are receiving a diagnosis for the first time.

The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness.

The richness of the high-quality data sets is crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project has completed recruitment of rare disease participants and cancer patients in early 2019.

Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to collect the following datasets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that Genomics' data applications are correct, do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their pseudonymised genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing participant data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of participants.

The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment)

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service (GMS). Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS.

Expected output

The project was established to sequence 100,000 genomes from around 85,000 NHS patients affected by a rare disease, or cancer.

The Project would also create a new genomic medicine service for the NHS – transforming the way people are cared for and bringing advanced diagnosis and personalised treatments to all those who need them.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

Recruitment of participants to the 100,000 Genomes Project was completed in 2018, with the 100,000th sequence achieved in December 2018. Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Trusted Research Environment (TRE).

Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the TRE on the dates shown above.

Benefits reported

• Patient benefit: providing clinical diagnosis and in time, new or more effective treatments for NHS patients. The project found that using WGS led to a new diagnosis for 25% of the participants. Of these new diagnoses, 14% found variations in regions of the genome that would be missed by other methods, including other types of non-whole genomic tests.

• New scientific insights and discovery: with the consent of patients, creating a database of 100,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers. This has enhanced genomic healthcare research by creating the largest genomic healthcare data resource in the world, which in turn will uncover answers for participants both now and in the future through genomic-level analysis of conditions.

• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. In addition, through the Genomics England Clinical Interpretation Partnership (GeCIP), creating a mechanism to both continually improve the accuracy and reliability of information fed back to patients and add to knowledge of the genetic basis of disease. This acceleration significantly contributed to delivering the Genomics Medicine Service (GMS) for the NHS, which makes whole genome sequencing part of routine healthcare.

• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics. After involving participants in all stages of the pioneering 100,000 Genomes Project and putting a trusted system in place contributed to a major dialogue led by Ipsos MORI and commissioned by Genomics England and co-funded by UK Research and Innovation’s Sciencewise programme in 2019 (post 100,000 genomes project completion) found the public are enthusiastic and optimistic about the potential for genomic medicine.

DARS-NIC-12784-R8W7V-v9.6 13 April 2022 to 12 July 2022
Title
Genomics England (MR1418) -Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
23
Files released
3

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v8.6

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v8.6
FieldWasBecame
Start date2021-04-012022-04-13
End date2022-03-312022-07-12
Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Cancer Registration Data: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Civil Registrations of Death: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Demographics: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Diagnostic Imaging Data Set (DID): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Emergency Care Data Set (ECDS): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
HES-ID to MPS-ID HES Accident and Emergency: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
HES-ID to MPS-ID HES Admitted Patient Care: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
HES-ID to MPS-ID HES Outpatients: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Hospital Episode Statistics Accident and Emergency (HES A and E): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Hospital Episode Statistics Admitted Patient Care (HES APC): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Hospital Episode Statistics Critical Care (HES Critical Care): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Hospital Episode Statistics Outpatients (HES OP): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
MRIS - Cause of Death Report: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
MRIS - Cohort Event Notification Report: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
MRIS - Flagging Current Status Report: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
MRIS - List Cleaning Report: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
MRIS - Members and Postings Report: legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Mental Health Minimum Data Set (MHMDS): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Mental Health Services Data Set (MHSDS): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Mental Health and Learning Disabilities Data Set (MHLDDS): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital
Patient Reported Outcome Measures (Linkable to HES): legal basisHealth and Social Care Act 2012 – s261(2)(c)Health and Social Care Act 2012 – s261(2)(c); Informed Patient consent to permit the receipt, processing and release of data by NHS Digital

Objective for processing

[3 paragraphs unchanged] The project was designed to create a new genomic medicine service for [7 words unchanged] create new capability for the clinical genomics research, both academic and industrial, though through the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure. [2 paragraphs unchanged] The richness of the high-quality data sets are is crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of whole genome sequencing (WGS) WGS data in the context of rich and extended phenotypes derived from electronic [31 words unchanged] in the context of other continuous traits, diseases and responses to therapy. [2 paragraphs unchanged] To achieve these goals Genomics wish to continue to collect the following data-sets: datasets: [5 paragraphs unchanged] Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that our Genomics' data applications are correct and correct, do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate. [1 paragraph unchanged] As the Discovery Forum is a collaborative venture, no fees are levied [35 words unchanged] been asked explicitly to give consent for commercial companies to access their de-identified pseudonymised genome and health data. [6 paragraphs unchanged] Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees. [4 paragraphs unchanged] Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees. To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service (GMS). Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS. The beneficiaries are: o Participants - through the work Genomics England do will ultimately influence their care; o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data; o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK. To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service. Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS.

Processing activities

[2 paragraphs unchanged] Following this the data are de-identified, pseudonymised, as all subsequent processing can be performed without direct identifiers. Genomics England [6 words unchanged] sensitive fields for each data set in line with details provided by the NHS Digital and following internal review of data sets. De-identification Pseudonymisation is a key facet of the Genomics England resource. De-identified Pseudonymised data are uploaded to a secure research environment hosted by Genomics England on a quarterly basis, where they are linked to participant genomes and primary clinical data. The second stage of processing involves the selection of a de-identified pseudonymised cohort of participants that fulfil a specific research request. Researchers are members [56 words unchanged] the research environment and any tools they wish to use for analysis. The third stage of processing is the analysis of the de-identified pseudonymised data sets within the research environment. Researchers perform all the analysis and processing within the environment: they do not extract de-identified pseudonymised data. Results data are placed in a secure folder for anonymisation verification before extraction. Genomics England provide NHS Digital with a cohort for linkage and they [16 words unchanged] to NHS Digital who provide the historical data for the extra cohort members members. The cohort is already flagged with NHS Digital so Genomics will only receive the historical data for the extra cohort members each quarter. The Research Environment (TRE): Environment: [2 paragraphs unchanged] Academic researchers access the Research Environment by applying to be a member of a Genomics England Clinical Interpretation Partnership (GeCIP) GeCIP domain. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following: [17 paragraphs unchanged] Each Discovery Forum member signs a Data Access Agreement with Genomics England. [79 words unchanged] project undertaken by the Company within the Research Environment must receive prior ARC Access Review Committee (ARC) approval. [2 paragraphs unchanged] The Access Review Committee (ARC) ARC provides an independent examination of requests for data access, with regards to [41 words unchanged] made up of participants and parents/carers involved in the 100,000 genomes project. [7 paragraphs unchanged] • whether the request aligns with the user’s Access Review Committee (ARC) ARC approval [8 paragraphs unchanged] The Research Environment contains External Data (for example Hospital Episodes Statistics [HES]) [51 words unchanged] access to the data. In some cases, particularly concerning the export of individual-level record-level data, these will be more conservative than those applied to 100,000 Genomes Data alone. [11 paragraphs unchanged] Genomics England provides NHS Digital with linking data in order to receive [11 words unchanged] by NHS Digital on a quarterly basis having been approved by the NHS Digital IGARD. Independent Group Advising on the Release of Data (IGARD). Genomics England identifies the linking data and agrees with NHS Digital the scope of the longitudinal data being provided. Genomics England determines the method of de-identification pseudonymisation and storage within the research environment Research Environment and secures this data for use by approved researchers only. Genomics England determines who these researchers are. [2 paragraphs unchanged] Access to deidentified pseudonymised data in the research environment Research Environment which will include longitudinal data sets (HES etc) provided by NHS Digital. Access to the Research Environment only allowed under access agreement. [2 paragraphs unchanged] UK Cloud, Amazon Web Services and Lifebit are additional data processors on this agreement. They provide the technical infrastructure to host the research environment that data is loaded into, and researchers access. Genomics England have ceased using UKCloud (UKC) as a data processor after migrating all data to Amazon Web Services (AWS). The data in UKC is planned to be destroyed in March 2022. A detailed plan has been created which will ensure the deletion runbook will be followed, which is aligned to Genomics England standard operating procedures. AWS and Lifebit are additional data processors on this agreement. They provide the technical infrastructure to host the research environment that data is loaded into, and researchers access. [3 paragraphs unchanged] o CloudOS will be hosted within GEL’s London AWS (Amazon Web Services) environment - All data is encrypted in transit and at rest. [9 paragraphs unchanged] An updated A Data Protection Impact Assessment will be was conducted in November 2021 prior to the transfer of the data to the new environment. environment and is being continually assessed by the data protection team.

Expected output

Genomics England’s target was to complete sequencing of 100,000 genomes by the end of 2018 – this was achieved by December 2018 - - https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000-genomes-sequenced/ The project was established to sequence 100,000 genomes from around 85,000 NHS patients affected by a rare disease, or cancer. By April 2017 over 41,000 whole genomes had been sequenced and, with the aid of clinical data, the first diagnoses were announced: The Project would also create a new genomic medicine service for the NHS – transforming the way people are cared for and bringing advanced diagnosis and personalised treatments to all those who need them. • First diagnoses from the pilot phase https://www.genomicsengland.co.uk/first-patients-diagnosed-through-the-100000-genomes-project/ Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem. • First children diagnosed through the project https://www.genomicsengland.co.uk/first-children-recieve-diagnoses-through-100000-genomes-project/ Recruitment of participants to the 100,000 Genomes Project was completed in 2018, with the 100,000th sequence achieved in December 2018. Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment. • Financial Times article: https://www.ft.com/content/d2e21cea-d684-11e6-944b-e7eb37a6aa8e By November 2018 109,611 samples had been collected, 91,100 genomes sequenced and 37,591 results returned to referring NHS Genomic Medicine Centres. Recruitment to the 100,000 genomes project is now completed, but analysis and return of results will continue throughout the period of this agreement. By May 2018 the Genomics England Research Environment was established to allow research access to de-identified genome and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with over 1,300 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. By the end of October 2018 the fifth quarterly release of new data into the Genomics England Research Environment was achieved. This release included 71,860 genomes and primary clinical data for 85,070 participants. It also included de-identified clinical data received from NHS Digital for 60,110 participants, totalling 4.2m records. 1,821 GeCIP members from 387 institutions and 33 research domains now requested and been given access to the Research Environment, and 70 research projects have been approved. Alongside this there are 89 members of the industry Discovery Forum, including 10 full members. Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital, into the Research Environment on the following dates, in order to continue support for, and to further develop, this ground-breaking resource: • 20th Aug 2020 • 20th Oct 2020 • 20th Dec 2020 Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment. [1 paragraph unchanged] All outputs will contain only data that is aggregated with small numbers suppressed in line with the HES Analysis Guide. (if using HES data)

Benefits reported

December 2017: The clinical benefit the 100,000 project brought to the participants acted as evidence to support the NHS to commit to whole genome sequencing as a standard of care for rare disease and cancer clinical indications: https://www.england.nhs.uk/genomics/nhs-genomic-med-service/ Over 41,000 Genomes were sequenced as of December 2017. Participant stories can be found at: https://www.genomicsengland.co.uk/alexs-story/ The yielded research benefits are being published on the Genomics England website’s publication page: https://www.genomicsengland.co.uk/about-gecip/publications/ and conference research outputs page: https://www.genomicsengland.co.uk/about-gecip/publications/further-research-outputs/ Further Genomics England has built upon its commitment to lead on Governments Government’s technology and innovation agenda by forging partnership with industry. Examples The creation of this included the Discovery Forum provides a new platform for collaboration and engagement between Genomics England, industry collaboration partners, academia, the NHS and the wider UK genomics landscape. Industry partners have a variety of outcomes measures and not all focus on publication for confirmation. We do however invite industry partners to present their work at conferences and keep up-to-date with leading life sciences companies Inivata and Thermo Fisher Scientific to improve understanding of cancer. their progress through our Strategic partnership directors. Public Health England has announced that Whole Genome Sequencing (WGS) is now being used to identify different strains of tuberculosis (TB). This is the first time that WGS has been used as a diagnostic solution for managing a disease on this scale anywhere in the world. The technique, developed in conjunction with the University of Oxford, allows faster and more accurate diagnoses, meaning patients can be treated with precisely the right medication more quickly. Genomics England has now engaged devolved nations and is recruiting participants from Scotland and Wales. May 2018: Over 60,000 genomes have been sequenced and over 12,000 clinical reports have been issued to referring NHS Genomic Medicine Centres. Thirty disease and cross-cutting research domains have had their plans approved and now have access to 100,000 the Genomes Project Research Environment. The number of users with access to the Genomics England Research Environment is now over 1,300. Twelve publications have arisen from or refer to the 100,000 Genomes Project during the last year, including: • The 100,000 Genomes Project: bringing whole genome sequencing to the NHS. Clare Turnbull et al. BMJ 2018; doi: https://doi.org/10.1136/bmj.k1687 (24 April 2018) • Identification of rare sequence variation underlying heritable pulmonary arterial hypertension. Nicholas W. Morrell et al. Nature Communications 2018;9; doi:10.1038/s41467-018-03672-4 (12 April 2018) • Introducing genomics into cancer care. Sue Hill BRJ Surg 2018;105(2):e14-e15 (17 January 2018) • Missense variants in the X-linked gene PRPS1 cause retinal degeneration in females. Alessia Fiorentino, Kaoru Fujinami, Gavin Arno et al. Hum Mutat 2017; doi:10.1002/humu.23349 (17 October 2017) See https://www.genomicsengland.co.uk/category/updates/ and https://www.genomicsengland.co.uk/aboutgecip/ publications/ for details of news and publications. Genomics England created the Discovery Forum in July 2017. It provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. November 2018: Over 109,000 samples have now been collected, 91,100 genomes sequenced and 37,591 clinical reports issued to referring NHS Genomic Medicine Centres. 71,860 genomes have been made available in the Genomics Engalnd Research Environment. 1,821 GeCIP members from 387 institutions and 33 research domains now have access to the research environment. There are now 89 members of the industry Discovery Forum, including 10 full members. Latest publications include: • Challenges in implementing genomic medicine: the 100,000 Genomes Project. Julian G. Barwell, Rory B.G. O’Sullivan, Laura K. Mansbridge, Joanna M. Lowry, Huw R. Dorkins. J Transl Genet Genom 2018;2:13; https://doi.org/10.20517/jtgg.2018.17 (11 Sep 2018) • What will follow the first hundred thousand genomes in the NHS? Malcolm Grant & John Paul Maytum. Per Med 2018; doi:10.2217/pme-2018-0025 (20 June 2018) • Clinical-grade validation of whole genome sequencing reveals robust detection of low-frequency variants and copy number alterations in CLL. Jenny Klintman, Katerina Barmpouti, Samantha JL Knight et al. Br J Haematol 2018; doi:10.1111/bjh.15406 (29 May 2018) February 2019: • Clinical and genetic variability in children with partial albinism. Campbell P, Ellingford JM, Parry NRA, et al. Nature Scientific Reports doi: 10.1038/s41598-019-51768-8 (12/11/2019) • Secondary C1q Deficiency in Activated PI3Kδ Syndrome Type 2 Ying Hong, Sira Nanthapisal, Ebun Omoyinmi, et al. Frontiers in Immunologydoi: 10.3389/fimmu.2019.02589 (11/11/2019) • Defective tubulin detyrosination causes structural brain abnormalities with cognitive deficiency in humans and mice Pagnamenta, Alistair T; Heemeryck, Pierre; Hilary C Martin, et al. Human Molecular Genetics doi: 10.1093/hmg/ddz186 (31/07/2019) • Germline selection shapes the landscape of human mitochondrial DNA Wei, Wei; Chinnery, Patrick; et al Science doi: 10.1126/science.aau6520 (01/04/2019) • Validating Whole Genome Sequencing (WGS) for clinical use in AML & ALL S. Henderson, A. Sosinsky, A. Hamblin, et al. British Journal of Haematology doi: 10.1111/bjh.15854 (27/03/2019) • Opportunities and Challenges for Molecular Understanding of Ciliopathies–The 100,000 Genomes Project Wheway, Gabrielle; Mitchison, Hannah H. Frontiers in Genetics doi: 10.3389/fgene.2019.00569 (11/03/2019) Whole-genome sequencing of rare disease patients in a national healthcare system Ouwehand, Willem H bioRxiv – pre-publication doi: 10.1101/507244 (01/01/2019) April 2020: • Genomic loci susceptible to systematic sequencing bias in clinical whole genomes Timothy M. Freeman, Genomics England Research Consortium, Dennis Wang, and Jason Harris Genome Research doi: 10.1101/gr.255349.119 (11/03/2020) • Mutational signature in colorectal cancer caused by genotoxic pks+ E. coli Cayetano Pleguezuelos-Manzano, Jens Puschhof, Axel Rosendahl Huber, et al. Nature doi: 10.1038/s41586-020-2080-8 (27/02/2020) • Expanding Clinical Presentations Due to Variations in THOC2 mRNA Nuclear Export Factor Raman Kumar, Elizabeth Palmer, Alison E. Gardner, et al. Frontiers in Molecular Neuroscience doi: 10.3389/fnmol.2020.00012 (11/02/2020) • Human and mouse essentiality screens as a resource for disease gene discovery Pilar Cacheiro, Violeta Muñoz-Fuentes, Stephen A. Murray, et al. Nature Communications doi: 10.1038/s41467-020-14284-2 (31/01/2020) • A restricted spectrum of missense KMT2D variants cause a multiple malformations disorder distinct from Kabuki syndrome Sara Cuvertino, Verity Hartill, Alice Colyer, et al. Genetics in Medicine doi: 10.1038/s41436-019-0743-3 (17/01/2020) One of the main outputs of the project was to provide a diverse and unique clinical dataset alongside genomic data to ensure continued learning and development from the already accrued knowledge. Continued data acquisition will allow researchers to follow clinical progression of disease and cancers. The dataset in terms of participants is largely stable, with no significant further additions. However, by migrating the data to a cloud based system, this allows further integration of data in developing cohort and genomic browser systems, allowing more in-depth real-time analysis The role of the 100,000 project was to gain an understanding of the rationale of whole genome sequencing for standard of care testing. The value of the 100,000 project now means that the NHS will provide whole genome sequencing as standard of care for the first time this year. The findings from the 100,000 project were fed back to clinicians to ensure a understanding of how to interpret genomic data and the subsequent translation of that into the concept of personalized medicine - as per the legitimate interests for Genomics England. The patient participant panel for the 100,000 Genomes project have supported grant and award nominations for Genomics England - celebrating their patient engagement.

Unchanged: Expected measurable benefits.

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics.

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, through the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to participants: for many they are receiving a diagnosis for the first time.

The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness.

The richness of the high-quality data sets is crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of WGS data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project has completed recruitment of rare disease participants and cancer patients in early 2019.

Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to collect the following datasets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that Genomics' data applications are correct, do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their pseudonymised genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing participant data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of participants.

The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment)

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service (GMS). Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS.

Expected output

The project was established to sequence 100,000 genomes from around 85,000 NHS patients affected by a rare disease, or cancer.

The Project would also create a new genomic medicine service for the NHS – transforming the way people are cared for and bringing advanced diagnosis and personalised treatments to all those who need them.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

Recruitment of participants to the 100,000 Genomes Project was completed in 2018, with the 100,000th sequence achieved in December 2018. Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.

Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the Research Environment on the dates shown above.

Benefits reported

The clinical benefit the 100,000 project brought to the participants acted as evidence to support the NHS to commit to whole genome sequencing as a standard of care for rare disease and cancer clinical indications: https://www.england.nhs.uk/genomics/nhs-genomic-med-service/

The yielded research benefits are being published on the Genomics England website’s publication page: https://www.genomicsengland.co.uk/about-gecip/publications/ and conference research outputs page: https://www.genomicsengland.co.uk/about-gecip/publications/further-research-outputs/

Further Genomics England has built upon its commitment to lead on Government’s technology and innovation agenda by forging partnership with industry. The creation of the Discovery Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners have a variety of outcomes measures and not all focus on publication for confirmation. We do however invite industry partners to present their work at conferences and keep up-to-date with their progress through our Strategic partnership directors.

DARS-NIC-12784-R8W7V-v8.6 1 April 2021 to 31 March 2022
Title
Genomics England (MR1418) -Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
23
Files released
681

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v7.4

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v7.4
FieldWasBecame
Start date2020-08-132021-04-01
End date2021-03-312022-03-31

Datasets: + HES-ID to MPS-ID HES Accident and Emergency; + HES-ID to MPS-ID HES Admitted Patient Care; + HES-ID to MPS-ID HES Outpatients

Unchanged: Objective for processing, Processing activities, Expected output, Expected measurable benefits, Benefits reported.

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics.

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, though the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to Participants: for many they are receiving a diagnosis for the first time.

The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness.

The richness of the high-quality data sets are crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of whole genome sequencing (WGS) data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project has completed recruitment of rare disease participants and cancer patients in early 2019.

Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to collect the following data-sets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that our data applications are correct and do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their de-identified genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.

The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment)

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees.

The beneficiaries are:

o Participants - through the work Genomics England do will ultimately influence their care;

o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.

To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service. Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS.

Expected output

Genomics England’s target was to complete sequencing of 100,000 genomes by the end of 2018 – this was achieved by December 2018 - - https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000-genomes-sequenced/

By April 2017 over 41,000 whole genomes had been sequenced and, with the aid of clinical data, the first diagnoses were announced:

• First diagnoses from the pilot phase https://www.genomicsengland.co.uk/first-patients-diagnosed-through-the-100000-genomes-project/

• First children diagnosed through the project https://www.genomicsengland.co.uk/first-children-recieve-diagnoses-through-100000-genomes-project/

• Financial Times article: https://www.ft.com/content/d2e21cea-d684-11e6-944b-e7eb37a6aa8e

By November 2018 109,611 samples had been collected, 91,100 genomes sequenced and 37,591 results returned to referring NHS Genomic Medicine Centres. Recruitment to the 100,000 genomes project is now completed, but analysis and return of results will continue throughout the period of this agreement.

By May 2018 the Genomics England Research Environment was established to allow research access to de-identified genome and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with over 1,300 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

By the end of October 2018 the fifth quarterly release of new data into the Genomics England Research Environment was achieved. This release included 71,860 genomes and primary clinical data for 85,070 participants. It also included de-identified clinical data received from NHS Digital for 60,110 participants, totalling 4.2m records. 1,821 GeCIP members from 387 institutions and 33 research domains now requested and been given access to the Research Environment, and 70 research projects have been approved. Alongside this there are 89 members of the industry Discovery Forum, including 10 full members.

Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital, into the Research Environment on the following dates, in order to continue support for, and to further develop, this ground-breaking resource:

• 20th Aug 2020

• 20th Oct 2020

• 20th Dec 2020

Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.

Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the Research Environment on the dates shown above.

All outputs will contain only data that is aggregated with small numbers suppressed in line with the HES Analysis Guide. (if using HES data)

Benefits reported

December 2017:

Over 41,000 Genomes were sequenced as of December 2017. Participant stories can be found at: https://www.genomicsengland.co.uk/alexs-story/

Genomics England has built upon its commitment to lead on Governments technology and innovation agenda by forging partnership with industry. Examples of this included a new industry collaboration with leading life sciences companies Inivata and Thermo Fisher Scientific to improve understanding of cancer.

Public Health England has announced that Whole Genome Sequencing (WGS) is now being used to identify different strains of tuberculosis (TB). This is the first time that WGS has been used as a diagnostic solution for managing a disease on this scale anywhere in the world. The technique, developed in conjunction with the University of Oxford, allows faster and more accurate diagnoses, meaning patients can be treated with precisely the right medication more quickly.

Genomics England has now engaged devolved nations and is recruiting participants from Scotland and Wales.

May 2018:

Over 60,000 genomes have been sequenced and over 12,000 clinical reports have been issued to referring NHS Genomic Medicine Centres. Thirty disease and cross-cutting research domains have had their plans approved and now have access to 100,000 the Genomes Project Research Environment. The number of users with access to the Genomics England Research Environment is now over 1,300.

Twelve publications have arisen from or refer to the 100,000 Genomes Project during the last year, including:

• The 100,000 Genomes Project: bringing whole genome sequencing to the NHS. Clare Turnbull et al. BMJ 2018; doi: https://doi.org/10.1136/bmj.k1687 (24 April 2018)

• Identification of rare sequence variation underlying heritable pulmonary arterial hypertension. Nicholas W. Morrell et al. Nature Communications 2018;9; doi:10.1038/s41467-018-03672-4 (12 April 2018)

• Introducing genomics into cancer care. Sue Hill BRJ Surg 2018;105(2):e14-e15 (17 January 2018)

• Missense variants in the X-linked gene PRPS1 cause retinal degeneration in females. Alessia Fiorentino, Kaoru Fujinami, Gavin Arno et al. Hum Mutat 2017; doi:10.1002/humu.23349 (17 October 2017)

See https://www.genomicsengland.co.uk/category/updates/ and https://www.genomicsengland.co.uk/aboutgecip/ publications/ for details of news and publications.

Genomics England created the Discovery Forum in July 2017. It provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

November 2018:

Over 109,000 samples have now been collected, 91,100 genomes sequenced and 37,591 clinical reports issued to referring NHS Genomic Medicine Centres. 71,860 genomes have been made available in the Genomics Engalnd Research Environment.

1,821 GeCIP members from 387 institutions and 33 research domains now have access to the research environment. There are now 89 members of the industry Discovery Forum, including 10 full members.

Latest publications include:

• Challenges in implementing genomic medicine: the 100,000 Genomes Project. Julian G. Barwell, Rory B.G. O’Sullivan, Laura K. Mansbridge, Joanna M. Lowry, Huw R. Dorkins. J Transl Genet Genom 2018;2:13; https://doi.org/10.20517/jtgg.2018.17 (11 Sep 2018)

• What will follow the first hundred thousand genomes in the NHS? Malcolm Grant & John Paul Maytum. Per Med 2018; doi:10.2217/pme-2018-0025 (20 June 2018)

• Clinical-grade validation of whole genome sequencing reveals robust detection of low-frequency variants and copy number alterations in CLL. Jenny Klintman, Katerina Barmpouti, Samantha JL Knight et al. Br J Haematol 2018; doi:10.1111/bjh.15406 (29 May 2018)

February 2019:

• Clinical and genetic variability in children with partial albinism. Campbell P, Ellingford JM, Parry NRA, et al. Nature Scientific Reports doi: 10.1038/s41598-019-51768-8 (12/11/2019)

• Secondary C1q Deficiency in Activated PI3Kδ Syndrome Type 2 Ying Hong, Sira Nanthapisal, Ebun Omoyinmi, et al. Frontiers in Immunologydoi: 10.3389/fimmu.2019.02589 (11/11/2019)

• Defective tubulin detyrosination causes structural brain abnormalities with cognitive deficiency in humans and mice Pagnamenta, Alistair T; Heemeryck, Pierre; Hilary C Martin, et al. Human Molecular Genetics doi: 10.1093/hmg/ddz186 (31/07/2019)

• Germline selection shapes the landscape of human mitochondrial DNA Wei, Wei; Chinnery, Patrick; et al Science doi: 10.1126/science.aau6520 (01/04/2019)

• Validating Whole Genome Sequencing (WGS) for clinical use in AML & ALL S. Henderson, A. Sosinsky, A. Hamblin, et al. British Journal of Haematology doi: 10.1111/bjh.15854 (27/03/2019)

• Opportunities and Challenges for Molecular Understanding of Ciliopathies–The 100,000 Genomes Project Wheway, Gabrielle; Mitchison, Hannah H. Frontiers in Genetics doi: 10.3389/fgene.2019.00569 (11/03/2019)

Whole-genome sequencing of rare disease patients in a national healthcare system Ouwehand, Willem H bioRxiv – pre-publication doi: 10.1101/507244 (01/01/2019)

April 2020:

• Genomic loci susceptible to systematic sequencing bias in clinical whole genomes Timothy M. Freeman, Genomics England Research Consortium, Dennis Wang, and Jason Harris Genome Research doi: 10.1101/gr.255349.119 (11/03/2020)

• Mutational signature in colorectal cancer caused by genotoxic pks+ E. coli Cayetano Pleguezuelos-Manzano, Jens Puschhof, Axel Rosendahl Huber, et al. Nature doi: 10.1038/s41586-020-2080-8 (27/02/2020)

• Expanding Clinical Presentations Due to Variations in THOC2 mRNA Nuclear Export Factor Raman Kumar, Elizabeth Palmer, Alison E. Gardner, et al. Frontiers in Molecular Neuroscience doi: 10.3389/fnmol.2020.00012 (11/02/2020)

• Human and mouse essentiality screens as a resource for disease gene discovery Pilar Cacheiro, Violeta Muñoz-Fuentes, Stephen A. Murray, et al. Nature Communications doi: 10.1038/s41467-020-14284-2 (31/01/2020)

• A restricted spectrum of missense KMT2D variants cause a multiple malformations disorder distinct from Kabuki syndrome Sara Cuvertino, Verity Hartill, Alice Colyer, et al. Genetics in Medicine doi: 10.1038/s41436-019-0743-3 (17/01/2020)

One of the main outputs of the project was to provide a diverse and unique clinical dataset alongside genomic data to ensure continued learning and development from the already accrued knowledge. Continued data acquisition will allow researchers to follow clinical progression of disease and cancers. The dataset in terms of participants is largely stable, with no significant further additions. However, by migrating the data to a cloud based system, this allows further integration of data in developing cohort and genomic browser systems, allowing more in-depth real-time analysis

The role of the 100,000 project was to gain an understanding of the rationale of whole genome sequencing for standard of care testing. The value of the 100,000 project now means that the NHS will provide whole genome sequencing as standard of care for the first time this year. The findings from the 100,000 project were fed back to clinicians to ensure a understanding of how to interpret genomic data and the subsequent translation of that into the concept of personalized medicine - as per the legitimate interests for Genomics England.

The patient participant panel for the 100,000 Genomes project have supported grant and award nominations for Genomics England - celebrating their patient engagement.

DARS-NIC-12784-R8W7V-v7.4 13 August 2020 to 31 March 2021
Title
Genomics England (MR1418) -Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
20
Files released
322

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v6.7

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v6.7
FieldWasBecame
Start date2020-04-012020-08-13

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics. [1 paragraph unchanged] Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem. [2 paragraphs unchanged] The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness. [1 paragraph unchanged] The 100,000 Genomes Project has completed recruitment of rare disease participants and will complete recruitment of cancer patients in early 2019. Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project. Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project. [5 paragraphs unchanged] MRIS: Cohort Event Notification, Cause of Death. Mortality data are essential for performing survival analyses and as a metric for success of medical care: analyses: this is crucial information for research in combination with other medical history. [36 words unchanged] analysis of medical timeline data and for the management of participant cohorts. MRIS: Members and Posting, List Cleaning Report, Flagging Current Status. Demographic Data: These identifiable data sets are vital for the safe and efficient management [23 words unchanged] participants are excluded from data collection, and that participants’ details are accurate. [3 paragraphs unchanged] The lawful basis for processing Participant Data under the General Data Protection The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants. Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants. The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants. [1 paragraph unchanged] · Creating a new genomic medicine service for the NHS; NHS (not a part of this agreement - will be part of a future amendment) [5 paragraphs unchanged] Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees. The beneficiaries are: o Participants - through the work Genomics England do will ultimately influence their care; o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data; o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK. To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service. Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS.

Processing activities

All organisations party to this agreement must comply with the Data Sharing Framework Contract requirements, including those regarding the use (and purposes of that use) by "Personnel" (as defined within the Data Sharing Framework Contract i.e.: employees, agents and contractors of the Data Recipient who may have access to that data). [5 paragraphs unchanged] The Research Environment (TRE): [18 paragraphs unchanged] Commercial researcher access to the research environment COMMERCIAL RESEARCHER ACCESS TO THE RESEARCH ENVIRONMENT. [29 paragraphs unchanged] Sub-licencing SUB LICENSING Genomics England has developed the Research Environment to allow registered third parties [41 words unchanged] Genomics England and other parties that dictate how the data may be used. accessed. Genomics England will always apply the requirements placed on them as holders [6 words unchanged] the Research Environment as a condition of having access to the data. The data is NOT for onward sharing outside of the Research Environment. UK Cloud Ltd is not being used for Cloud Storage. Data control: All organisations party to this agreement must comply with the Data Sharing Framework Contract requirements, including those regarding the use (and purposes of that use) by “Personnel” (as defined within the Data Sharing Framework Contract ie: employees, agents and contractors of the Data Recipient who may have access to that data). Genomics England in its provision of whole genome sequencing are applying to NHS Digital for secondary clinical data to link to the genomic data. With regards to data provided by NHS Digital, Genomics England are the sole data controller. Genomics England provides NHS Digital with linking data in order to receive longitudinal data sets. These data sets are delivered to Genomics England by NHS Digital on a quarterly basis having been approved by the NHS Digital IGARD. Genomics England identifies the linking data and agrees with NHS Digital the scope of the longitudinal data being provided. Genomics England determines the method of de-identification and storage within the research environment and secures this data for use by approved researchers only. Genomics England determines who these researchers are. Genomics England is the Data Controller for longitudinal data sets processed in the Genomics England Research Library. Researchers in academic, educational or commercial organisations Access to deidentified data in the research environment which will include longitudinal data sets (HES etc) provided by NHS Digital. Access to the Research Environment only allowed under access agreement. The individual researchers are Data Controllers when carrying out research within the research environment. Research which is approved using the data will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/ Data Processors: UK Cloud, Amazon Web Services and Lifebit are additional data processors on this agreement. They provide the technical infrastructure to host the research environment that data is loaded into, and researchers access. o Only summary level data can be removed from the environment. o Approved researchers will only be able to access Lifebit’s Platform as a service (PaaS) Cloud Operating System (OS) through a virtual desktop. o Secondary data will be ingested into CloudOS. o CloudOS will be hosted within GEL’s London AWS (Amazon Web Services) environment - All data is encrypted in transit and at rest. o CloudOS controls access to the secondary data. o A security and Data Protection Impact Assessment (DPIA) will be conducted prior to loading live Lifebit: Lifebit has been selected as platform partner to deliver the Research Environment after reviewing several proposals. The UK-based Subject Matter Expert (SME) offered a proven and innovative technology solution offering a blend of robustness and ease of use. Lifebit CloudOS provides a secure and collaborative workspace to enable researchers to easily perform genomic data analysis. The platform will deliver an intuitive, integrated and collaborative user experience and enables fast, effective research outcomes across a wide range of academic and biotech/pharma researchers with varying levels of technical competency. Data Minimisation: This will be limited to the selected cohort and additions and deletions will be updated regularly. [3 paragraphs unchanged] An updated Data Protection Impact Assessment will be conducted prior to the transfer of the data to the new environment.

Expected output

[9 paragraphs unchanged] • 28th Feb 19 • 20th Aug 2020 • 30th Apr 19 • 20th Oct 2020 • 31st Jul 19 • 20th Dec 2020 • 31st Oct 19 [3 paragraphs unchanged]

Benefits reported

Update December 2017 2017: [1 paragraph unchanged] Genomics England has built upon its commitment to lead on Governments technology and innovation agenda by forging partnership with industry. Examples of this include included a new industry collaboration with leading life sciences companies Inivata and Thermo Fisher Scientific to improve understanding of cancer. [2 paragraphs unchanged] Update May 2018 2018: [8 paragraphs unchanged] Update November 2018 2018: [6 paragraphs unchanged] Update 2019 February 2019: [7 paragraphs unchanged] Update Apr 2020 April 2020: [5 paragraphs unchanged] One of the main outputs of the project was to provide a diverse and unique clinical dataset alongside genomic data to ensure continued learning and development from the already accrued knowledge. Continued data acquisition will allow researchers to follow clinical progression of disease and cancers. The dataset in terms of participants is largely stable, with no significant further additions. However, by migrating the data to a cloud based system, this allows further integration of data in developing cohort and genomic browser systems, allowing more in-depth real-time analysis The role of the 100,000 project was to gain an understanding of the rationale of whole genome sequencing for standard of care testing. The value of the 100,000 project now means that the NHS will provide whole genome sequencing as standard of care for the first time this year. The findings from the 100,000 project were fed back to clinicians to ensure a understanding of how to interpret genomic data and the subsequent translation of that into the concept of personalized medicine - as per the legitimate interests for Genomics England. The patient participant panel for the 100,000 Genomes project have supported grant and award nominations for Genomics England - celebrating their patient engagement.

Unchanged: Expected measurable benefits.

Objective for processing

A genome is the body’s instruction manual - a copy is stored in almost every healthy cell in your body. The study of that genome and all the technologies needed to analyse and interpret it is called genomics.

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

Combining genomic sequence data with medical records has created a ground-breaking research resource. Researchers are currently studying how best to use genomics in healthcare and how best to interpret the data to help patients. The causes, diagnosis and treatment of disease is also being investigated. Revealing which variants cause disease is also helping companies find new targeted medicines. Kick-starting a UK genomics industry was another key aim of the project and the UK now has a vibrant genomics ecosystem.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, though the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to Participants: for many they are receiving a diagnosis for the first time.

The way in which the NHS is able to link a whole lifetime of medical records with a person’s genome data and the fact it can do this on a large scale is unique. The richness of this data can help to understand disease and to tease apart the complex relationship between our genes, what happens to us in our lives and illness.

The richness of the high-quality data sets are crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of whole genome sequencing (WGS) data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project has completed recruitment of rare disease participants and cancer patients in early 2019.

Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to collect the following data-sets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

Mortality data are essential for performing survival analyses: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

Demographic Data: These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that our data applications are correct and do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their de-identified genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is Legitimate Interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.

The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancers.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS (not a part of this agreement - will be part of a future amendment)

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

Patients and participants will be at the heart of this programme and include the participant panel, a 30 strong panel involved in Genomics England research committees.

The beneficiaries are:

o Participants - through the work Genomics England do will ultimately influence their care;

o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.

To confirm - this agreement does not cover (at present) the use of data for the Genomic Medicine Service. Genomics England will submit an amendment with relevant documentation so that the use of NHS Digital data can be approved for the GMS.

Expected output

Genomics England’s target was to complete sequencing of 100,000 genomes by the end of 2018 – this was achieved by December 2018 - - https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000-genomes-sequenced/

By April 2017 over 41,000 whole genomes had been sequenced and, with the aid of clinical data, the first diagnoses were announced:

• First diagnoses from the pilot phase https://www.genomicsengland.co.uk/first-patients-diagnosed-through-the-100000-genomes-project/

• First children diagnosed through the project https://www.genomicsengland.co.uk/first-children-recieve-diagnoses-through-100000-genomes-project/

• Financial Times article: https://www.ft.com/content/d2e21cea-d684-11e6-944b-e7eb37a6aa8e

By November 2018 109,611 samples had been collected, 91,100 genomes sequenced and 37,591 results returned to referring NHS Genomic Medicine Centres. Recruitment to the 100,000 genomes project is now completed, but analysis and return of results will continue throughout the period of this agreement.

By May 2018 the Genomics England Research Environment was established to allow research access to de-identified genome and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with over 1,300 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

By the end of October 2018 the fifth quarterly release of new data into the Genomics England Research Environment was achieved. This release included 71,860 genomes and primary clinical data for 85,070 participants. It also included de-identified clinical data received from NHS Digital for 60,110 participants, totalling 4.2m records. 1,821 GeCIP members from 387 institutions and 33 research domains now requested and been given access to the Research Environment, and 70 research projects have been approved. Alongside this there are 89 members of the industry Discovery Forum, including 10 full members.

Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital, into the Research Environment on the following dates, in order to continue support for, and to further develop, this ground-breaking resource:

• 20th Aug 2020

• 20th Oct 2020

• 20th Dec 2020

Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.

Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the Research Environment on the dates shown above.

All outputs will contain only data that is aggregated with small numbers suppressed in line with the HES Analysis Guide. (if using HES data)

Benefits reported

December 2017:

Over 41,000 Genomes were sequenced as of December 2017. Participant stories can be found at: https://www.genomicsengland.co.uk/alexs-story/

Genomics England has built upon its commitment to lead on Governments technology and innovation agenda by forging partnership with industry. Examples of this included a new industry collaboration with leading life sciences companies Inivata and Thermo Fisher Scientific to improve understanding of cancer.

Public Health England has announced that Whole Genome Sequencing (WGS) is now being used to identify different strains of tuberculosis (TB). This is the first time that WGS has been used as a diagnostic solution for managing a disease on this scale anywhere in the world. The technique, developed in conjunction with the University of Oxford, allows faster and more accurate diagnoses, meaning patients can be treated with precisely the right medication more quickly.

Genomics England has now engaged devolved nations and is recruiting participants from Scotland and Wales.

May 2018:

Over 60,000 genomes have been sequenced and over 12,000 clinical reports have been issued to referring NHS Genomic Medicine Centres. Thirty disease and cross-cutting research domains have had their plans approved and now have access to 100,000 the Genomes Project Research Environment. The number of users with access to the Genomics England Research Environment is now over 1,300.

Twelve publications have arisen from or refer to the 100,000 Genomes Project during the last year, including:

• The 100,000 Genomes Project: bringing whole genome sequencing to the NHS. Clare Turnbull et al. BMJ 2018; doi: https://doi.org/10.1136/bmj.k1687 (24 April 2018)

• Identification of rare sequence variation underlying heritable pulmonary arterial hypertension. Nicholas W. Morrell et al. Nature Communications 2018;9; doi:10.1038/s41467-018-03672-4 (12 April 2018)

• Introducing genomics into cancer care. Sue Hill BRJ Surg 2018;105(2):e14-e15 (17 January 2018)

• Missense variants in the X-linked gene PRPS1 cause retinal degeneration in females. Alessia Fiorentino, Kaoru Fujinami, Gavin Arno et al. Hum Mutat 2017; doi:10.1002/humu.23349 (17 October 2017)

See https://www.genomicsengland.co.uk/category/updates/ and https://www.genomicsengland.co.uk/aboutgecip/ publications/ for details of news and publications.

Genomics England created the Discovery Forum in July 2017. It provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

November 2018:

Over 109,000 samples have now been collected, 91,100 genomes sequenced and 37,591 clinical reports issued to referring NHS Genomic Medicine Centres. 71,860 genomes have been made available in the Genomics Engalnd Research Environment.

1,821 GeCIP members from 387 institutions and 33 research domains now have access to the research environment. There are now 89 members of the industry Discovery Forum, including 10 full members.

Latest publications include:

• Challenges in implementing genomic medicine: the 100,000 Genomes Project. Julian G. Barwell, Rory B.G. O’Sullivan, Laura K. Mansbridge, Joanna M. Lowry, Huw R. Dorkins. J Transl Genet Genom 2018;2:13; https://doi.org/10.20517/jtgg.2018.17 (11 Sep 2018)

• What will follow the first hundred thousand genomes in the NHS? Malcolm Grant & John Paul Maytum. Per Med 2018; doi:10.2217/pme-2018-0025 (20 June 2018)

• Clinical-grade validation of whole genome sequencing reveals robust detection of low-frequency variants and copy number alterations in CLL. Jenny Klintman, Katerina Barmpouti, Samantha JL Knight et al. Br J Haematol 2018; doi:10.1111/bjh.15406 (29 May 2018)

February 2019:

• Clinical and genetic variability in children with partial albinism. Campbell P, Ellingford JM, Parry NRA, et al. Nature Scientific Reports doi: 10.1038/s41598-019-51768-8 (12/11/2019)

• Secondary C1q Deficiency in Activated PI3Kδ Syndrome Type 2 Ying Hong, Sira Nanthapisal, Ebun Omoyinmi, et al. Frontiers in Immunologydoi: 10.3389/fimmu.2019.02589 (11/11/2019)

• Defective tubulin detyrosination causes structural brain abnormalities with cognitive deficiency in humans and mice Pagnamenta, Alistair T; Heemeryck, Pierre; Hilary C Martin, et al. Human Molecular Genetics doi: 10.1093/hmg/ddz186 (31/07/2019)

• Germline selection shapes the landscape of human mitochondrial DNA Wei, Wei; Chinnery, Patrick; et al Science doi: 10.1126/science.aau6520 (01/04/2019)

• Validating Whole Genome Sequencing (WGS) for clinical use in AML & ALL S. Henderson, A. Sosinsky, A. Hamblin, et al. British Journal of Haematology doi: 10.1111/bjh.15854 (27/03/2019)

• Opportunities and Challenges for Molecular Understanding of Ciliopathies–The 100,000 Genomes Project Wheway, Gabrielle; Mitchison, Hannah H. Frontiers in Genetics doi: 10.3389/fgene.2019.00569 (11/03/2019)

Whole-genome sequencing of rare disease patients in a national healthcare system Ouwehand, Willem H bioRxiv – pre-publication doi: 10.1101/507244 (01/01/2019)

April 2020:

• Genomic loci susceptible to systematic sequencing bias in clinical whole genomes Timothy M. Freeman, Genomics England Research Consortium, Dennis Wang, and Jason Harris Genome Research doi: 10.1101/gr.255349.119 (11/03/2020)

• Mutational signature in colorectal cancer caused by genotoxic pks+ E. coli Cayetano Pleguezuelos-Manzano, Jens Puschhof, Axel Rosendahl Huber, et al. Nature doi: 10.1038/s41586-020-2080-8 (27/02/2020)

• Expanding Clinical Presentations Due to Variations in THOC2 mRNA Nuclear Export Factor Raman Kumar, Elizabeth Palmer, Alison E. Gardner, et al. Frontiers in Molecular Neuroscience doi: 10.3389/fnmol.2020.00012 (11/02/2020)

• Human and mouse essentiality screens as a resource for disease gene discovery Pilar Cacheiro, Violeta Muñoz-Fuentes, Stephen A. Murray, et al. Nature Communications doi: 10.1038/s41467-020-14284-2 (31/01/2020)

• A restricted spectrum of missense KMT2D variants cause a multiple malformations disorder distinct from Kabuki syndrome Sara Cuvertino, Verity Hartill, Alice Colyer, et al. Genetics in Medicine doi: 10.1038/s41436-019-0743-3 (17/01/2020)

One of the main outputs of the project was to provide a diverse and unique clinical dataset alongside genomic data to ensure continued learning and development from the already accrued knowledge. Continued data acquisition will allow researchers to follow clinical progression of disease and cancers. The dataset in terms of participants is largely stable, with no significant further additions. However, by migrating the data to a cloud based system, this allows further integration of data in developing cohort and genomic browser systems, allowing more in-depth real-time analysis

The role of the 100,000 project was to gain an understanding of the rationale of whole genome sequencing for standard of care testing. The value of the 100,000 project now means that the NHS will provide whole genome sequencing as standard of care for the first time this year. The findings from the 100,000 project were fed back to clinicians to ensure a understanding of how to interpret genomic data and the subsequent translation of that into the concept of personalized medicine - as per the legitimate interests for Genomics England.

The patient participant panel for the 100,000 Genomes project have supported grant and award nominations for Genomics England - celebrating their patient engagement.

DARS-NIC-12784-R8W7V-v6.7 1 April 2020 to 31 March 2021
Title
Genomics England (MR1418) -Renewal Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
20
Files released
327

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-12784-R8W7V-v5.7

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-12784-R8W7V-v5.7
FieldWasBecame
TitleGenomics England (MR1418) - Amendment and Updated Request for tranche of data across multiple data sets.Genomics England (MR1418) -Renewal Request for tranche of data across multiple data sets.
Start date2019-02-012020-04-01
End date2020-03-312021-03-31

Datasets: + Cancer Registration Data; + Civil Registrations of Death; + Demographics; + Emergency Care Data Set (ECDS)

Processing activities

[5 paragraphs unchanged] There will be no data linkage undertaken with NHS Digital data provided under this agreement that is not already noted in the agreement. All organisations party to this agreement must comply with the Data Sharing Framework Contract requirements, including those regarding the use (and purposes of that use) by “Personnel” (as defined within the Data Sharing Framework Contract ie: employees, agents and contractors of the Data Recipient who may have access to that data). PROMS terms and Conditions will be adhered to. PROMS data is only available for non-commercial purposes, such as academic research, or in connection with delivering services to the NHS [52 paragraphs unchanged] All organisations party to this agreement must comply with the Data Sharing Framework Contract requirements, including those regarding the use (and purposes of that use) by “Personnel” (as defined within the Data Sharing Framework Contract ie: employees, agents and contractors of the Data Recipient who may have access to that data). There will be no data linkage undertaken with NHS Digital data provided under this agreement that is not already noted in the agreement. PROMS terms and Conditions will be adhered to. PROMS data is only available for non-commercial purposes, such as academic research, or in connection with delivering services to the NHS

Expected output

[15 paragraphs unchanged] All outputs will contain only data that is aggregated with small numbers suppressed in line with the HES Analysis Guide. (if using HES data)

Benefits reported

[21 paragraphs unchanged] Update 2019 • Clinical and genetic variability in children with partial albinism. Campbell P, Ellingford JM, Parry NRA, et al. Nature Scientific Reports doi: 10.1038/s41598-019-51768-8 (12/11/2019) • Secondary C1q Deficiency in Activated PI3Kδ Syndrome Type 2 Ying Hong, Sira Nanthapisal, Ebun Omoyinmi, et al. Frontiers in Immunologydoi: 10.3389/fimmu.2019.02589 (11/11/2019) • Defective tubulin detyrosination causes structural brain abnormalities with cognitive deficiency in humans and mice Pagnamenta, Alistair T; Heemeryck, Pierre; Hilary C Martin, et al. Human Molecular Genetics doi: 10.1093/hmg/ddz186 (31/07/2019) • Germline selection shapes the landscape of human mitochondrial DNA Wei, Wei; Chinnery, Patrick; et al Science doi: 10.1126/science.aau6520 (01/04/2019) • Validating Whole Genome Sequencing (WGS) for clinical use in AML & ALL S. Henderson, A. Sosinsky, A. Hamblin, et al. British Journal of Haematology doi: 10.1111/bjh.15854 (27/03/2019) • Opportunities and Challenges for Molecular Understanding of Ciliopathies–The 100,000 Genomes Project Wheway, Gabrielle; Mitchison, Hannah H. Frontiers in Genetics doi: 10.3389/fgene.2019.00569 (11/03/2019) Whole-genome sequencing of rare disease patients in a national healthcare system Ouwehand, Willem H bioRxiv – pre-publication doi: 10.1101/507244 (01/01/2019) Update Apr 2020 • Genomic loci susceptible to systematic sequencing bias in clinical whole genomes Timothy M. Freeman, Genomics England Research Consortium, Dennis Wang, and Jason Harris Genome Research doi: 10.1101/gr.255349.119 (11/03/2020) • Mutational signature in colorectal cancer caused by genotoxic pks+ E. coli Cayetano Pleguezuelos-Manzano, Jens Puschhof, Axel Rosendahl Huber, et al. Nature doi: 10.1038/s41586-020-2080-8 (27/02/2020) • Expanding Clinical Presentations Due to Variations in THOC2 mRNA Nuclear Export Factor Raman Kumar, Elizabeth Palmer, Alison E. Gardner, et al. Frontiers in Molecular Neuroscience doi: 10.3389/fnmol.2020.00012 (11/02/2020) • Human and mouse essentiality screens as a resource for disease gene discovery Pilar Cacheiro, Violeta Muñoz-Fuentes, Stephen A. Murray, et al. Nature Communications doi: 10.1038/s41467-020-14284-2 (31/01/2020) • A restricted spectrum of missense KMT2D variants cause a multiple malformations disorder distinct from Kabuki syndrome Sara Cuvertino, Verity Hartill, Alice Colyer, et al. Genetics in Medicine doi: 10.1038/s41436-019-0743-3 (17/01/2020)

Changed only in punctuation, spacing or capitalisation: Expected measurable benefits.

Unchanged: Objective for processing.

Objective for processing

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, though the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to Participants: for many they are receiving a diagnosis for the first time.

The richness of the high-quality data sets are crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of whole genome sequencing (WGS) data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project has completed recruitment of rare disease participants and will complete recruitment of cancer patients in early 2019. Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to collect the following data-sets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

MRIS: Cohort Event Notification, Cause of Death. Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

MRIS: Members and Posting, List Cleaning Report, Flagging Current Status. These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that our data applications are correct and do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their de-identified genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing Participant Data under the General Data Protection

Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.

It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS;

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

Expected output

Genomics England’s target was to complete sequencing of 100,000 genomes by the end of 2018 – this was achieved by December 2018 - - https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000-genomes-sequenced/

By April 2017 over 41,000 whole genomes had been sequenced and, with the aid of clinical data, the first diagnoses were announced:

• First diagnoses from the pilot phase https://www.genomicsengland.co.uk/first-patients-diagnosed-through-the-100000-genomes-project/

• First children diagnosed through the project https://www.genomicsengland.co.uk/first-children-recieve-diagnoses-through-100000-genomes-project/

• Financial Times article: https://www.ft.com/content/d2e21cea-d684-11e6-944b-e7eb37a6aa8e

By November 2018 109,611 samples had been collected, 91,100 genomes sequenced and 37,591 results returned to referring NHS Genomic Medicine Centres. Recruitment to the 100,000 genomes project is now completed, but analysis and return of results will continue throughout the period of this agreement.

By May 2018 the Genomics England Research Environment was established to allow research access to de-identified genome and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with over 1,300 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

By the end of October 2018 the fifth quarterly release of new data into the Genomics England Research Environment was achieved. This release included 71,860 genomes and primary clinical data for 85,070 participants. It also included de-identified clinical data received from NHS Digital for 60,110 participants, totalling 4.2m records. 1,821 GeCIP members from 387 institutions and 33 research domains now requested and been given access to the Research Environment, and 70 research projects have been approved. Alongside this there are 89 members of the industry Discovery Forum, including 10 full members.

Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital, into the Research Environment on the following dates, in order to continue support for, and to further develop, this ground-breaking resource:

• 28th Feb 19

• 30th Apr 19

• 31st Jul 19

• 31st Oct 19

Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.

Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the Research Environment on the dates shown above.

All outputs will contain only data that is aggregated with small numbers suppressed in line with the HES Analysis Guide. (if using HES data)

Benefits reported

Update December 2017

Over 41,000 Genomes were sequenced as of December 2017. Participant stories can be found at: https://www.genomicsengland.co.uk/alexs-story/

Genomics England has built upon its commitment to lead on Governments technology and innovation agenda by forging partnership with industry. Examples of this include a new industry collaboration with leading life sciences companies Inivata and Thermo Fisher Scientific to improve understanding of cancer.

Public Health England has announced that Whole Genome Sequencing (WGS) is now being used to identify different strains of tuberculosis (TB). This is the first time that WGS has been used as a diagnostic solution for managing a disease on this scale anywhere in the world. The technique, developed in conjunction with the University of Oxford, allows faster and more accurate diagnoses, meaning patients can be treated with precisely the right medication more quickly.

Genomics England has now engaged devolved nations and is recruiting participants from Scotland and Wales.

Update May 2018

Over 60,000 genomes have been sequenced and over 12,000 clinical reports have been issued to referring NHS Genomic Medicine Centres. Thirty disease and cross-cutting research domains have had their plans approved and now have access to 100,000 the Genomes Project Research Environment. The number of users with access to the Genomics England Research Environment is now over 1,300.

Twelve publications have arisen from or refer to the 100,000 Genomes Project during the last year, including:

• The 100,000 Genomes Project: bringing whole genome sequencing to the NHS. Clare Turnbull et al. BMJ 2018; doi: https://doi.org/10.1136/bmj.k1687 (24 April 2018)

• Identification of rare sequence variation underlying heritable pulmonary arterial hypertension. Nicholas W. Morrell et al. Nature Communications 2018;9; doi:10.1038/s41467-018-03672-4 (12 April 2018)

• Introducing genomics into cancer care. Sue Hill BRJ Surg 2018;105(2):e14-e15 (17 January 2018)

• Missense variants in the X-linked gene PRPS1 cause retinal degeneration in females. Alessia Fiorentino, Kaoru Fujinami, Gavin Arno et al. Hum Mutat 2017; doi:10.1002/humu.23349 (17 October 2017)

See https://www.genomicsengland.co.uk/category/updates/ and https://www.genomicsengland.co.uk/aboutgecip/ publications/ for details of news and publications.

Genomics England created the Discovery Forum in July 2017. It provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

Update November 2018

Over 109,000 samples have now been collected, 91,100 genomes sequenced and 37,591 clinical reports issued to referring NHS Genomic Medicine Centres. 71,860 genomes have been made available in the Genomics Engalnd Research Environment.

1,821 GeCIP members from 387 institutions and 33 research domains now have access to the research environment. There are now 89 members of the industry Discovery Forum, including 10 full members.

Latest publications include:

• Challenges in implementing genomic medicine: the 100,000 Genomes Project. Julian G. Barwell, Rory B.G. O’Sullivan, Laura K. Mansbridge, Joanna M. Lowry, Huw R. Dorkins. J Transl Genet Genom 2018;2:13; https://doi.org/10.20517/jtgg.2018.17 (11 Sep 2018)

• What will follow the first hundred thousand genomes in the NHS? Malcolm Grant & John Paul Maytum. Per Med 2018; doi:10.2217/pme-2018-0025 (20 June 2018)

• Clinical-grade validation of whole genome sequencing reveals robust detection of low-frequency variants and copy number alterations in CLL. Jenny Klintman, Katerina Barmpouti, Samantha JL Knight et al. Br J Haematol 2018; doi:10.1111/bjh.15406 (29 May 2018)

Update 2019

• Clinical and genetic variability in children with partial albinism. Campbell P, Ellingford JM, Parry NRA, et al. Nature Scientific Reports doi: 10.1038/s41598-019-51768-8 (12/11/2019)

• Secondary C1q Deficiency in Activated PI3Kδ Syndrome Type 2 Ying Hong, Sira Nanthapisal, Ebun Omoyinmi, et al. Frontiers in Immunologydoi: 10.3389/fimmu.2019.02589 (11/11/2019)

• Defective tubulin detyrosination causes structural brain abnormalities with cognitive deficiency in humans and mice Pagnamenta, Alistair T; Heemeryck, Pierre; Hilary C Martin, et al. Human Molecular Genetics doi: 10.1093/hmg/ddz186 (31/07/2019)

• Germline selection shapes the landscape of human mitochondrial DNA Wei, Wei; Chinnery, Patrick; et al Science doi: 10.1126/science.aau6520 (01/04/2019)

• Validating Whole Genome Sequencing (WGS) for clinical use in AML & ALL S. Henderson, A. Sosinsky, A. Hamblin, et al. British Journal of Haematology doi: 10.1111/bjh.15854 (27/03/2019)

• Opportunities and Challenges for Molecular Understanding of Ciliopathies–The 100,000 Genomes Project Wheway, Gabrielle; Mitchison, Hannah H. Frontiers in Genetics doi: 10.3389/fgene.2019.00569 (11/03/2019)

Whole-genome sequencing of rare disease patients in a national healthcare system Ouwehand, Willem H bioRxiv – pre-publication doi: 10.1101/507244 (01/01/2019)

Update Apr 2020

• Genomic loci susceptible to systematic sequencing bias in clinical whole genomes Timothy M. Freeman, Genomics England Research Consortium, Dennis Wang, and Jason Harris Genome Research doi: 10.1101/gr.255349.119 (11/03/2020)

• Mutational signature in colorectal cancer caused by genotoxic pks+ E. coli Cayetano Pleguezuelos-Manzano, Jens Puschhof, Axel Rosendahl Huber, et al. Nature doi: 10.1038/s41586-020-2080-8 (27/02/2020)

• Expanding Clinical Presentations Due to Variations in THOC2 mRNA Nuclear Export Factor Raman Kumar, Elizabeth Palmer, Alison E. Gardner, et al. Frontiers in Molecular Neuroscience doi: 10.3389/fnmol.2020.00012 (11/02/2020)

• Human and mouse essentiality screens as a resource for disease gene discovery Pilar Cacheiro, Violeta Muñoz-Fuentes, Stephen A. Murray, et al. Nature Communications doi: 10.1038/s41467-020-14284-2 (31/01/2020)

• A restricted spectrum of missense KMT2D variants cause a multiple malformations disorder distinct from Kabuki syndrome Sara Cuvertino, Verity Hartill, Alice Colyer, et al. Genetics in Medicine doi: 10.1038/s41436-019-0743-3 (17/01/2020)

DARS-NIC-12784-R8W7V-v5.7 1 February 2019 to 31 March 2020
Title
Genomics England (MR1418) - Amendment and Updated Request for tranche of data across multiple data sets.
Commercial
Yes
Sublicensing
Yes
Datasets
16
Files released
377

Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Diagnostic Imaging Data Set (DID); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); MRIS - Cause of Death Report; MRIS - Cohort Event Notification Report; MRIS - Flagging Current Status Report; MRIS - List Cleaning Report; MRIS - Members and Postings Report; Patient Reported Outcome Measures (Linkable to HES)

Objective for processing

Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.

The project was designed to create a new genomic medicine service for the NHS. It was also designed to create new capability for the clinical genomics research, both academic and industrial, though the creation of a unique medical data set combining genomic sequence data with medical records within a secure research infrastructure.

Genomics England has sought to obtain information from participants’ medical records that span their entire lifetime. The DNA sequence, and information from patients’ health records and provided by the participants, are collected and stored securely as a resource for use by approved researchers for scientific and medical purposes during the life, and after the death, of participants. Diagnoses arising from the sequencing and analysis of the participants’ DNA are being fed back to Participants: for many they are receiving a diagnosis for the first time.

The richness of the high-quality data sets are crucial to the success of the 100,000 Genomes Project in delivering value to the NHS. The evaluation of whole genome sequencing (WGS) data in the context of rich and extended phenotypes derived from electronic health records adds significant value. The richness of the Project dataset allows Genomics England to move beyond the primary phenotype that led to the patient’s enrolment to evaluate the genome sequence in the context of other continuous traits, diseases and responses to therapy.

The 100,000 Genomes Project has completed recruitment of rare disease participants and will complete recruitment of cancer patients in early 2019. Genomics England will now continue to gather and provide genomic and clinical data for this cohort to continue the diagnostic and research aims of the project.

To achieve these goals Genomics wish to continue to collect the following data-sets:

Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.

Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants’ phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients’ histories on individual and cohort level and their relationship with genomic alterations.

Mental Health Data sets: Minimum Data Set, Mental Health and Learning Disabilities Data Set, and Mental Health Services Data Set. The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.

Patient Reported Outcome Measures (PROMS). In combination with other data, this dataset allows correlation of outcomes and pain related measures to genetic markers and interventions in some key diseases. More generally they have utility in non-hypothesis driven research, including research into electronic health records.

MRIS: Cohort Event Notification, Cause of Death. Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.

MRIS: Members and Posting, List Cleaning Report, Flagging Current Status. These identifiable data sets are vital for the safe and efficient management of the participant cohort. They are used to ensure that our data applications are correct and do not cover duplicate participants, that withdrawn participants are excluded from data collection, and that participants’ details are accurate.

Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape. Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help the 100,000 Genomes Project and its partners to turn research findings into treatments, diagnostics and benefits for patients as soon as possible. There are currently over 100 members of the Forum.

As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their de-identified genome and health data.

The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a ‘critical friend’ and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England’s landmark data set.

The lawful basis for processing Participant Data under the General Data Protection

Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.

It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.

The processing is necessary to support and enable Genomics England’s legitimate interests in:

· Creating a new genomic medicine service for the NHS;

· Enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of rare diseases and cancer.

The beneficiaries are:

· Participants through the work we do will influence their care;

· researchers and industry by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;

· and the wider public by accelerating the uptake of genomic medicine making it available to patients in the UK.

Expected output

Genomics England’s target was to complete sequencing of 100,000 genomes by the end of 2018 – this was achieved by December 2018 - - https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000-genomes-sequenced/

By April 2017 over 41,000 whole genomes had been sequenced and, with the aid of clinical data, the first diagnoses were announced:

• First diagnoses from the pilot phase https://www.genomicsengland.co.uk/first-patients-diagnosed-through-the-100000-genomes-project/

• First children diagnosed through the project https://www.genomicsengland.co.uk/first-children-recieve-diagnoses-through-100000-genomes-project/

• Financial Times article: https://www.ft.com/content/d2e21cea-d684-11e6-944b-e7eb37a6aa8e

By November 2018 109,611 samples had been collected, 91,100 genomes sequenced and 37,591 results returned to referring NHS Genomic Medicine Centres. Recruitment to the 100,000 genomes project is now completed, but analysis and return of results will continue throughout the period of this agreement.

By May 2018 the Genomics England Research Environment was established to allow research access to de-identified genome and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with over 1,300 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

By the end of October 2018 the fifth quarterly release of new data into the Genomics England Research Environment was achieved. This release included 71,860 genomes and primary clinical data for 85,070 participants. It also included de-identified clinical data received from NHS Digital for 60,110 participants, totalling 4.2m records. 1,821 GeCIP members from 387 institutions and 33 research domains now requested and been given access to the Research Environment, and 70 research projects have been approved. Alongside this there are 89 members of the industry Discovery Forum, including 10 full members.

Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital, into the Research Environment on the following dates, in order to continue support for, and to further develop, this ground-breaking resource:

• 28th Feb 19

• 30th Apr 19

• 31st Jul 19

• 31st Oct 19

Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.

Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants into the Research Environment on the dates shown above.

Benefits reported

Update December 2017

Over 41,000 Genomes were sequenced as of December 2017. Participant stories can be found at: https://www.genomicsengland.co.uk/alexs-story/

Genomics England has built upon its commitment to lead on Governments technology and innovation agenda by forging partnership with industry. Examples of this include a new industry collaboration with leading life sciences companies Inivata and Thermo Fisher Scientific to improve understanding of cancer.

Public Health England has announced that Whole Genome Sequencing (WGS) is now being used to identify different strains of tuberculosis (TB). This is the first time that WGS has been used as a diagnostic solution for managing a disease on this scale anywhere in the world. The technique, developed in conjunction with the University of Oxford, allows faster and more accurate diagnoses, meaning patients can be treated with precisely the right medication more quickly.

Genomics England has now engaged devolved nations and is recruiting participants from Scotland and Wales.

Update May 2018

Over 60,000 genomes have been sequenced and over 12,000 clinical reports have been issued to referring NHS Genomic Medicine Centres. Thirty disease and cross-cutting research domains have had their plans approved and now have access to 100,000 the Genomes Project Research Environment. The number of users with access to the Genomics England Research Environment is now over 1,300.

Twelve publications have arisen from or refer to the 100,000 Genomes Project during the last year, including:

• The 100,000 Genomes Project: bringing whole genome sequencing to the NHS. Clare Turnbull et al. BMJ 2018; doi: https://doi.org/10.1136/bmj.k1687 (24 April 2018)

• Identification of rare sequence variation underlying heritable pulmonary arterial hypertension. Nicholas W. Morrell et al. Nature Communications 2018;9; doi:10.1038/s41467-018-03672-4 (12 April 2018)

• Introducing genomics into cancer care. Sue Hill BRJ Surg 2018;105(2):e14-e15 (17 January 2018)

• Missense variants in the X-linked gene PRPS1 cause retinal degeneration in females. Alessia Fiorentino, Kaoru Fujinami, Gavin Arno et al. Hum Mutat 2017; doi:10.1002/humu.23349 (17 October 2017)

See https://www.genomicsengland.co.uk/category/updates/ and https://www.genomicsengland.co.uk/aboutgecip/ publications/ for details of news and publications.

Genomics England created the Discovery Forum in July 2017. It provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.

Update November 2018

Over 109,000 samples have now been collected, 91,100 genomes sequenced and 37,591 clinical reports issued to referring NHS Genomic Medicine Centres. 71,860 genomes have been made available in the Genomics Engalnd Research Environment.

1,821 GeCIP members from 387 institutions and 33 research domains now have access to the research environment. There are now 89 members of the industry Discovery Forum, including 10 full members.

Latest publications include:

• Challenges in implementing genomic medicine: the 100,000 Genomes Project. Julian G. Barwell, Rory B.G. O’Sullivan, Laura K. Mansbridge, Joanna M. Lowry, Huw R. Dorkins. J Transl Genet Genom 2018;2:13; https://doi.org/10.20517/jtgg.2018.17 (11 Sep 2018)

• What will follow the first hundred thousand genomes in the NHS? Malcolm Grant & John Paul Maytum. Per Med 2018; doi:10.2217/pme-2018-0025 (20 June 2018)

• Clinical-grade validation of whole genome sequencing reveals robust detection of low-frequency variants and copy number alterations in CLL. Jenny Klintman, Katerina Barmpouti, Samantha JL Knight et al. Br J Haematol 2018; doi:10.1111/bjh.15406 (29 May 2018)

Register history

When this agreement appeared in, or was edited in, each monthly edition of the register. Built by comparing every edition this site holds, the earliest of which is July 2021.

"Amended in place" means NHS England changed the record without issuing a new version number. The register publishes no changelog for those edits; this site infers them by comparing editions. An edit is attributed to the edition it first appears in, not to the date it was made.

Cite this page

NHS England (2026) Data Uses Register, September 2026 edition, agreement DARS-NIC-12784-R8W7V, “Genomics England: Use of data within the National Genomics Research Library (NGRL)”. Read via NHS Data Access Explorer (unofficial), https://healthdatauses.uk/agreements/dars-nic-12784-r8w7v/ (accessed [date]).

This address stays the same, but the page is rebuilt with each monthly edition, so the citation names the edition it shows. Every edition's data is kept in the facts store.

Source: datausesregister_september2026.xlsx, September 2026 edition of the NHS England Data Uses Register. Search that workbook for DARS-NIC-12784-R8W7V to see the original rows.