GenOMICC COVID-19 Study
Genomics England · Research
In term In term in the September 2026 edition: the latest version runs to 29 January 2028.
- Reference
- DARS-NIC-374190-D0N1M
- Current version
- v10.2
- Term of current version
- 30 January 2026 to 29 January 2028
- Start date
- 21 July 2020
- Data controller
- Sole Data Controller
- Commercial purposes
- Yes
- Sublicensing
- Yes
- Files released to date
- 2,769
Why the data was released
Objective for processing
Genomics England requires access to NHS England data for the purpose of the following work programme:
The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19.
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS England and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England National Genomics Research Library (NGRL) to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and its outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By utilising the GenOMICC consortium’s prospective study design that leverages existing recruitment infrastructure in critical care, Genomics England will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
The following NHS England Data will be accessed:
• Hospital Episode Statistics
o Admitted Patient Care
o Accident & Emergency
o Critical Care
o Outpatients
• Emergency Care Data Set (ECDS)
These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
• Diagnostic Imaging Dataset (DID) necessary to provide invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
• Civil Registrations of Death- Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
• Cancer Registration - necessary to see the incidence of cancer within the cohort
• Covid-19 SGSS First Positives (Second Generation Surveillance System)
• Covid-19 Vaccination Status
• Covid-19 Hospitalization in England Surveillance System
These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
• Mental Health Services Data Set- The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
• Covid-19 General Practise Extraction Service assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
The level of the Data will be
• Identifiable- many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to the NGRL hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data
The Data will be minimised as follows.
• Limited to a study cohort identified by Genomics England of;
• ~100,000 patients who consented to participate. ~30,000 of which recruited by Geonomics meeting the study's COVID-19 eligibility criteria, and ~71,243 as a control cohort.
• Genomics England request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.
Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data.
The University of Edinburgh is responsible for acquisition of primary clinical data. That relates to data acquired at patient registration. Genomics England in its provision of whole genome sequencing are applying to NHS England for secondary clinical data to link to the genomic data.
Staff and academics from the University of Edinburgh are required to join a Genomics England Clinical Interpretation Partnership (GeCIP) to access the secondary clinical data from NHS England within the NGRL. Please see the processing activities of this DSA for further detail on GeCIP.
The lawful basis for processing personal data under the UK GDPR is:
Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
The lawful basis for processing special category data under the UK GDPR is:
Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The funding is provided the Department of Health and Social Care. The funding is specifically for the purposes described.
The funder will have no ability to suppress or otherwise limit the publication of findings.
Lifebit provides IT support to Genomics England.
Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.
SUB-LICENCING:
Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.
It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.
The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.
All requests for data access will be subject to the following considerations:
• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).
• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).
• Provision of a signed Genomics England data access agreement to the Access Review Committee.
• Prioritisation of access according to resource availability.
• Facilitation of high-quality health research
Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:
• Diagnostic discovery
• Pre-clinical research
• Clinical Trials Referral
• Real World Evidence / Market Access
Approval process for Commercial organisations for access to the NGRL:
Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.
The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.
Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.
Approval process for GeCIP users (academic) of the NGRL:
• Researcher visits Genomics England website to enrol as a GECIP member
• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)
• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.
• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.
• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.
• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.
All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:
• Identification of diagnoses originally missed by the standardised pipeline
• Feedback of new diagnoses to patients
• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone
Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.
Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.
AIRLOCK POLICY:
The following rules are applied to all airlock requests:
1. All relevant details of the summary results to be transferred must be provided with every request.
2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.
3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.
4. Summary results requested for transfer are assessed using the following criteria:
a. whether the request aligns with the users ARC approval in full;
b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;
c. any data security implications;
d. any disclosure risks;
e. the technical feasibility and associated cost of the request;
f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;
g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.
The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:
• Technical Lead
• User Community Representative
• Bioinformatics Director
• Caldicott Guardian
• Chief Scientist representative
Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.
GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
• UK academic research institutions (e.g., universities, research institutions etc.)
• NHS trusts or authorities
• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project
• Foreign universities and research institutions that carry out significant research activity
• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))
• Foreign healthcare organisations (private or public) that undertake significant research activity
To be eligible for data access as a GeCIP member, applicants must meet these requirements:
• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.
• Their host institution has verified that they are affiliated with that institution.
• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).
• The GeCIP domain lead has approved the application.
• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.
• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).
The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.
GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.
All personnel accessing the Data have been appropriately trained in data protection and confidentiality.
The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:
> National Cancer Registration and Analysis Service (NCRAS)
and uncurated NCRAS data
> Secure Anonymised Information Linkage (SAIL) data; Welsh data
> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)
> NHS Trusts data
> Data feeds from the Intensive Care National Audit Registry (ICNARC)
> The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
The Data will not be linked with any other data.
The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.
To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.
Data accessed under sub licence will not be re-identified.
Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.
Data shared through the Airlock process is aggregate data only and is therefore not personal data so does not require a legal basis under the UK GDPR.
A release register detailing any sub licences and onward sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse
Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract.
In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users.
NHS England will require the ability to audit the sub licensee.
Processing activities
Genomics England will transfer data to NHS England. The data will consist of identifying details NHS Number, Date of Birth, Gender, and a unique person ID for the cohort to be linked with NHS England data.
NHS England will provide the relevant records from the datasets listed in this agreement to Genomics England. The Data will
• contain directly identifying data items including Names, Postcode, Cause of Deaths, Place of Birth, Cancer Registration Number, which are required to provide maximum insight, and therefore maximum value to the researchers accessing the Data.
The NHS England Data is pseudonymised within the AWS cloud and is then loaded into the NGRL. Raw, identifiable files are kept in a secure location on AWS.
First stage of processing – quality verification
• This ensures that the data set is complete, accurate and complies with the NHS data dictionary or relevant specification. Participant identifiers in the dataset are verified against Genomics England's participant details and any updates required to identifiable data fields, e.g., dates of birth, are highlighted. Finally, the data set is reviewed against recent participant withdrawals so that any withdrawals notified after the data application was made can be removed from the data sets.
Second Stage of processing- selection of a de-identified cohort of participants that fulfil a specific research request
• Researchers are members of a Genomics England Clinical Interpretation Partnership (GeCIP) or the Discovery Forum. Research requests are assessed to ensure that they are included in the approved use purposes set out in the Genomics England Protocol and fall within the scope of the relevant GeCIP or the Discovery Forum. Researchers declare any data they wish to bring into the NGRL and any tools they wish to use for analysis.
Third Stage of processing- analysis of the de-identified data sets within the NGRL
• Researchers perform all the analysis and processing within the environment: they do not extract de-identified data. Results data are placed in a secure folder for anonymisation verification before extraction.
The Data will not be transferred to any other location.
The Data will be stored on the NGRL and the AWS Cloud at Genomics England.
Genomics England stores NGRL data on the Cloud provided by Amazon Web Services (AWS).
The Data will be accessed by authorised personnel via remote access.
The Controller(s) must confirm and provide evidence upon audit by NHS England that access via any remote device complies with the data security obligations within this DSA and the Data Sharing Framework Contract.
For remote access:
- Remote access will only be from secure locations situated within the territory of use (as further restricted elsewhere within the DSA if so done) stated within this DSA;
- Access controls granting users the minimum level of access required are in place;
- Remote access is only via secure connections (e.g., VPNs or secure protocols) to protect data;
- Multifactor authentication (MFA) is required for remote access;
- Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for remote access;
- All remote access is undertaken within the scope of the organisation’s DSPT (or other security arrangements as per this DSA) and complies with the organisation’s remote access policy.
The above applies in addition to any condition set out elsewhere within the DSA (e.g. who may carry out processing, and for what purpose).
The Data will be processed worldwide, subject to any further restrictions elsewhere within the DSA.
Data will be released into the NGRL where it will be linked with data feeds from the Intensive Care National Audit Registry (ICNARC) and The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
Access to confidential patient identifiable Data is restricted to an extremely limited number of employees of Genomics England, accessible on AWS.
Substantive employees of Genomics England and researchers who are a member of the GeCIP and Discovery Forum will process the Data for the purposes described above.
Expected output
Combining genomic sequence data from COVID participants together with their medical records has created a ground-breaking research resource. Researchers are currently studying this data to be able to deliver outputs that will have tangible benefits to healthcare and patients. Understanding the symptoms and impact of COVID on people with differing genetic make-ups is hoped to lead to successful treatment of the disease being investigated.
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the NGRL to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. Specific examples of journal outputs are as follows:
• A first update on mapping the human genetic architecture of COVID-19 (2022) Nature 608(7921):E1-E10. doi: 10.1038/s41586-022-04826-7
• Mapping the human genetic architecture of COVID-19 (2021) Nature 600(7889):472-477. doi: 10.1038/s41586-021-03767-x
• Whole-genome sequencing reveals host factors underlying critical COVID-19 (2022) Nature 607(7917):97-103. doi: 10.1038/s41586-022-04576-6
• Genetic mechanisms of critical illness in COVID-19 (2021) Nature 591(7848):92-98. doi: 10.1038/s41586-020-03065-y
Expected measurable benefits
Access to the data will enable the research to discover new rare and common variants alongside new biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
Specifically the short term (6 months) and medium term benefits and outcomes from this programme of research anticipated are;
• Variants enable Polygenic Risk Score to predict greatest risk and avoid ITU
• Pre-morbid clinical conditions or biomarkers of risk and rapid NHS uptake to avoid ITU
• Identify novel therapies or precision interventions for rapid national trials
• Longitudinal life course sequel of COVID-19 for pandemic planning
PATIENT BENEFIT
o Providing improved clinical understanding of disease progression in COVID-19
o Correlation to disease progression and pre-morbid status
o Identification of susceptibility genes
o Develop a biomarker test(s) to predict an individual͛'s response to SARS-CoV-2 exposure, considering both COVID-19 severity and vulnerability to infection.
o Identify targets that can be used in to inform development of new treatments
• New scientific insights and discovery:
o with the consent of patients, creating a database of 35,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers.
o Correlation of host and viral genomic data
o Potential to provide improved testing for future pandemics
o Aide researchers to identify novel targets for vaccines and therapy
o Identification of highly penetrant rare variants in genes and pathways relating to viral susceptibility or immunodeficiency.
o Genome wide association studies (GWAS) using common variants to identify genes and pathways associated with viral response. These analyses will be aligned with other COVID-19 research consortia.
o Rare variant burden analysis to identify genes and pathways enriched in rare variants associated with viral response
• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. WGS could potentially provide the most accurate diagnostic test for COVID 19.
• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices.
• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics.
Benefits reported so far
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS England. Understanding this significant value is why Genomics England are so keen to add NHS England data to its clinical data source for the GenOMICC study.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated August 2022:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
May 2021:
• Identified 4 targets that can utilise and repurpose currently used medications for treatment and management of severe covid.
• Found evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19. This mechanism may also be amenable to targeted treatment with existing drugs.
These results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
Mapping the human genetic architecture of COVID-19
Dec 2021
• Contributed to the meta-analyses that consist of up to 49,562 patients with COVID-19 from 46 studies across 19 countries.
• Identified 13 genome-wide significant loci that are associated with SARS-CoV-2 infection or severe manifestations of COVID-19. Several of these loci correspond to previously documented associations to lung or autoimmune and inflammatory diseases.
• They represent potentially actionable mechanisms in response to infection.
• There is evidence of a causal role for smoking and body-mass index for severe COVID-19 although not for type II diabetes.
This working model of international collaboration underscores what is possible for future genetic discoveries in emerging pandemics, or indeed for any complex human disease.
May 2022
• Identified TMPRSS2 variant has a protective effect against severe COVID-19
• This is a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19
Whole genome sequencing of 7,491 GenOMICC participants admitted with severe COVID
Jul 2022
• Identified 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2)
• Found evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease.
• Found evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets.
The results show that comparisons between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Aug 2022
• Contributed to the meta-analyses bringing together 60 studies from 25 countries for 3 COVID-19 phenotypes.
• Between-study heterogeneity rather than differences across ancestries are a more likely explanation for the observed heterogeneity in the effect sizes across studies
• Identified that the variant rs35705950:G>T located in the promoter of MUC5B (11p15.5) is protective against hospitalization
• Identified that rs190509934:T>C, which is upstream of ACE2, is associated with decreased susceptibility risk. Recent results have shown that the rs190509934:T>C variant lowers ACE2 expression, which in turn confers protection against SARS-CoV-2 infection12.
The biological insights gained by this expansion of the COVID-19 Host Genetic Initiative showed that increasing sample size and diversity remain a fruitful activity to better understand the human genetic architecture of COVID-19.
The first genomes from severe volunteers came through in June 2020.
By June 2021, interim findings from the study were published. These findings have already helped doctors make better decisions when treating patients with COVID-19, resulting in better outcomes. In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy.
Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
Datasets on the current version
Legal basis for provision: Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c); Health and Social Care Act 2012 - s261(5)(d)
| Dataset | Type of data | Sensitivity | Frequency | Confidential data |
|---|---|---|---|---|
| Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Cancer Registration Data | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Civil Registrations of Death | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Community Services Data Set (CSDS) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| COVID-19 Hospitalization in England Surveillance System | Identifiable | Non-Sensitive | Ongoing | Consent (Reasonable Expectation) |
| COVID-19 SGSS First Positives (Second Generation Surveillance System) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| COVID-19 Vaccination Status | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Diagnostic Imaging Data Set (DID) | Identifiable | Non-Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Emergency Care Data Set (ECDS) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| HES-ID to MPS-ID HES Accident and Emergency | Anonymised - ICO Code Compliant | Non-Sensitive | One-Off | Consent (Reasonable Expectation) |
| HES-ID to MPS-ID HES Admitted Patient Care | Anonymised - ICO Code Compliant | Non-Sensitive | One-Off | Consent (Reasonable Expectation) |
| HES-ID to MPS-ID HES Outpatients | Anonymised - ICO Code Compliant | Non-Sensitive | One-Off | Consent (Reasonable Expectation) |
| Hospital Episode Statistics Accident and Emergency (HES A and E) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Hospital Episode Statistics Admitted Patient Care (HES APC) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Hospital Episode Statistics Critical Care (HES Critical Care) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Hospital Episode Statistics Outpatients (HES OP) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
| Mental Health Services Data Set (MHSDS) | Identifiable | Sensitive | Ongoing | Consent (Reasonable Expectation) |
Files released
Files released counts only files released externally by DARS. Access granted in NHS England's own systems, such as its Secure Data Environment, is not included.
This agreement permits sublicensing: the applicant may pass data on to others. Anything passed on is not recorded in this register.
Patient opt-outs were not applied to any of the 2,769 files released under this agreement, across every version. About opt-outs
Files released against version 10.2 of this agreement, summarised by dataset.
| Dataset | Files | First released | Last released | Opt-outs applied |
|---|---|---|---|---|
| Community Services Data Set (CSDS) | 35 | June 2026 | July 2026 | No |
| Hospital Episode Statistics Admitted Patient Care (HES APC) | 21 | May 2026 | June 2026 | No |
| Hospital Episode Statistics Outpatients (HES OP) | 15 | May 2026 | June 2026 | No |
| Hospital Episode Statistics Critical Care (HES Critical Care) | 9 | May 2026 | June 2026 | No |
| Hospital Episode Statistics Accident and Emergency (HES A and E) | 7 | May 2026 | May 2026 | No |
| Diagnostic Imaging Data Set (DID) | 5 | June 2026 | August 2026 | No |
| Emergency Care Data Set (ECDS) | 4 | May 2026 | August 2026 | No |
| COVID-19 Vaccination Status | 2 | May 2026 | August 2026 | No |
| Cancer Registration Data | 2 | May 2026 | August 2026 | No |
| Civil Registrations of Death | 2 | May 2026 | August 2026 | No |
Version history
The register lists each renewal of this agreement as a separate row. This site has 11 versions.
DARS-NIC-374190-D0N1M-v10.2 30 January 2026 to 29 January 2028
- Title
- GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 18
- Files released
- 102
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); COVID-19 Vaccination Status; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS)
What changed from DARS-NIC-374190-D0N1M-v9.3
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2026-01-30 |
Datasets: + HES-ID to MPS-ID HES Accident and Emergency
Unchanged: Objective for processing, Processing activities, Expected output, Expected measurable benefits, Benefits reported.
DARS-NIC-374190-D0N1M-v9.3 30 December 2025 to 29 January 2028
- Title
- GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 17
- Files released
- 0
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); COVID-19 Vaccination Status; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS)
What changed from DARS-NIC-374190-D0N1M-v8.2
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2025-12-30 | |
| End date | 2028-01-29 | |
| Community Services Data Set (CSDS): legal basis | Health and Social Care Act 2012 - s261(5)(d) | |
| Hospital Episode Statistics Admitted Patient Care (HES APC): legal basis | Health and Social Care Act 2012 - s261(5)(d) | |
| Hospital Episode Statistics Critical Care (HES Critical Care): legal basis | Health and Social Care Act 2012 - s261(5)(d) | |
| Hospital Episode Statistics Outpatients (HES OP): legal basis | Health and Social Care Act 2012 - s261(5)(d) |
Datasets:
+ HES-ID to MPS-ID HES Admitted Patient Care; + HES-ID to MPS-ID HES Outpatients · − Secondary Uses Service Payment By Results Spells
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic
[11 words unchanged]
been based on the strength of the clinical data provided by NHS
Digital.
England.
Understanding this significant value is why Genomics England are so keen to add NHS
Digital
England
data to its clinical data source for the GenOMICC study.
[33 paragraphs unchanged]
Unchanged: Objective for processing, Processing activities, Expected output, Expected measurable benefits.
Objective for processing
Genomics England requires access to NHS England data for the purpose of the following work programme:
The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19.
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS England and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England National Genomics Research Library (NGRL) to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and its outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By utilising the GenOMICC consortium’s prospective study design that leverages existing recruitment infrastructure in critical care, Genomics England will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
The following NHS England Data will be accessed:
• Hospital Episode Statistics
o Admitted Patient Care
o Accident & Emergency
o Critical Care
o Outpatients
• Emergency Care Data Set (ECDS)
These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
• Diagnostic Imaging Dataset (DID) necessary to provide invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
• Civil Registrations of Death- Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
• Cancer Registration - necessary to see the incidence of cancer within the cohort
• Covid-19 SGSS First Positives (Second Generation Surveillance System)
• Covid-19 Vaccination Status
• Covid-19 Hospitalization in England Surveillance System
These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
• Mental Health Services Data Set- The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
• Covid-19 General Practise Extraction Service assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
The level of the Data will be
• Identifiable- many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to the NGRL hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data
The Data will be minimised as follows.
• Limited to a study cohort identified by Genomics England of;
• ~100,000 patients who consented to participate. ~30,000 of which recruited by Geonomics meeting the study's COVID-19 eligibility criteria, and ~71,243 as a control cohort.
• Genomics England request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.
Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data.
The University of Edinburgh is responsible for acquisition of primary clinical data. That relates to data acquired at patient registration. Genomics England in its provision of whole genome sequencing are applying to NHS England for secondary clinical data to link to the genomic data.
Staff and academics from the University of Edinburgh are required to join a Genomics England Clinical Interpretation Partnership (GeCIP) to access the secondary clinical data from NHS England within the NGRL. Please see the processing activities of this DSA for further detail on GeCIP.
The lawful basis for processing personal data under the UK GDPR is:
Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
The lawful basis for processing special category data under the UK GDPR is:
Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The funding is provided the Department of Health and Social Care. The funding is specifically for the purposes described.
The funder will have no ability to suppress or otherwise limit the publication of findings.
Lifebit provides IT support to Genomics England.
Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.
SUB-LICENCING:
Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.
It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.
The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.
All requests for data access will be subject to the following considerations:
• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).
• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).
• Provision of a signed Genomics England data access agreement to the Access Review Committee.
• Prioritisation of access according to resource availability.
• Facilitation of high-quality health research
Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:
• Diagnostic discovery
• Pre-clinical research
• Clinical Trials Referral
• Real World Evidence / Market Access
Approval process for Commercial organisations for access to the NGRL:
Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.
The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.
Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.
Approval process for GeCIP users (academic) of the NGRL:
• Researcher visits Genomics England website to enrol as a GECIP member
• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)
• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.
• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.
• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.
• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.
All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:
• Identification of diagnoses originally missed by the standardised pipeline
• Feedback of new diagnoses to patients
• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone
Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.
Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.
AIRLOCK POLICY:
The following rules are applied to all airlock requests:
1. All relevant details of the summary results to be transferred must be provided with every request.
2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.
3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.
4. Summary results requested for transfer are assessed using the following criteria:
a. whether the request aligns with the users ARC approval in full;
b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;
c. any data security implications;
d. any disclosure risks;
e. the technical feasibility and associated cost of the request;
f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;
g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.
The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:
• Technical Lead
• User Community Representative
• Bioinformatics Director
• Caldicott Guardian
• Chief Scientist representative
Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.
GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
• UK academic research institutions (e.g., universities, research institutions etc.)
• NHS trusts or authorities
• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project
• Foreign universities and research institutions that carry out significant research activity
• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))
• Foreign healthcare organisations (private or public) that undertake significant research activity
To be eligible for data access as a GeCIP member, applicants must meet these requirements:
• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.
• Their host institution has verified that they are affiliated with that institution.
• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).
• The GeCIP domain lead has approved the application.
• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.
• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).
The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.
GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.
All personnel accessing the Data have been appropriately trained in data protection and confidentiality.
The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:
> National Cancer Registration and Analysis Service (NCRAS)
and uncurated NCRAS data
> Secure Anonymised Information Linkage (SAIL) data; Welsh data
> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)
> NHS Trusts data
> Data feeds from the Intensive Care National Audit Registry (ICNARC)
> The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
The Data will not be linked with any other data.
The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.
To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.
Data accessed under sub licence will not be re-identified.
Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.
Data shared through the Airlock process is aggregate data only and is therefore not personal data so does not require a legal basis under the UK GDPR.
A release register detailing any sub licences and onward sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse
Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract.
In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users.
NHS England will require the ability to audit the sub licensee.
Expected output
Combining genomic sequence data from COVID participants together with their medical records has created a ground-breaking research resource. Researchers are currently studying this data to be able to deliver outputs that will have tangible benefits to healthcare and patients. Understanding the symptoms and impact of COVID on people with differing genetic make-ups is hoped to lead to successful treatment of the disease being investigated.
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the NGRL to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. Specific examples of journal outputs are as follows:
• A first update on mapping the human genetic architecture of COVID-19 (2022) Nature 608(7921):E1-E10. doi: 10.1038/s41586-022-04826-7
• Mapping the human genetic architecture of COVID-19 (2021) Nature 600(7889):472-477. doi: 10.1038/s41586-021-03767-x
• Whole-genome sequencing reveals host factors underlying critical COVID-19 (2022) Nature 607(7917):97-103. doi: 10.1038/s41586-022-04576-6
• Genetic mechanisms of critical illness in COVID-19 (2021) Nature 591(7848):92-98. doi: 10.1038/s41586-020-03065-y
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS England. Understanding this significant value is why Genomics England are so keen to add NHS England data to its clinical data source for the GenOMICC study.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated August 2022:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
May 2021:
• Identified 4 targets that can utilise and repurpose currently used medications for treatment and management of severe covid.
• Found evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19. This mechanism may also be amenable to targeted treatment with existing drugs.
These results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
Mapping the human genetic architecture of COVID-19
Dec 2021
• Contributed to the meta-analyses that consist of up to 49,562 patients with COVID-19 from 46 studies across 19 countries.
• Identified 13 genome-wide significant loci that are associated with SARS-CoV-2 infection or severe manifestations of COVID-19. Several of these loci correspond to previously documented associations to lung or autoimmune and inflammatory diseases.
• They represent potentially actionable mechanisms in response to infection.
• There is evidence of a causal role for smoking and body-mass index for severe COVID-19 although not for type II diabetes.
This working model of international collaboration underscores what is possible for future genetic discoveries in emerging pandemics, or indeed for any complex human disease.
May 2022
• Identified TMPRSS2 variant has a protective effect against severe COVID-19
• This is a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19
Whole genome sequencing of 7,491 GenOMICC participants admitted with severe COVID
Jul 2022
• Identified 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2)
• Found evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease.
• Found evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets.
The results show that comparisons between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Aug 2022
• Contributed to the meta-analyses bringing together 60 studies from 25 countries for 3 COVID-19 phenotypes.
• Between-study heterogeneity rather than differences across ancestries are a more likely explanation for the observed heterogeneity in the effect sizes across studies
• Identified that the variant rs35705950:G>T located in the promoter of MUC5B (11p15.5) is protective against hospitalization
• Identified that rs190509934:T>C, which is upstream of ACE2, is associated with decreased susceptibility risk. Recent results have shown that the rs190509934:T>C variant lowers ACE2 expression, which in turn confers protection against SARS-CoV-2 infection12.
The biological insights gained by this expansion of the COVID-19 Host Genetic Initiative showed that increasing sample size and diversity remain a fruitful activity to better understand the human genetic architecture of COVID-19.
The first genomes from severe volunteers came through in June 2020.
By June 2021, interim findings from the study were published. These findings have already helped doctors make better decisions when treating patients with COVID-19, resulting in better outcomes. In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy.
Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
DARS-NIC-374190-D0N1M-v8.2 15 January 2025 to 14 January 2026
- Title
- GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 16
- Files released
- 0
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); COVID-19 Vaccination Status; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v7.2
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2025-01-15 | |
| End date | 2026-01-14 |
Objective for processing
[107 paragraphs unchanged]
The Data will be processed worldwide.
[32 paragraphs unchanged]
Data shared through the Airlock process is aggregate data only and is
t herefore
therefore
not personal data so does not require a legal basis under the UK GDPR.
[4 paragraphs unchanged]
Processing activities
[23 paragraphs unchanged]
The Data will be processed
worldwide.
worldwide, subject to any further restrictions elsewhere within the DSA.
[3 paragraphs unchanged]
Expected measurable benefits
[10 paragraphs unchanged]
o Develop a biomarker test(s) to predict an
individual͛s
individual͛'s
response to SARS-CoV-2 exposure, considering both COVID-19 severity and vulnerability to infection.
[12 paragraphs unchanged]
Unchanged: Expected output, Benefits reported.
Objective for processing
Genomics England requires access to NHS England data for the purpose of the following work programme:
The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19.
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS England and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England National Genomics Research Library (NGRL) to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and its outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By utilising the GenOMICC consortium’s prospective study design that leverages existing recruitment infrastructure in critical care, Genomics England will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
The following NHS England Data will be accessed:
• Hospital Episode Statistics
o Admitted Patient Care
o Accident & Emergency
o Critical Care
o Outpatients
• Emergency Care Data Set (ECDS)
These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
• Diagnostic Imaging Dataset (DID) necessary to provide invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
• Civil Registrations of Death- Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
• Cancer Registration - necessary to see the incidence of cancer within the cohort
• Covid-19 SGSS First Positives (Second Generation Surveillance System)
• Covid-19 Vaccination Status
• Covid-19 Hospitalization in England Surveillance System
These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
• Mental Health Services Data Set- The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
• Covid-19 General Practise Extraction Service assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
The level of the Data will be
• Identifiable- many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to the NGRL hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data
The Data will be minimised as follows.
• Limited to a study cohort identified by Genomics England of;
• ~100,000 patients who consented to participate. ~30,000 of which recruited by Geonomics meeting the study's COVID-19 eligibility criteria, and ~71,243 as a control cohort.
• Genomics England request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.
Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data.
The University of Edinburgh is responsible for acquisition of primary clinical data. That relates to data acquired at patient registration. Genomics England in its provision of whole genome sequencing are applying to NHS England for secondary clinical data to link to the genomic data.
Staff and academics from the University of Edinburgh are required to join a Genomics England Clinical Interpretation Partnership (GeCIP) to access the secondary clinical data from NHS England within the NGRL. Please see the processing activities of this DSA for further detail on GeCIP.
The lawful basis for processing personal data under the UK GDPR is:
Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
The lawful basis for processing special category data under the UK GDPR is:
Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The funding is provided the Department of Health and Social Care. The funding is specifically for the purposes described.
The funder will have no ability to suppress or otherwise limit the publication of findings.
Lifebit provides IT support to Genomics England.
Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.
SUB-LICENCING:
Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.
It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.
The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.
All requests for data access will be subject to the following considerations:
• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).
• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).
• Provision of a signed Genomics England data access agreement to the Access Review Committee.
• Prioritisation of access according to resource availability.
• Facilitation of high-quality health research
Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:
• Diagnostic discovery
• Pre-clinical research
• Clinical Trials Referral
• Real World Evidence / Market Access
Approval process for Commercial organisations for access to the NGRL:
Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.
The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.
Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.
Approval process for GeCIP users (academic) of the NGRL:
• Researcher visits Genomics England website to enrol as a GECIP member
• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)
• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.
• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.
• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.
• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.
All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:
• Identification of diagnoses originally missed by the standardised pipeline
• Feedback of new diagnoses to patients
• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone
Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.
Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.
AIRLOCK POLICY:
The following rules are applied to all airlock requests:
1. All relevant details of the summary results to be transferred must be provided with every request.
2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.
3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.
4. Summary results requested for transfer are assessed using the following criteria:
a. whether the request aligns with the users ARC approval in full;
b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;
c. any data security implications;
d. any disclosure risks;
e. the technical feasibility and associated cost of the request;
f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;
g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.
The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:
• Technical Lead
• User Community Representative
• Bioinformatics Director
• Caldicott Guardian
• Chief Scientist representative
Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.
GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
• UK academic research institutions (e.g., universities, research institutions etc.)
• NHS trusts or authorities
• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project
• Foreign universities and research institutions that carry out significant research activity
• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))
• Foreign healthcare organisations (private or public) that undertake significant research activity
To be eligible for data access as a GeCIP member, applicants must meet these requirements:
• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.
• Their host institution has verified that they are affiliated with that institution.
• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).
• The GeCIP domain lead has approved the application.
• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.
• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).
The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.
GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.
All personnel accessing the Data have been appropriately trained in data protection and confidentiality.
The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:
> National Cancer Registration and Analysis Service (NCRAS)
and uncurated NCRAS data
> Secure Anonymised Information Linkage (SAIL) data; Welsh data
> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)
> NHS Trusts data
> Data feeds from the Intensive Care National Audit Registry (ICNARC)
> The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
The Data will not be linked with any other data.
The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.
To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.
Data accessed under sub licence will not be re-identified.
Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.
Data shared through the Airlock process is aggregate data only and is therefore not personal data so does not require a legal basis under the UK GDPR.
A release register detailing any sub licences and onward sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse
Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract.
In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users.
NHS England will require the ability to audit the sub licensee.
Expected output
Combining genomic sequence data from COVID participants together with their medical records has created a ground-breaking research resource. Researchers are currently studying this data to be able to deliver outputs that will have tangible benefits to healthcare and patients. Understanding the symptoms and impact of COVID on people with differing genetic make-ups is hoped to lead to successful treatment of the disease being investigated.
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the NGRL to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. Specific examples of journal outputs are as follows:
• A first update on mapping the human genetic architecture of COVID-19 (2022) Nature 608(7921):E1-E10. doi: 10.1038/s41586-022-04826-7
• Mapping the human genetic architecture of COVID-19 (2021) Nature 600(7889):472-477. doi: 10.1038/s41586-021-03767-x
• Whole-genome sequencing reveals host factors underlying critical COVID-19 (2022) Nature 607(7917):97-103. doi: 10.1038/s41586-022-04576-6
• Genetic mechanisms of critical illness in COVID-19 (2021) Nature 591(7848):92-98. doi: 10.1038/s41586-020-03065-y
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated August 2022:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
May 2021:
• Identified 4 targets that can utilise and repurpose currently used medications for treatment and management of severe covid.
• Found evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19. This mechanism may also be amenable to targeted treatment with existing drugs.
These results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
Mapping the human genetic architecture of COVID-19
Dec 2021
• Contributed to the meta-analyses that consist of up to 49,562 patients with COVID-19 from 46 studies across 19 countries.
• Identified 13 genome-wide significant loci that are associated with SARS-CoV-2 infection or severe manifestations of COVID-19. Several of these loci correspond to previously documented associations to lung or autoimmune and inflammatory diseases.
• They represent potentially actionable mechanisms in response to infection.
• There is evidence of a causal role for smoking and body-mass index for severe COVID-19 although not for type II diabetes.
This working model of international collaboration underscores what is possible for future genetic discoveries in emerging pandemics, or indeed for any complex human disease.
May 2022
• Identified TMPRSS2 variant has a protective effect against severe COVID-19
• This is a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19
Whole genome sequencing of 7,491 GenOMICC participants admitted with severe COVID
Jul 2022
• Identified 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2)
• Found evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease.
• Found evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets.
The results show that comparisons between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Aug 2022
• Contributed to the meta-analyses bringing together 60 studies from 25 countries for 3 COVID-19 phenotypes.
• Between-study heterogeneity rather than differences across ancestries are a more likely explanation for the observed heterogeneity in the effect sizes across studies
• Identified that the variant rs35705950:G>T located in the promoter of MUC5B (11p15.5) is protective against hospitalization
• Identified that rs190509934:T>C, which is upstream of ACE2, is associated with decreased susceptibility risk. Recent results have shown that the rs190509934:T>C variant lowers ACE2 expression, which in turn confers protection against SARS-CoV-2 infection12.
The biological insights gained by this expansion of the COVID-19 Host Genetic Initiative showed that increasing sample size and diversity remain a fruitful activity to better understand the human genetic architecture of COVID-19.
The first genomes from severe volunteers came through in June 2020.
By June 2021, interim findings from the study were published. These findings have already helped doctors make better decisions when treating patients with COVID-19, resulting in better outcomes. In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy.
Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
DARS-NIC-374190-D0N1M-v7.2 4 October 2024 to 3 January 2025
- Title
- GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 16
- Files released
- 0
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); COVID-19 Vaccination Status; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v6.4
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2024-10-04 | |
| End date | 2025-01-03 |
Processing activities
[20 paragraphs unchanged]
- Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for
the
remote access;
[6 paragraphs unchanged]
Unchanged: Objective for processing, Expected output, Expected measurable benefits, Benefits reported.
Objective for processing
Genomics England requires access to NHS England data for the purpose of the following work programme:
The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19.
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS England and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England National Genomics Research Library (NGRL) to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and its outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By utilising the GenOMICC consortium’s prospective study design that leverages existing recruitment infrastructure in critical care, Genomics England will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
The following NHS England Data will be accessed:
• Hospital Episode Statistics
o Admitted Patient Care
o Accident & Emergency
o Critical Care
o Outpatients
• Emergency Care Data Set (ECDS)
These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
• Diagnostic Imaging Dataset (DID) necessary to provide invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
• Civil Registrations of Death- Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
• Cancer Registration - necessary to see the incidence of cancer within the cohort
• Covid-19 SGSS First Positives (Second Generation Surveillance System)
• Covid-19 Vaccination Status
• Covid-19 Hospitalization in England Surveillance System
These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
• Mental Health Services Data Set- The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
• Covid-19 General Practise Extraction Service assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
The level of the Data will be
• Identifiable- many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to the NGRL hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data
The Data will be minimised as follows.
• Limited to a study cohort identified by Genomics England of;
• ~100,000 patients who consented to participate. ~30,000 of which recruited by Geonomics meeting the study's COVID-19 eligibility criteria, and ~71,243 as a control cohort.
• Genomics England request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.
Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data.
The University of Edinburgh is responsible for acquisition of primary clinical data. That relates to data acquired at patient registration. Genomics England in its provision of whole genome sequencing are applying to NHS England for secondary clinical data to link to the genomic data.
Staff and academics from the University of Edinburgh are required to join a Genomics England Clinical Interpretation Partnership (GeCIP) to access the secondary clinical data from NHS England within the NGRL. Please see the processing activities of this DSA for further detail on GeCIP.
The lawful basis for processing personal data under the UK GDPR is:
Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
The lawful basis for processing special category data under the UK GDPR is:
Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The funding is provided the Department of Health and Social Care. The funding is specifically for the purposes described.
The funder will have no ability to suppress or otherwise limit the publication of findings.
Lifebit provides IT support to Genomics England.
Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.
SUB-LICENCING:
Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.
It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.
The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.
All requests for data access will be subject to the following considerations:
• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).
• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).
• Provision of a signed Genomics England data access agreement to the Access Review Committee.
• Prioritisation of access according to resource availability.
• Facilitation of high-quality health research
Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:
• Diagnostic discovery
• Pre-clinical research
• Clinical Trials Referral
• Real World Evidence / Market Access
Approval process for Commercial organisations for access to the NGRL:
Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.
The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.
Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.
Approval process for GeCIP users (academic) of the NGRL:
• Researcher visits Genomics England website to enrol as a GECIP member
• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)
• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.
• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.
• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.
• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.
All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:
• Identification of diagnoses originally missed by the standardised pipeline
• Feedback of new diagnoses to patients
• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone
Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.
Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.
AIRLOCK POLICY:
The following rules are applied to all airlock requests:
1. All relevant details of the summary results to be transferred must be provided with every request.
2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.
3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.
4. Summary results requested for transfer are assessed using the following criteria:
a. whether the request aligns with the users ARC approval in full;
b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;
c. any data security implications;
d. any disclosure risks;
e. the technical feasibility and associated cost of the request;
f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;
g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.
The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:
• Technical Lead
• User Community Representative
• Bioinformatics Director
• Caldicott Guardian
• Chief Scientist representative
The Data will be processed worldwide.
Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.
GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
• UK academic research institutions (e.g., universities, research institutions etc.)
• NHS trusts or authorities
• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project
• Foreign universities and research institutions that carry out significant research activity
• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))
• Foreign healthcare organisations (private or public) that undertake significant research activity
To be eligible for data access as a GeCIP member, applicants must meet these requirements:
• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.
• Their host institution has verified that they are affiliated with that institution.
• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).
• The GeCIP domain lead has approved the application.
• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.
• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).
The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.
GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.
All personnel accessing the Data have been appropriately trained in data protection and confidentiality.
The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:
> National Cancer Registration and Analysis Service (NCRAS)
and uncurated NCRAS data
> Secure Anonymised Information Linkage (SAIL) data; Welsh data
> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)
> NHS Trusts data
> Data feeds from the Intensive Care National Audit Registry (ICNARC)
> The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
The Data will not be linked with any other data.
The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.
To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.
Data accessed under sub licence will not be re-identified.
Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.
Data shared through the Airlock process is aggregate data only and is t herefore not personal data so does not require a legal basis under the UK GDPR.
A release register detailing any sub licences and onward sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse
Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract.
In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users.
NHS England will require the ability to audit the sub licensee.
Expected output
Combining genomic sequence data from COVID participants together with their medical records has created a ground-breaking research resource. Researchers are currently studying this data to be able to deliver outputs that will have tangible benefits to healthcare and patients. Understanding the symptoms and impact of COVID on people with differing genetic make-ups is hoped to lead to successful treatment of the disease being investigated.
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the NGRL to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. Specific examples of journal outputs are as follows:
• A first update on mapping the human genetic architecture of COVID-19 (2022) Nature 608(7921):E1-E10. doi: 10.1038/s41586-022-04826-7
• Mapping the human genetic architecture of COVID-19 (2021) Nature 600(7889):472-477. doi: 10.1038/s41586-021-03767-x
• Whole-genome sequencing reveals host factors underlying critical COVID-19 (2022) Nature 607(7917):97-103. doi: 10.1038/s41586-022-04576-6
• Genetic mechanisms of critical illness in COVID-19 (2021) Nature 591(7848):92-98. doi: 10.1038/s41586-020-03065-y
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated August 2022:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
May 2021:
• Identified 4 targets that can utilise and repurpose currently used medications for treatment and management of severe covid.
• Found evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19. This mechanism may also be amenable to targeted treatment with existing drugs.
These results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
Mapping the human genetic architecture of COVID-19
Dec 2021
• Contributed to the meta-analyses that consist of up to 49,562 patients with COVID-19 from 46 studies across 19 countries.
• Identified 13 genome-wide significant loci that are associated with SARS-CoV-2 infection or severe manifestations of COVID-19. Several of these loci correspond to previously documented associations to lung or autoimmune and inflammatory diseases.
• They represent potentially actionable mechanisms in response to infection.
• There is evidence of a causal role for smoking and body-mass index for severe COVID-19 although not for type II diabetes.
This working model of international collaboration underscores what is possible for future genetic discoveries in emerging pandemics, or indeed for any complex human disease.
May 2022
• Identified TMPRSS2 variant has a protective effect against severe COVID-19
• This is a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19
Whole genome sequencing of 7,491 GenOMICC participants admitted with severe COVID
Jul 2022
• Identified 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2)
• Found evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease.
• Found evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets.
The results show that comparisons between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Aug 2022
• Contributed to the meta-analyses bringing together 60 studies from 25 countries for 3 COVID-19 phenotypes.
• Between-study heterogeneity rather than differences across ancestries are a more likely explanation for the observed heterogeneity in the effect sizes across studies
• Identified that the variant rs35705950:G>T located in the promoter of MUC5B (11p15.5) is protective against hospitalization
• Identified that rs190509934:T>C, which is upstream of ACE2, is associated with decreased susceptibility risk. Recent results have shown that the rs190509934:T>C variant lowers ACE2 expression, which in turn confers protection against SARS-CoV-2 infection12.
The biological insights gained by this expansion of the COVID-19 Host Genetic Initiative showed that increasing sample size and diversity remain a fruitful activity to better understand the human genetic architecture of COVID-19.
The first genomes from severe volunteers came through in June 2020.
By June 2021, interim findings from the study were published. These findings have already helped doctors make better decisions when treating patients with COVID-19, resulting in better outcomes. In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy.
Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
DARS-NIC-374190-D0N1M-v6.4 15 February 2024 to 14 August 2024
- Title
- GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 16
- Files released
- 49
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); COVID-19 Vaccination Status; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v5.5
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Title | GenOMICC COVID-19 Study | |
| Start date | 2024-02-15 | |
| End date | 2024-08-14 |
Datasets:
− Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; − Demographics; − HES-ID to MPS-ID HES Accident and Emergency; − HES-ID to MPS-ID HES Admitted Patient Care; − HES-ID to MPS-ID HES Outpatients
Objective for processing
This Agreement is seeking approval to request continued supply of data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
Genomics England requires access to NHS England data for the purpose of the following work programme:
The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19.
[3 paragraphs unchanged]
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS
Digital
England
and Health Data Research UK (HDRUK). Genomics England will include deep immune
[14 words unchanged]
upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
[2 paragraphs unchanged]
6. To provide access to these data sets via the Genomics England
Trusted
National Genomics
Research
Environment (TRE)
Library (NGRL)
to international and national academia and industry and facilitate international collaboration on COVID-19.
[1 paragraph unchanged]
8. To engage and involve public and patients in setting strategy and
[7 words unchanged]
outputs. This will initially be based upon the 100,000 Genomes Project Participant
Panel.
Panel
[4 paragraphs unchanged]
Genomics England Background: (For Context)
The following NHS England Data will be accessed:
Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.
• Hospital Episode Statistics
Data will be released into the Genomics England TRE where it will be linked with associated clinical data as well as additional data sources, which will include NHS Digital data requested through this agreement, COVID-19 testing feeds and viral genomics from Public Health England and data feeds from the Intensive Care National Audit Registry (ICNARC).
o Admitted Patient Care
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable TRE for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new TRE (COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables effective COVID-19 research outcomes across a wide range of academic and Biotech/Pharma researchers with varying levels of technical competency. This offers a major upgrade to Genomics England’s current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling, and enable shared workspaces between Genomics England and Partners. COVID-RE must serve the immediate COVID-19 research effort and may also advance the transformation of Genomics England’s platform infrastructure, which is a key enabler to the research community.
o Accident & Emergency
There are two key value streams provided by the COVID-RE. Firstly ‘raw data to analytics ready data’ stream which must permit the collection of data from multiple locations from unstructured, through semi and fully structured forms and transforms them the appropriate data model, data store based on their ongoing use and availability. Secondly the ‘discovery to insight’ stream that supports researchers by providing the capability and the framework to support their User Journey from understanding the data available to them through to providing the necessary analytics and publishing tools.
o Critical Care
The newly established COVID-RE will provide the following:
o Outpatients
• Seamless interface with cloud storage and compute capabilities.
• Emergency Care Data Set (ECDS)
• A unified data platform containing datastores appropriate for all the required clinical and genomic data types.
These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
• Data integration capability to deliver analytics-ready datasets into domain-specific data stores across batch and streaming integration patterns.
• Diagnostic Imaging Dataset (DID) necessary to provide invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
• Standards-based access services providing secure, fast, flexible, robust and auditable access to TRE data assets.
• Civil Registrations of Death- Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
• Applications to support the research user journey from an exploration of data sets, through cohort building, analysis and publication - providing tools appropriate to a variety of user requirements (ways of working) and programming competency.
• Cancer Registration - necessary to see the incidence of cancer within the cohort
• Applications and workflows can be delivered natively within the platform, however COVID-RE must enable access to container-based applications and, via a set of hardened APIs, to a relevant external service (as long as security is maintained).
• Covid-19 SGSS First Positives (Second Generation Surveillance System)
Details of the sublicense model via the Genomics England TRE are supplied within the processing activities section of this agreement.
• Covid-19 Vaccination Status
Regarding genome data sharing with sublicensees - As an example, if an individual’s whole-exome or whole-genome sequence forms part of a database but all other records and identifiers have been fully deleted so that all that remains is sequence data, this data could easily be ‘individuated’ or singled out. However, it can be argued that it cannot be connected to a natural person without something further which can relate it to them.
• Covid-19 Hospitalization in England Surveillance System
In this perspective, the Global Alliance for Genomics and Health (GA4GH) has grouped together to provide unified strategies for addressing the major challenges of this data revolution. Genomics England are key members in this alliance and a recent publication (https://www.cell.com/cell-genomics/fulltext/S2666-979X(21)00036-7) presents the GA4GH suite of secure, interoperable technical standards and policy frameworks
These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
• GA4GH and the PHG foundation both suggest that genetic, and particularly genomic information should not be viewed as inherently or directly identifying without some further link to or impact on an individual
• Mental Health Services Data Set- The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
• To that end, Genomics' TRE is a leading light in security and standards as per HDR UK (https://zenodo.org/record/5767586#.YcPE9hPP27M) and an author is an employee of Genomics England.
• Covid-19 General Practise Extraction Service assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
• Furthermore, Genomics' Airlock policy is clear and explicit about the level of data that can be extracted and does not permit row level/ individual data extraction
The level of the Data will be
The additional clinical data is key for researchers to be able to understand and infer clinically relevant and actionable findings. The aim is to provide a high quality, diverse clinical dataset, detailing each participant’s journey and to understand pre-existing conditions and also the early behaviours in the disease course. Genomics England currently have an agreement to receive NHS Digital data for the extant 100,000 Genomes Project participants (DARS-NIC-12784) and have seen first-hand the depth and quality of the data and how it has aided researchers.
• Identifiable- many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to the NGRL hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data
To that aim, the significant gap in the current data collection for the COVID-19 project can be addressed by the non-standard NHS Digital data feed.
The Data will be minimised as follows.
The Secondary Use Services Admitted Patient Care Feed (SUS APC) would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
• Limited to a study cohort identified by Genomics England of;
The associated COVID-19 datasets (including NHS111 and CV19 Testing Data in particular which will be requested under a future version of this agreement) will add to the detail in patient journey course, flagging for instance how participants were monitoring their symptoms and was there an associated poor prognosis.
• ~100,000 patients who consented to participate. ~30,000 of which recruited by Geonomics meeting the study's COVID-19 eligibility criteria, and ~71,243 as a control cohort.
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to approx. 36,000 participants.
• Genomics England request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.
The data would be available for analysis alongside the extant Genomics England data set of 100,000 Genomes Project participants and would be made available to approved researchers worldwide as per existing governance procedures. This 100,000 Genomes Project data set will act as an appropriate matched control group. Data on the 100,000 Genomes Cohort will be provided to NHS Digital for linkage - (as they will act as a control cohort), as well as new additions for the GenOMICC Study.
Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data.
The below detail pertains to the background of the study and susceptibility to infection sets out why Genomics are investigating. The origin of the GenOMICC study was focused on smaller cohorts of participants. However, in collaboration with the COG-UK group, this is now expanded to whole genome sequence of approx. 36,000 affected individuals as set out above.
The University of Edinburgh is responsible for acquisition of primary clinical data. That relates to data acquired at patient registration. Genomics England in its provision of whole genome sequencing are applying to NHS England for secondary clinical data to link to the genomic data.
GenOMICC Study Background:
Staff and academics from the University of Edinburgh are required to join a Genomics England Clinical Interpretation Partnership (GeCIP) to access the secondary clinical data from NHS England within the NGRL. Please see the processing activities of this DSA for further detail on GeCIP.
Susceptibility to infection is profoundly heritable (Sorensen et al. 1988). Patients who develop life-threatening illness following infection with usually innocuous pathogens, such as influenza (Miller et al. 2010), are genetically different from the rest of the population (Albright et al. 2008). Understanding the genetic mechanisms of susceptibility may yield new therapeutic targets (Baillie 2014) that can be used to make susceptible patients more similar to individuals who are resistant to, or tolerant of, specific pathogens.
The lawful basis for processing personal data under the UK GDPR is:
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV [Human Immunodeficiency Virus] (Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions (e.g., small molecules to inhibit an enzyme or receptor that is dysfunctional in resistant individuals) would therefore be protective to the host in a similar way to antibiotics, with the advantage that it is conceptually more difficult for any one pathogen to evolve resistance to such a therapy.
Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.
A second, more challenging problem arises in patients who become critically ill following infection. The patterns of immune-mediated organ dysfunction, immunoparesis, and death are very similar in severe infections and sterile systemic injuries (such as burns, haemorrhage, pancreatitis and trauma). Ultimately, death is a consequence of the host response to injury (Angus and Poll 2013), through final common pathways of organ failure that are clinically and biochemically evident, and unrelated to the original precipitant.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
Broadly, the severity of critical illness follows directly from the severity and duration of the initial insult. In bacterial sepsis, early antibiotics are the mainstay of therapy; in influenza, early antivirals; in haemorrhage, early resuscitation; in trauma, urgent action to prevent secondary injury. There are no therapies with which to modulate the host response to systemic injury.
There is a lack of direct evidence of heritability for outcomes of critical illness, due in part to difficulties in defining and quantifying the heterogeneous multi-organ dysfunction syndrome (MODS), and in part due to the rapid pace of change in critical care medicine, making it impossible to tackle this question in long term outcome studies. However, clinical and biological evidence support the hypothesis that the pathogenesis of MODS is immune in origin (Angus and Poll 2013). Hence, predictions can be made from the extensive knowledge of other immune conditions. Whether or not MODS is considered to be an autoimmune or infectious condition is moot: these conditions share a great deal of similarity in genetic predispositions, cell types and mechanisms of pathogenesis. It is therefore very likely that propensity to survive MODS has a heritable component, and there is some direct evidence in support of this hypothesis (Rautanen et al. 2015). If this is the case, then the identity of the specific variants that contribute to outcome could potentially be utilised to design therapies to promote survival after the onset of MODS.
This study aims to identify genetic predisposition to specific syndromes of critical illness. Specifically, susceptibility to life-threatening infections caused by an identified pathogen, and susceptibility to death following the onset of organ failure due to sepsis or sterile injury. In order to maximise the probability of identifying host genetic loci associated with susceptibility, Genomics England will restrict some analyses to younger individuals in good general health and lacking in known predisposing factors.
The same principle was used to determine an upper age limit for inclusion for some analyses. With advancing age, there is an increase in undiagnosed co-morbidity, frailty, and susceptibility to serious complications of infection or critical injury. There is therefore an increase in the probability of susceptibility to, and mortality from, critical illness that is consequent upon non-genetic factors.
Participation Group:
Patients will be identified and recruited in hospital during acute illness. Potential participants will be identified through hospital workers upon presentation at recruiting sites. The disease processes under study have a high mortality, so it is desirable to recruit patients as early as possible in the disease process.
Participants or an appropriate parent/guardian/consultee will be approached by staff trained in consent procedures that protect the rights of the patient and adhere to the ethical principles within the Declaration of Helsinki. Staff will explain the details of the study to the participant or parent/guardian/consultee and allow them time to discuss and ask questions. The staff will review the informed consent form with the person giving consent (or assent) and endeavour to ensure understanding of the contents, including study procedures, risks, benefits, and the right to withdraw. Participants who agree to participate (or their parent/guardian or consultee who declares their wishes to do so) will be asked to sign and date an informed consent form.
In view of the importance of early sampling, participants or their parent/guardian/consultee will be permitted to consent and begin to participate in the study immediately if they wish to do so. Those who prefer more time to consider participation will be approached again after an agreed time, normally one day, to discuss further.
Patients who meet the inclusion/exclusion criteria and who have given informed consent to participate directly, or have been consented by a parent/guardian or whose wishes have been declared by a consultee, will be enrolled to the study.
Samples and data will be collected according to available resources and the weight of the patient will be measured for children under 12 in order to prevent excessive volume sampling. Samples required for medical management will at all times have priority over samples taken for research tests. Aliquots or samples for research purposes should never compromise the quality or quantity of samples required for medical management. Wherever practical, taking research samples should be timed to coincide with clinical sampling. The research team will be responsible for sharing the sampling protocol with health care workers supporting patient management in order to minimise disruption to routine care and avoid unnecessary procedures.
Consent will be sought from patients who survive critical illness and regain capacity to give consent. At each of the follow-up sessions, the investigator gathering data will determine whether the patient has regained capacity. In the event that a patient continues to be incapacitated beyond the follow-up period, the local investigator will plan a subsequent capacity check at a specific date, after an interval to be determined by the nature of the incapacity. The planned dates of capacity checks on incapacitated survivors will be stored locally in the site file, together with a record of the outcome of each check.
Patients who decline to participate at this stage will be removed from the study. Where the patient cannot be recontacted despite best endeavours, they will remain in the study.
Withdrawal:
Participants are freely able to decline participation in this study or to withdraw from participation at any point without suffering any implied or explicit disadvantage. All patients will be treated according to standard practice regardless of whether they participate.
The following options of withdrawal will be made available to participants:
1. Partial withdrawal. Data WILL continue to be updated and used for research, but no further contact will be made with the participant
2. Full withdrawal.
• no further contact will be made with the participant;
• data will not be updated from health records;
• data will not be removed from research that is underway or has already been done, and an audit record will be maintained to confirm participation.
Consent version and life course follow-up
There are a series of consent versions as documented below:
• V1.08 (March 2020) - there are circa 1500 participants on this consent, allows for COVID research but not for longitudinal follow-up. Genomics are exploring reconsenting and the Confidentiality Advisory Group (CAG) as options for this cohort. Currently, these individuals will NOT be submitted to NHS Digital
• V2.1 (April 2020) - Included retention of NHS number and linkage to longitudinal follow -up, NHS Digital are named on the Participant Information Sheet (PIS). These individuals will be submitted to NHS Digital for data linkage
• V2.4 (July 2020) - In addition to already stated in v2.1, this covers community recruitment of participants and movement of identifiers. Sharing data with NHS Digital is again stated. These individuals will be submitted to NHS Digital for data linkage
PIS and Consent Version 2.1 and 2.4 are identical in their reference to accessing data from secondary data suppliers and include retention and collection of identifiable information and both name NHS Digital as a source of linkage. All individuals consented are collected with information including version of consent recruited under, thus genomics England can distinguish individuals based on consent. The most up to date protocol is listed on the website (https://genomicc.org/protocol/)
Datasets requested from NHS Digital:
This agreement permits access to the data sets requested, in line with COVID-19 related purpose. For participants recruited under the consultee process, the restrictions set out in Reg 3( 1) COPI will also apply.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data for Pandemic Planning and Research (GDPPR) extracts to assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
Hospital Episode Statistics (HES): Outpatients (OP), Admitted Patient Care (APC), Critical Care (CC) and Accident & Emergency (AE)/ Emergency Care Data Set (ECDS). These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Secondary Uses Service datasets. The minimal latency in availability of these datasets is highly desirable for the research objectives set out in this project.
Mental Health Data sets: The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
Cancer Registration Data sets: To see the incidence of cancer within the cohort
Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
COVID datasets: These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
Assessment of the datasets has been undertaken and NHS Digital are satisfied that they are necessary for the COVID-19 work being undertaken. Genomics have confirmed that all research which is approved from the GenoMICC study using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Genomics England Industry access:
Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help turn research findings into treatments, diagnostics and benefits for patients as soon as possible.
As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations to access the COVID data, however they are charged based on storage and compute, such as running their own bioinformatics pipeline. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected.
The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a 'critical friend' and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's landmark data set.
The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
[1 paragraph unchanged]
Additionally, as Genomics England will be processing health data - a special category of personal data, they will also be processing data under Article 9 (2)(j) as processing is necessary for archiving purposes in the public interest. Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The lawful basis for processing special category data under the UK GDPR is:
Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
[4 paragraphs unchanged]
The funding is provided the Department of Health and Social Care. The funding is specifically for the purposes described.
The funder will have no ability to suppress or otherwise limit the publication of findings.
Lifebit provides IT support to Genomics England.
Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.
SUB-LICENCING:
Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.
It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.
The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.
All requests for data access will be subject to the following considerations:
• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).
• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).
• Provision of a signed Genomics England data access agreement to the Access Review Committee.
• Prioritisation of access according to resource availability.
• Facilitation of high-quality health research
Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:
• Diagnostic discovery
• Pre-clinical research
• Clinical Trials Referral
• Real World Evidence / Market Access
Approval process for Commercial organisations for access to the NGRL:
Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.
The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.
Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.
Approval process for GeCIP users (academic) of the NGRL:
• Researcher visits Genomics England website to enrol as a GECIP member
• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)
• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.
• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.
• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.
• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.
All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:
• Identification of diagnoses originally missed by the standardised pipeline
• Feedback of new diagnoses to patients
• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone
Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.
Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.
AIRLOCK POLICY:
The following rules are applied to all airlock requests:
1. All relevant details of the summary results to be transferred must be provided with every request.
2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.
3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.
4. Summary results requested for transfer are assessed using the following criteria:
a. whether the request aligns with the users ARC approval in full;
b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;
c. any data security implications;
d. any disclosure risks;
e. the technical feasibility and associated cost of the request;
f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;
g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.
The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:
• Technical Lead
• User Community Representative
• Bioinformatics Director
• Caldicott Guardian
• Chief Scientist representative
The Data will be processed worldwide.
Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.
GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
• UK academic research institutions (e.g., universities, research institutions etc.)
• NHS trusts or authorities
• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project
• Foreign universities and research institutions that carry out significant research activity
• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))
• Foreign healthcare organisations (private or public) that undertake significant research activity
To be eligible for data access as a GeCIP member, applicants must meet these requirements:
• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.
• Their host institution has verified that they are affiliated with that institution.
• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).
• The GeCIP domain lead has approved the application.
• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.
• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).
The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.
GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.
All personnel accessing the Data have been appropriately trained in data protection and confidentiality.
The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:
> National Cancer Registration and Analysis Service (NCRAS)
and uncurated NCRAS data
> Secure Anonymised Information Linkage (SAIL) data; Welsh data
> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)
> NHS Trusts data
> Data feeds from the Intensive Care National Audit Registry (ICNARC)
> The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
The Data will not be linked with any other data.
The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.
To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.
Data accessed under sub licence will not be re-identified.
Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.
Data shared through the Airlock process is aggregate data only and is t herefore not personal data so does not require a legal basis under the UK GDPR.
A release register detailing any sub licences and onward sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse
Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract.
In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users.
NHS England will require the ability to audit the sub licensee.
Processing activities
Genomics England provide NHS Digital with a cohort for linkage, and they receive data from NHS Digital on a monthly basis. Every month, Genomics England provide an updated cohort to NHS Digital who provide the historical data for the extra cohort members. The cohort is already flagged with NHS Digital so Genomics England will only receive the historical data for the extra cohort members each month.
Genomics England will transfer data to NHS England. The data will consist of identifying details NHS Number, Date of Birth, Gender, and a unique person ID for the cohort to be linked with NHS England data.
The first stage of processing focuses on quality verification. This ensures that the data set is complete, accurate and complies with the NHS data dictionary or relevant specification. Participant identifiers in the dataset are verified against Genomics England's participant details and any updates required to identifiable data fields, e.g., dates of birth, are highlighted. Finally, the data set is reviewed against recent participant withdrawals so that any withdrawals notified after the data application was made can be removed from the data sets.
NHS England will provide the relevant records from the datasets listed in this agreement to Genomics England. The Data will
Following this the data are de-identified, as all subsequent processing can be performed without direct identifiers. Genomics England has compiled lists of identifiable and sensitive fields for each data set in line with details provided by NHS Digital and following internal review of data sets. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to a secure Trusted Research Environment (TRE) hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data.
• contain directly identifying data items including Names, Postcode, Cause of Deaths, Place of Birth, Cancer Registration Number, which are required to provide maximum insight, and therefore maximum value to the researchers accessing the Data.
The second stage of processing involves the selection of a de-identified cohort of participants that fulfil a specific research request. Researchers are members of a Genomics England Clinical Interpretation Partnership (GeCIP) or the Discovery Forum. Research requests are assessed to ensure that they are included in the approved use purposes set out in the Genomics England Protocol and fall within the scope of the relevant GeCIP or the Discovery Forum. Researchers declare any data they wish to bring into the TRE and any tools they wish to use for analysis.
The NHS England Data is pseudonymised within the AWS cloud and is then loaded into the NGRL. Raw, identifiable files are kept in a secure location on AWS.
The third stage of processing is the analysis of the de-identified data sets within the TRE. Researchers perform all the analysis and processing within the environment: they do not extract de-identified data. Results data are placed in a secure folder for anonymisation verification before extraction.
First stage of processing – quality verification
There will be no data linkage undertaken with NHS Digital data provided under this agreement that is not already noted in the agreement.
• This ensures that the data set is complete, accurate and complies with the NHS data dictionary or relevant specification. Participant identifiers in the dataset are verified against Genomics England's participant details and any updates required to identifiable data fields, e.g., dates of birth, are highlighted. Finally, the data set is reviewed against recent participant withdrawals so that any withdrawals notified after the data application was made can be removed from the data sets.
THE TRUSTED RESEARCH ENVIRONMENT (TRE)
Second Stage of processing- selection of a de-identified cohort of participants that fulfil a specific research request
All research analysis on the Genomics England dataset will only be carried out via a secure analysis environment hosted within the Genomics England data centre - the Genomics England TRE. Analytical tools and applications are available within the TREt. No sequencing or clinical data are made available for download, users cannot copy or paste out of the TRE, and there is limited internet access within it (i.e. whitelisted sites). Movement of files into and out of the TRE is governed via an 'Airlock' Policy.
• Researchers are members of a Genomics England Clinical Interpretation Partnership (GeCIP) or the Discovery Forum. Research requests are assessed to ensure that they are included in the approved use purposes set out in the Genomics England Protocol and fall within the scope of the relevant GeCIP or the Discovery Forum. Researchers declare any data they wish to bring into the NGRL and any tools they wish to use for analysis.
Academic researchers access the TRE by applying to be a member of a Genomics England Clinical Interpretation Partnership (GeCIP) domain. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
Third Stage of processing- analysis of the de-identified data sets within the NGRL
o UK academic research institutions (e.g. universities, research institutions etc.)
• Researchers perform all the analysis and processing within the environment: they do not extract de-identified data. Results data are placed in a secure folder for anonymisation verification before extraction.
o NHS trusts or authorities
The Data will not be transferred to any other location.
o UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project
The Data will be stored on the NGRL and the AWS Cloud at Genomics England.
o Foreign universities and research institutions that carry out significant research activity
Genomics England stores NGRL data on the Cloud provided by Amazon Web Services (AWS).
o UK and foreign governmental departments that carry out significant research activity (e.g. Medical Research Council, National Institute for Health, Public Health England)
The Data will be accessed by authorised personnel via remote access.
o Foreign healthcare organisations (private or public) that undertake significant research activity Membership is not open to those who are self-employed or employed by:
The Controller(s) must confirm and provide evidence upon audit by NHS England that access via any remote device complies with the data security obligations within this DSA and the Data Sharing Framework Contract.
o private UK healthcare institutions
For remote access:
o commercial companies.
- Remote access will only be from secure locations situated within the territory of use (as further restricted elsewhere within the DSA if so done) stated within this DSA;
o To be eligible for data access as a GeCIP member, applicants must meet these requirements:
- Access controls granting users the minimum level of access required are in place;
o Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.
- Remote access is only via secure connections (e.g., VPNs or secure protocols) to protect Data;
o Their host institution has verified that they are affiliated with that institution.
- Multifactor authentication (MFA) is required for remote access;
o The applicant's GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee [ARC] (see below).
- Device security, including up-to-date software and operating systems, antivirus software, and enabled firewalls are utilised for the remote access;
o The GeCIP domain lead has approved the application.
- All remote access is undertaken within the scope of the organisation’s DSPT (or other security arrangements as per this DSA) and complies with the organisation’s remote access policy.
Following approval, GeCIP researchers must sign a specific agreement covering their behaviour and working practice within the data infrastructure. Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
The above applies in addition to any condition set out elsewhere within the DSA (e.g. who may carry out processing, and for what purpose).
Commercial researcher access to the TRE:
The Data will be processed worldwide.
Genomics England operates a membership-based forum - the Discovery Forum - which is open to a range of companies world-wide and allows access to the TRE. It provides a platform for collaboration between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Data will be released into the NGRL where it will be linked with data feeds from the Intensive Care National Audit Registry (ICNARC) and The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
Each Discovery Forum member signs a Data Access Agreement with Genomics England. This states the research purposes which the company is authorised to carry out and stipulates the number of genomes sequences that can be accessed. It covers the Company's behaviour and working practices: in particular it binds users to Genomics England's Airlock Policy, Information Governance, IT Security and Data Protection Polices. Companies need to nominate named individuals to be their Researchers who must complete information governance training before accessing data. Once the Data Access Agreement is in place, each research project undertaken by the Company within the TRE must receive prior ARC approval.
Access to confidential patient identifiable Data is restricted to an extremely limited number of employees of Genomics England, accessible on AWS.
Discovery Forum members access the TRE in a similar manner to GeCIP Researchers: all research is carried out within the TRE, and any movement of results out of the environment occurs only through the Airlock Process.
Substantive employees of Genomics England and researchers who are a member of the GeCIP and Discovery Forum will process the Data for the purposes described above.
THE ACCESS REVIEW COMMITTEE (ARC)
ARC provides an independent examination of requests for data access, with regards to the acceptable uses of the Genomics England dataset which are outlined in The National Genomic Research Library Protocol the 100,000 Genomes Project Protocol and Data Access and Acceptable Uses Policy. The ARC comprises external scientific experts, patient representatives and members of Genomics England's Participant Panel which is made up of participants and parents/carers involved in the 100,000 Genomes Project.
THE AIRLOCK PROCESS
The Genomics England Trusted Research Environment has been developed with the intention that all data analysis is carried out within it and that the only data to leave it are summary results. An Airlock process has been established which enables material (data, files, tools etc) to be moved in or out of the Research Environment in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access. Removal of summary results therefore requires an Airlock request.
The following rules are applied to all Airlock requests:
1. All relevant details of the summary results to be transferred must be provided with every request.
2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.
3. All summary results transferred will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.
4. Summary results requested for transfer are assessed using the following criteria:
• whether the request aligns with the users ARC approval in full
• whether the request can clearly be demonstrated to be aligned with a registered project in the Research Environment
• any data security implications
• any disclosure risks
• the technical feasibility and associated cost of the request
• when importing data, its scientific value to the community of researchers within the Research Environment, and when and how it will be shared
• when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.
The Airlock process is governed by the Airlock Policy, which defines the process and governance of the Airlock process. A set of Airlock Policy Guidelines presents the rules-of-thumb/principles that will be referenced by both the researcher (during preparation of analysis results) and the output checker (during output-checking).
Analysed results are inspected to ensure they cannot be used to disclose the identity of the participants. Checking of summary results by the Airlock Review Team is governed by a set of principles that guide individual decisions. By using a principles-based approach where each case is assessed individually the security of the dataset is maintained by exporting only appropriate data. Review of transfer requests resulting in public-sharing/publication of data will be checked more stringently. Any approved Airlock export can only be used for the specific use detailed in the original request. The Research Environment contains pseudonymised longitudinal data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be used and what can be exported. Where an export contains pseudonymised longitudinal data, Genomics England will always apply the requirements placed on them as conditions of having access to the data.
All Airlock requests go through a robust approval process and the Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. Where a precedent has been set and a clear set of principles and rules are in place for types of research, the Airlock Manager can approve the request. For more complicated requests or where no precedent has been set these will go to the Airlock Committee and be reviewed by the Airlock Review Team.
The Airlock Review Team is a delegation of the Genomics England Chief Scientist responsible for oversight of all airlock requests in accordance with the Airlock Policy and the groups Terms of Reference. It comprises:
• Technical Lead
• User Community Representative
• Bioinformatics Director
• Caldicott Guardian
• Chief Scientist representative
SUB-LICENCING
Genomics England has developed the TRE to allow registered third parties to access pseudonymised versions of the data that it holds, for the purposes of approved research. The TRE contains External Data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be accessed. Genomics England will always apply the requirements placed on them as holders of External Data to users of the TRE as a condition of having access to the data. The data is NOT for onward sharing outside of the TRE.
DATA CONTROLLERS
For clarity, the University of Edinburgh is responsible for acquisition of primary clinical data. That relates to data acquired at patient registration.
Genomics England in its provision of whole genome sequencing are applying to NHS Digital for secondary clinical data to link to the genomic data. With regards to data provided by NHS Digital, Genomics England are the sole data controller. To this end, staff and academics from the University of Edinburgh are required to join a GeCIP to access the secondary clinical data from NHS Digital within the TRE.
Genomics England provides NHS Digital with linking data in order to receive longitudinal data sets. These data sets are delivered to Genomics England by NHS Digital on a monthly basis having been approved by the NHS Digital IGARD. Genomics England identifies the linking data and agrees with NHS Digital the scope of the longitudinal data being provided. Genomics England determines the method of de-identification and storage within the TRE and secures this data for use by approved researchers only. Genomics England determines who these researchers are.
Genomics England is the Data Controller for longitudinal data sets processed in the Genomics England Research Library.
Researchers in academic, educational or commercial organisations
Access to deidentified data in the TRE which will include longitudinal data sets (HES etc) provided by NHS Digital. Access to the TRE only allowed under access agreement.
The individual researchers are Data Controllers when carrying out research within the TRE. Data disseminated under this agreement for COVID-19 purposes will be restricted to the GEL Covid TRE. Only COVID-19 research approved studies will be granted access to the data. All research which is granted access for COVID-19 purposes must be employed or engaged for the purposes of the health service as the request for data is to support research that has been set as a priority by the CMO. Research which is approved using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
DATA PROCESSORS
o Only summary level data can be removed from the environment.
o Approved researchers will only be able to access Lifebit’s PaaS [Platform-as-a-service] CloudOS [Operating System] through a virtual desktop.
o Secondary data will be ingested into CloudOS.
o CloudOS will be hosted within GEL’s London AWS (Amazon Web Services) environment - All data is encrypted in transit and at rest.
o CloudOS controls access to the secondary data.
o A security and DPIA (Data Protection Impact Assessment) will be conducted prior to loading live
Genomics England have ceased using UKCloud (UKC) as a data processor after migrating all data to Amazon Web Services (AWS). The data in UKC was destroyed in March 2022. A detailed plan has been created which ensures the deletion runbook was followed, which is aligned to Genomics England standard operating procedures.
Lifebit:
Lifebit has been selected as platform partner to deliver the TRE after reviewing several proposals. The UK-based SME (Small Medium Enterprise) offered a proven and innovative technology solution offering a blend of robustness and ease of use. Lifebit CloudOS (Operating System) provides a secure and collaborative workspace to enable researchers to easily perform COVID-19 genomic data analysis. The platform will deliver an intuitive, integrated and collaborative user experience and enables fast, effective COVID-19 research outcomes across a wide range of academic and biotech/pharma researchers with varying levels of technical competency.
DATA MINIMISATION
Genomics England's Trusted Research Environment (TRE) aligns to the current NHS guidance and is currently cited as an example of best practice. The detail to support this is below.
In the HDR UK TRE Principles and Best Practices paper from December 2021, the Genomics England model is explicitly described (page 17) and this document has a foreword authored by the Director of Data Policy, NHSX and Director of Tech Policy, NHSX. https://www.hdruk.ac.uk/news/new-principles-published-to-improve-public-confidence-in-access-and-use-of-data-for-health-research-through-trusted-research-environments
Conceptually, in terms of the 5 safes framework, the safes work together to protect the privacy of the individual. The TRE model makes 4 of the safes: safe-setting, safe-people, safe-projects and safe-outputs very strong, meaning that it’s possible to allow the criteria for safe-data to be relaxed (i.e., de-identification only) while maintaining the overall level of privacy protection. The TRE model effectively moves data minimisation to the output stage - ’safe outputs’, i.e., the airlock, where the airlock managers and the airlock committee do review requests to ensure summary data is minimised to only what is necessary to demonstrate externally a particular research result. The data released by the Airlock project is aggregated summary data. Given these measures, there has even been a discussion about data within TREs being regarded as "functionally anonymous” because of these safeguards, which would put them outside the constraints of GDPR.
The risk is further minimised as all the researchers are under contractual obligation not to re-identify individuals from the data. By allowing researchers access to all the data it allows "hypothesis free research" and means they may discover things they didn't suspect.
The full dataset is “adequate and relevant” for all researchers to access and the TRE framework provides the required limitation. Research in the RE is around discovering genome phenome relationships and given the complexity of the genome and how little we understand about it, it’s impossible to predict what phenotype terms may be involved, so researchers need to be able to explore all of them, rather than repeatedly requesting different datasets.
Each individual participant has agreed and is aware that their data is in the research environment and accessed by researchers.
Access to the data is necessary for this purpose to ensure the research can be carried out and the most value extracted from the data. We do have a wide range of controls outlined above to make sure the risk is minimised.
Both Article 89 and Article 25(1) of the UK GDPR support this approach.
A Data Protection Impact Assessment was conducted in November 2021 prior to the transfer of the data to the new environment and is being continually assessed and updated by the data protection team.
COHORT SIZE
Briefly, there will be effectively 2 cohorts, one is the 100K Project (CONTROL COHORT) which is about 92,000 participants. This will be a near enough static list.
The second cohort is the COVID-19 recruited cohort participants.
Expected output
[1 paragraph unchanged]
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the
TRE
NGRL
to support this ongoing research. Each research group who accesses this data
[12 words unchanged]
aim to deliver against. Specific examples of journal outputs are as follows:
[4 paragraphs unchanged]
Expected measurable benefits
Access to the data will enable the research to discover new rare and common variants alongside new
multi-omic
biomarkers that underpin host response to infection, allow investigation of the impact
[71 words unchanged]
these resources and the presence of parents helps compensate for this problem.
[22 paragraphs unchanged]
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Trusted Research Environment (TRE).
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England TRE where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw. This is a ground-breaking and far-reaching study which will harness the UK͛s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
The TRE also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public Health England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the TRE monthly in order to continue support for, and to further develop, this ground-breaking resource.
Unchanged: Benefits reported.
Objective for processing
Genomics England requires access to NHS England data for the purpose of the following work programme:
The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19.
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS England and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England National Genomics Research Library (NGRL) to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and its outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By utilising the GenOMICC consortium’s prospective study design that leverages existing recruitment infrastructure in critical care, Genomics England will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
The following NHS England Data will be accessed:
• Hospital Episode Statistics
o Admitted Patient Care
o Accident & Emergency
o Critical Care
o Outpatients
• Emergency Care Data Set (ECDS)
These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
• Diagnostic Imaging Dataset (DID) necessary to provide invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
• Civil Registrations of Death- Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
• Cancer Registration - necessary to see the incidence of cancer within the cohort
• Covid-19 SGSS First Positives (Second Generation Surveillance System)
• Covid-19 Vaccination Status
• Covid-19 Hospitalization in England Surveillance System
These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
• Mental Health Services Data Set- The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
• Covid-19 General Practise Extraction Service assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
The level of the Data will be
• Identifiable- many indirect identifiable Data items have been requested because they provide valuable Data that can help researchers make new scientific and medical discoveries. All directly identifiable Data items will either be removed or transformed according to best practice agreed with NHSE. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to the NGRL hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data
The Data will be minimised as follows.
• Limited to a study cohort identified by Genomics England of;
• ~100,000 patients who consented to participate. ~30,000 of which recruited by Geonomics meeting the study's COVID-19 eligibility criteria, and ~71,243 as a control cohort.
• Genomics England request full history of patient Data to provide maximum insight, and therefore maximum value to the researchers accessing the Data. Because of the wide scope of the proposal, there are no other alternative or less intrusive ways of achieving the purpose described.
Genomics England is the controller as the organisation responsible for ensuring that the Data will only be processed for the purpose described above. Genomics England also process the Data.
The University of Edinburgh is responsible for acquisition of primary clinical data. That relates to data acquired at patient registration. Genomics England in its provision of whole genome sequencing are applying to NHS England for secondary clinical data to link to the genomic data.
Staff and academics from the University of Edinburgh are required to join a Genomics England Clinical Interpretation Partnership (GeCIP) to access the secondary clinical data from NHS England within the NGRL. Please see the processing activities of this DSA for further detail on GeCIP.
The lawful basis for processing personal data under the UK GDPR is:
Article 6(1)(f) - processing is necessary for the purposes of the legitimate interests pursued by the controller or by a third party.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
The lawful basis for processing special category data under the UK GDPR is:
Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The funding is provided the Department of Health and Social Care. The funding is specifically for the purposes described.
The funder will have no ability to suppress or otherwise limit the publication of findings.
Lifebit provides IT support to Genomics England.
Amazon Web Services (AWS) provides IT back up services to Genomics England and will store copies of the Data as contracted by Genomics England.
SUB-LICENCING:
Genomics Clinical Interpretation Partners (GeCIP) members (Academic research organisations), and members of the Discovery Forum (Commercial organisations) will also have access to the pseudonymised Data within the NGRL, subject to internal approval by Genomics England. NHS England Data is combined with the genomic and sample data within the NGRL, providing a more comprehensive medical history, and going forward, a more comprehensive patient journey which will be a valuable resource for medical research. All applications have to provide health and social care benefits and are reviewed by a panel (the Access Review Committee (ARC)) before access is granted.
It is anticipated that the volume of sub-licences will be 150-200 per year. The GeCIP sub licence agreement is indefinite, until it is terminated by either the GeCIP member or Genomics.
The Data Access Agreement for Discovery Forum Members has a specified term, normally 12 months, at which point the company and Genomics can choose to renew or not.
All requests for data access will be subject to the following considerations:
• Protection of data subjects (honouring commitments made to them, acting within the scope of consent and according to conditions of Research Ethics Committee approval).
• Compliance with legal and regulatory requirements General Data Protection Regulation 2018, Data Protection Bill 2017, Freedom of Information Act 2000, NHS Act 2006, Health and Social Care Act 2012, the Common Law Duty of Confidentiality, Human Tissue Act 2004 and applicable requirements from organisations affiliated with the Health Research Authority, including Research Ethics Committees and the Confidentiality Advisory Group (CAG).
• Provision of a signed Genomics England data access agreement to the Access Review Committee.
• Prioritisation of access according to resource availability.
• Facilitation of high-quality health research
Commercial partnerships are crucial to achieving the aims of the NGRL and are achieved through the Discovery Forum. As with the non-commercial academic research led by GeCIP, commercial research aims to bring benefit to the patients and, through the use of the Data, inform development of platforms and tools for future diagnostic discovery. Commercial research can be broadly categorised into four themes that answer different questions along the typical Research and Discovery Biopharmaceutical Pipeline. At a high level they are divided into:
• Diagnostic discovery
• Pre-clinical research
• Clinical Trials Referral
• Real World Evidence / Market Access
Approval process for Commercial organisations for access to the NGRL:
Discovery Forum applications from a commercial organisation would be reviewed for suitability by the Partnership Development (PD)Team. The PD Team consider the credentials of the applying organisation including consideration of adverse public perception and reputational risk from approving data access for that organisation. If the PD Team feel appropriate, they are then passed on to be scrutinised by the independent Access Review Committee (ARC). ARC is constituted of Participant Panel members and senior individuals from various scientific and medical backgrounds. ARC assess the company’s research proposal, including patient/participant involvement, potential future value to patients/the NHS and the ethics of the proposal.
The ARC will assess whether there has been any Patient and Public Involvement and Engagement (PPIE) informing the research questions and design. For many commercial applications that are exploring early-stage research and development (R&D), for example, target identification and validation, there will not have been any PPIE because the research may be tied to exploring fundamental biological mechanisms and pathways rather than particular conditions or phenotypes. If there has been PPIE, the ARC will determine whether it has adequately informed the research questions and design, and whether there is a commitment to ongoing PPIE and transparency following the outcomes of the research. Although PPIE is not a requirement of applications, ARC encourage applicants to consider at what stage in their R&D process it would be appropriate to consult with patient advocacy and participation groups.
Genomics England will only work with companies that are aligned with its strategy and mission to bring the benefits of genomic medicine to everyone. The Partnerships Development team will assess whether a company seeking access to NGRL data is working in the cancer or rare disease diagnostics and therapeutics space, or supporting UK Government strategic scientific initiatives – if not, Genomics England would not permit an application to ARC in the first place. All applications must conform to the acceptable uses set out in the REC-approved NGRL protocol. If the research proposal is for later stage research that has a clear pathway to intended patient or health system benefit, the ARC would expect to see this articulated as part of the rationale for seeking access to NGRL data. Given the early stage of much commercial genomics research, not all accepted applications will be able to demonstrate a clear explanation of the expected healthcare benefits.
Approval process for GeCIP users (academic) of the NGRL:
• Researcher visits Genomics England website to enrol as a GECIP member
• Completion of onboarding process; Verification by their institution (institution will be required to sign a Genomics participation agreement and appoint a membership secretary), verification of their self-stated qualifications and areas of research interest by the GEL Scientific Manager to join their domain of choice, take the IG and GECIP rules training course and pass test with at least 80%. They are then able to access the NGRL and the Research Portal (the area where prospective GeCIP applicants can apply and register their research project)
• Within 3 months of gaining access they need to either submit a research proposal for Genomics England approval, which currently has to fit with the Detailed Research Plan for their domain, or join another registered project. Otherwise they will lose access.
• On an annual basis, complete a survey sent out by Genomics England giving details of their research progress and any outputs, to aid reporting to ARC.
• Any data they wish to either import or export to/from the NGRL has to be approved by Airlock (Airlock policy is described below) as not being personally identifiable.
• If a researcher has not accessed the NGRL, the Research Portal, or logged into their GEL account to gain access to either of the previous for 6 months their account will be deactivated.
All research activities undertaken in the NGRL aim to enrich the existing dataset via one or multiple routes:
• Identification of diagnoses originally missed by the standardised pipeline
• Feedback of new diagnoses to patients
• Mobilising samples which can help to identify diagnoses that were missed through analyses of WGS alone
Researchers can access pseudonymised Data through NGRL under sub licence. The only Data allowed to be exported are summary results. An airlock policy has been established which enables material (data, files, tools etc) to be moved in or out of the NGRL in a controlled and supervised manner; facilitating research and discovery, while maintaining control of security and access.
Data accessed under sub licence is only granted to named individuals identified to Genomics England who agree to comply with the Airlock policy, Information Governance and IT Security Policy. Before being provided with credentials necessary to access the NGRL a Company Researcher must complete information governance training which shall be provided by Genomics England.
AIRLOCK POLICY:
The following rules are applied to all airlock requests:
1. All relevant details of the summary results to be transferred must be provided with every request.
2. All summary results transferred must be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any summary results rejected along with the reason for the rejection.
3. All imports will be checked for viruses and malware and those failing this test will be rejected. It is the responsibility of the requestors to resolve such issues before re-submitting the file for transfer.
4. Summary results requested for transfer are assessed using the following criteria:
a. whether the request aligns with the users ARC approval in full;
b. whether the request can clearly be demonstrated to be aligned with a registered project in the NGRL;
c. any data security implications;
d. any disclosure risks;
e. the technical feasibility and associated cost of the request;
f. when importing data, its scientific value to the community of researchers within the NGRL, and when and how it will be shared;
g. when importing data, checks will be performed to ensure that the data importer owns the data and holds the correct consents and approvals.
The Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. For more complicated requests or where no precedent has been set these will go to the airlock committee for review and a decision. The airlock committee is a delegation of the Genomics England Chief Scientist who responsible for oversight of all airlock requests in accordance with the airlock policy. The committee comprises of:
• Technical Lead
• User Community Representative
• Bioinformatics Director
• Caldicott Guardian
• Chief Scientist representative
The Data will be processed worldwide.
Access is restricted to substantive employees of Genomics England, Genomics Clinical Interpretation Partners (GeCIP) members, and members of the Discovery Forum, who have authorisation from the Principal Investigator.
GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
• UK academic research institutions (e.g., universities, research institutions etc.)
• NHS trusts or authorities
• UK and foreign charitable organisations directly related to the focus of the 100,000 Genomes Project
• Foreign universities and research institutions that carry out significant research activity
• UK and foreign governmental departments that carry out significant research activity (e.g., Medical Research Council (MRC), National Institute of Health (NIH), Public Health England (PHE))
• Foreign healthcare organisations (private or public) that undertake significant research activity
To be eligible for data access as a GeCIP member, applicants must meet these requirements:
• Their host institution has signed a GeCIP Participation Agreement, which outlines the key principles that members of each institution must adhere to, including the Intellectual Property and Publication Policy.
• Their host institution has verified that they are affiliated with that institution.
• The applicant’s GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee (see below).
• The GeCIP domain lead has approved the application.
• Following approval, GeCIP researchers must sign a specific agreement (‘GeCIP rules’) covering their behaviour and working practice within the data infrastructure.
• Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
All applications have to provide health and social care benefits in England and are reviewed by a panel (the Access Review Committee (ARC)).
The ARC provides an independent examination of requests for data access. The ARC comprises external scientific experts, patient representatives and members of Genomics England’s Participant Panel.
GeCIP users will be granted access to all data and knowledge held within the NGRL. Each GeCIP domain will have access to its own private shared area of the NGRL for data storage and collaboration. The secure virtual desktop infrastructure will provide the ‘workspace’ for clinical teams, research groups and trainees to undertake their work.
All personnel accessing the Data have been appropriately trained in data protection and confidentiality.
The Data will be linked at person record level with the patient’s genetic data within the NGRL. This includes the following data:
> National Cancer Registration and Analysis Service (NCRAS)
and uncurated NCRAS data
> Secure Anonymised Information Linkage (SAIL) data; Welsh data
> Patient samples (e.g., blood, saliva, tissue, RNA, plasma and serum)
> NHS Trusts data
> Data feeds from the Intensive Care National Audit Registry (ICNARC)
> The UK Health Security Agency (UKHSA) viral genomic data and associated metadata
The Data will not be linked with any other data.
The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised Dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Dataset.
To protect patient confidentiality, access to the NGRL will be granted only for specific, approved purposes in accordance with informed consent. Any attempted use beyond the specified purpose may lead to exclusion and possible legal action, where appropriate.
Data accessed under sub licence will not be re-identified.
Genomics England rely on GDPR Article 6 (1)(f) for the personal data and Article 9(2)(j) for the special category data shared within the NGRL.
Data shared through the Airlock process is aggregate data only and is t herefore not personal data so does not require a legal basis under the UK GDPR.
A release register detailing any sub licences and onward sharing can be found here: https://research.genomicsengland.co.uk/research-registry/browse
Genomics will take responsibility for the actions and omissions of all sub licences and breach of a sub licence will automatically be regarded as breach of the Data Sharing Framework Contract.
In the event of termination or expiry of the Data Sharing Framework Contract between NHS England and the applicant, Data from NHS England will be removed from the NGRL, preventing access to the Data for all users.
NHS England will require the ability to audit the sub licensee.
Expected output
Combining genomic sequence data from COVID participants together with their medical records has created a ground-breaking research resource. Researchers are currently studying this data to be able to deliver outputs that will have tangible benefits to healthcare and patients. Understanding the symptoms and impact of COVID on people with differing genetic make-ups is hoped to lead to successful treatment of the disease being investigated.
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the NGRL to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. Specific examples of journal outputs are as follows:
• A first update on mapping the human genetic architecture of COVID-19 (2022) Nature 608(7921):E1-E10. doi: 10.1038/s41586-022-04826-7
• Mapping the human genetic architecture of COVID-19 (2021) Nature 600(7889):472-477. doi: 10.1038/s41586-021-03767-x
• Whole-genome sequencing reveals host factors underlying critical COVID-19 (2022) Nature 607(7917):97-103. doi: 10.1038/s41586-022-04576-6
• Genetic mechanisms of critical illness in COVID-19 (2021) Nature 591(7848):92-98. doi: 10.1038/s41586-020-03065-y
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated August 2022:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
May 2021:
• Identified 4 targets that can utilise and repurpose currently used medications for treatment and management of severe covid.
• Found evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19. This mechanism may also be amenable to targeted treatment with existing drugs.
These results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
Mapping the human genetic architecture of COVID-19
Dec 2021
• Contributed to the meta-analyses that consist of up to 49,562 patients with COVID-19 from 46 studies across 19 countries.
• Identified 13 genome-wide significant loci that are associated with SARS-CoV-2 infection or severe manifestations of COVID-19. Several of these loci correspond to previously documented associations to lung or autoimmune and inflammatory diseases.
• They represent potentially actionable mechanisms in response to infection.
• There is evidence of a causal role for smoking and body-mass index for severe COVID-19 although not for type II diabetes.
This working model of international collaboration underscores what is possible for future genetic discoveries in emerging pandemics, or indeed for any complex human disease.
May 2022
• Identified TMPRSS2 variant has a protective effect against severe COVID-19
• This is a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19
Whole genome sequencing of 7,491 GenOMICC participants admitted with severe COVID
Jul 2022
• Identified 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2)
• Found evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease.
• Found evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets.
The results show that comparisons between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Aug 2022
• Contributed to the meta-analyses bringing together 60 studies from 25 countries for 3 COVID-19 phenotypes.
• Between-study heterogeneity rather than differences across ancestries are a more likely explanation for the observed heterogeneity in the effect sizes across studies
• Identified that the variant rs35705950:G>T located in the promoter of MUC5B (11p15.5) is protective against hospitalization
• Identified that rs190509934:T>C, which is upstream of ACE2, is associated with decreased susceptibility risk. Recent results have shown that the rs190509934:T>C variant lowers ACE2 expression, which in turn confers protection against SARS-CoV-2 infection12.
The biological insights gained by this expansion of the COVID-19 Host Genetic Initiative showed that increasing sample size and diversity remain a fruitful activity to better understand the human genetic architecture of COVID-19.
The first genomes from severe volunteers came through in June 2020.
By June 2021, interim findings from the study were published. These findings have already helped doctors make better decisions when treating patients with COVID-19, resulting in better outcomes. In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy.
Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
DARS-NIC-374190-D0N1M-v5.5 19 October 2022 to 18 October 2023
- Title
- R26 - GENOMICS ENGLAND: GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 21
- Files released
- 1,304
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); COVID-19 Vaccination Status; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v4.1
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2022-10-19 | |
| End date | 2023-10-18 | |
| Demographics: legal basis | Not stated |
Objective for processing
December 2021 Amendment- Version 4.1 of the agreement seeks to change the way in which the common law duty of confidentiality for the cohort is being addressed moving from COPI to consent and in addition requests access to vaccination data.
[7 paragraphs unchanged]
6. To provide access to these data sets via the Genomics England
Trusted
Research Environment
(TRE)
to international and national academia and industry and facilitate international collaboration on COVID-19.
[1 paragraph unchanged]
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and
it’s
its
outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel.
[1 paragraph unchanged]
The variable response to COVID-19 suggests that, as with susceptibility to other
[23 words unchanged]
for Rare Disease it is known that rare variants cause immunodeficiency. By
undertaking a
utilising the GenOMICC consortium’s
prospective study design that leverages existing recruitment infrastructure in critical care,
together with
Genomics
England, NHS England, Devolved Nations and Public Health
England
(PHE) infrastructure this research
will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
[4 paragraphs unchanged]
Data will be released into the Genomics England
Research Environment
TRE
where it will be linked with associated clinical data as well as
[22 words unchanged]
England and data feeds from the Intensive Care National Audit Registry (ICNARC).
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable
Research Environment
TRE
for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new
Research Environment
TRE
(COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables
[13 words unchanged]
with varying levels of technical competency. This offers a major upgrade to
our
Genomics England’s
current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling,
[25 words unchanged]
England’s platform infrastructure, which is a key enabler to the research community.
[8 paragraphs unchanged]
Details of the sublicense model via the Genomics England
Research Environment
TRE
are supplied within the processing activities section of this agreement.
[3 paragraphs unchanged]
• To that end,
our research environment
Genomics' TRE
is a leading light in security and standards as per HDR UK (https://zenodo.org/record/5767586#.YcPE9hPP27M) and an author is an employee of Genomics England.
• Furthermore,
our
Genomics'
Airlock policy is clear and explicit about the level of data that can be extracted and does not permit row level/ individual data extraction
[3 paragraphs unchanged]
The associated COVID-19
data-sets
datasets
(including NHS111 and CV19 Testing Data in particular which will be requested
[20 words unchanged]
participants were monitoring their symptoms and was there an associated poor prognosis.
[5 paragraphs unchanged]
The genetic mechanisms of susceptibility to infection are likely to be highly
[20 words unchanged]
et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions
(e.g.
(e.g.,
small molecules to inhibit an enzyme or receptor that is dysfunctional in
[23 words unchanged]
difficult for any one pathogen to evolve resistance to such a therapy.
[23 paragraphs unchanged]
• V1.08 (March 2020) - there are circa 1500 participants on this
consent ,
consent,
allows for COVID research but not for longitudinal follow-up. Genomics are exploring reconsenting and
CAG
the Confidentiality Advisory Group (CAG)
as options for this cohort. Currently, these individuals will NOT be submitted to NHS Digital
• V2.1 (April 2020) - Included retention of NHS number and linkage to longitudinal follow -up, NHS Digital are named on the
PIS.
Participant Information Sheet (PIS).
These individuals will be submitted to NHS Digital for data linkage
[3 paragraphs unchanged]
The current application to
This agreement permits
access to the data sets
requested are
requested,
in line with COVID-19 related
purpose, in terms of
purpose. For participants recruited under the consultee process,
the restrictions set out in Reg
3(1) COPI. However, based on the now REC approved protocol, we are looking to move this application to Consent basis.
3( 1) COPI will also apply.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data
[22 words unchanged]
comorbidities and underlying medical conditions that were not captured in acute settings.
We
Genomics England
understand this dataset
remains under COVID regulations and
will
remain
be used
only for COVID research going forward.
[1 paragraph unchanged]
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes,
e.g.
e.g.,
tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
[18 paragraphs unchanged]
Processing activities
All organisations party to this agreement must comply with the Data Sharing Framework Contract requirements, including those regarding the use (and purposes of that use) by "Personnel" (as defined within the Data Sharing Framework Contract i.e.: employees, agents and contractors of the Data Recipient who may have access to that data).
Genomics England provide NHS Digital with a cohort for linkage, and they receive data from NHS Digital on a monthly basis. Every month, Genomics England provide an updated cohort to NHS Digital who provide the historical data for the extra cohort members. The cohort is already flagged with NHS Digital so Genomics England will only receive the historical data for the extra cohort members each month.
Genomics England provide NHS Digital with a cohort for linkage and they receive data from NHS Digital on a monthly basis. Every month, Genomics England provide an updated cohort to NHS Digital who provide the historical data for the extra cohort members. The cohort is already flagged with NHS Digital so Genomics England will only receive the historical data for the extra cohort members each month.
The first stage of processing focuses on quality verification. This ensures that the data set is complete, accurate and complies with the NHS data dictionary or relevant specification. Participant identifiers in the dataset are verified against Genomics England's participant details and any updates required to identifiable data fields, e.g., dates of birth, are highlighted. Finally, the data set is reviewed against recent participant withdrawals so that any withdrawals notified after the data application was made can be removed from the data sets.
The first stage of processing focuses on quality verification. This ensures that the data set is complete, accurate and complies with the NHS data dictionary or relevant specification. Participant identifiers in the dataset are verified against Genomics England's participant details and any updates required to identifiable data fields, e.g. dates of birth, are highlighted. Finally, the data set is reviewed against recent participant withdrawals so that any withdrawals notified after the data application was made can be removed from the data sets.
Following this the data are de-identified, as all subsequent processing can be performed without direct identifiers. Genomics England has compiled lists of identifiable and sensitive fields for each data set in line with details provided by NHS Digital and following internal review of data sets. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to a secure Trusted Research Environment (TRE) hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data.
Following this the data are de-identified, as all subsequent processing can be performed without direct identifiers. Genomics England has compiled lists of identifiable and sensitive fields for each data set in line with details provided by NHS Digital and following internal review of data sets. De-identification is a key facet of the Genomics England resource. De-identified data are uploaded to a secure research environment hosted by Genomics England on a monthly basis, where they are linked to participant genomes and primary clinical data.
The second stage of processing involves the selection of a de-identified cohort of participants that fulfil a specific research request. Researchers are members of a Genomics England Clinical Interpretation Partnership (GeCIP) or the Discovery Forum. Research requests are assessed to ensure that they are included in the approved use purposes set out in the Genomics England Protocol and fall within the scope of the relevant GeCIP or the Discovery Forum. Researchers declare any data they wish to bring into the TRE and any tools they wish to use for analysis.
The second stage of processing involves the selection of a de-identified cohort of participants that fulfil a specific research request. Researchers are members of a Genomics England Clinical Interpretation Partnership (GeCIP) or the Discovery Forum. Research requests are assessed to ensure that they are included in the approved use purposes set out in the Genomics England Protocol, and fall within the scope of the relevant GeCIP or the Discovery Forum. Researchers declare any data they wish to bring into the research environment and any tools they wish to use for analysis.
The third stage of processing is the analysis of the de-identified data sets within the TRE. Researchers perform all the analysis and processing within the environment: they do not extract de-identified data. Results data are placed in a secure folder for anonymisation verification before extraction.
The third stage of processing is the analysis of the de-identified data sets within the research environment. Researchers perform all the analysis and processing within the environment: they do not extract de-identified data. Results data are placed in a secure folder for anonymisation verification before extraction.
[1 paragraph unchanged]
The Research Environment (TRE):
THE TRUSTED RESEARCH ENVIRONMENT (TRE)
All research analysis on the Genomics England dataset will only be carried
[5 words unchanged]
environment hosted within the Genomics England data centre - the Genomics England
Research Environment.
TRE.
Analytical tools and applications are available within the
Research Environment.
TREt.
No sequencing or clinical data are made available for download, users cannot copy or paste out of the
Research Environment,
TRE,
and there is limited internet access within it (i.e. whitelisted sites). Movement of files into and out of the
Research Environment
TRE
is governed via an 'Airlock' Policy.
Academic researcher access to the Research Environment:
Academic researchers access the TRE by applying to be a member of a Genomics England Clinical Interpretation Partnership (GeCIP) domain. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
Academic researchers access the Research Environment by applying to be a member of a Genomics England Clinical Interpretation Partnership (GeCIP) domain. GeCIP membership is open to any individual, student or member of staff, who is affiliated with a host institution which include the following:
[14 paragraphs unchanged]
Commercial researcher access to the
research environment:
TRE:
Genomics England operates a membership-based forum - the Discovery Forum - which is open to a range of companies world-wide and allows access to the
Research Environment.
TRE.
It provides a platform for collaboration between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Each Discovery Forum member signs a Data Access Agreement with Genomics England.
[74 words unchanged]
is in place, each research project undertaken by the Company within the
Research Environment
TRE
must receive prior ARC approval.
Discovery Forum members access the
Research Environment
TRE
in a similar manner to GeCIP Researchers: all research is carried out within the
Research Environment,
TRE,
and any movement of results out of the environment occurs only through the Airlock Process.
The Access Review Committee:
THE ACCESS REVIEW COMMITTEE (ARC)
The Access Review Committee (ARC)
ARC
provides an independent examination of requests for data access, with regards to
[47 words unchanged]
made up of participants and parents/carers involved in the 100,000 Genomes Project.
The Airlock Process:
THE AIRLOCK PROCESS
The Genomics England
Trusted
Research Environment has been developed with the intention that all data analysis is carried out within it and that the only data to leave it are
analytical
summary
results. An Airlock process has been established which enables material (data, files, tools
etc.)
etc)
to be moved in or out of the Research Environment in a
[5 words unchanged]
research and discovery, while maintaining control of security and access. Removal of
summary
results therefore requires an Airlock request.
[1 paragraph unchanged]
o
1.
All relevant details of the
files
summary results
to be transferred must be provided with every request.
o
2.
All
files
summary results
transferred
may
must
be checked by Genomics England to ensure compliance with the relevant policies. Users will be notified of any
files
summary results
rejected along with the reason for the rejection.
o
3.
All
files
summary results
transferred will be checked for viruses and malware and those failing this
[9 words unchanged]
the requestors to resolve such issues before re-submitting the file for transfer.
o Files
4. Summary results
requested for transfer are assessed using the following criteria:
• whether the request aligns with the
user's
users
ARC approval
in full
[6 paragraphs unchanged]
The Airlock process is governed by the Airlock
Policy (attached),
Policy,
which defines the process and governance of the Airlock process. A set
[14 words unchanged]
researcher (during preparation of analysis results) and the output checker (during output-checking).
Analysed results are inspected to ensure they cannot be used to disclose the identity of the
participant.
participants.
Checking of
statistical output
summary results
by the Airlock Review Team is governed by a
generalizable
set of principles that guide individual
decisions and ensure flexible evaluation of the Genomics England dataset.
decisions.
By using a principles-based approach where each case is assessed individually the security of the dataset is maintained by exporting only
'safe'
appropriate
data. Review of transfer requests resulting in public-sharing/publication of data will be
[7 words unchanged]
can only be used for the specific use detailed in the original
export.
request. The Research Environment contains pseudonymised longitudinal data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be used and what can be exported. Where an export contains pseudonymised longitudinal data, Genomics England will always apply the requirements placed on them as conditions of having access to the data.
The Research Environment contains External Data (for example, Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be used and what can be exported. Where an export contains External Data, Genomics England will always apply the requirements placed on them as conditions of having access to the data. In some cases, particularly concerning the export of individual-level data, these will be more conservative than those applied to 100,000 Genomes Data alone.
All Airlock requests go through a robust approval process and the Airlock Manager has formal delegated approval to approve requests where there is precedent from previous Airlock Review Committees. Where a precedent has been set and a clear set of principles and rules are in place for types of research, the Airlock Manager can approve the request. For more complicated requests or where no precedent has been set these will go to the Airlock Committee and be reviewed by the Airlock Review Team.
The Airlock Review Team is a delegation of the Genomics England Chief Scientist responsible for oversight of all airlock requests in accordance with the Airlock Policy and the
group's
groups
Terms of Reference. It comprises:
o Senior Information Risk Office (SIRO)
• Technical Lead
o Technical Lead
• User Community Representative
o User Community Representative
• Bioinformatics Director
o Bioinformatics Director
• Caldicott Guardian
o Caldicott Guardian
• Chief Scientist representative
o Chief Scientist
SUB-LICENCING
Sub-licencing:
Genomics England has developed the TRE to allow registered third parties to access pseudonymised versions of the data that it holds, for the purposes of approved research. The TRE contains External Data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be accessed. Genomics England will always apply the requirements placed on them as holders of External Data to users of the TRE as a condition of having access to the data. The data is NOT for onward sharing outside of the TRE.
Genomics England has developed the Research Environment to allow registered third parties to access pseudonymised versions of the data that it holds, for the purposes of approved research. The Research Environment contains External Data (for example Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts and data sharing agreements between Genomics England and other parties that dictate how the data may be accessed. Genomics England will always apply the requirements placed on them as holders of External Data to users of the Research Environment as a condition of having access to the data. The data is NOT for onward sharing outside of the Research Environment.
DATA CONTROLLERS
Data control:
[1 paragraph unchanged]
Genomics England in its provision of whole genome sequencing are applying to
[43 words unchanged]
GeCIP to access the secondary clinical data from NHS Digital within the
Research Environment.
TRE.
Genomics England provides NHS Digital with linking data in order to receive
[44 words unchanged]
provided. Genomics England determines the method of de-identification and storage within the
research environment
TRE
and secures this data for use by approved researchers only. Genomics England determines who these researchers are.
[2 paragraphs unchanged]
Access to deidentified data in the
research environment
TRE
which will include longitudinal data sets (HES etc) provided by NHS Digital. Access to the
Research Environment
TRE
only allowed under access agreement.
The individual researchers are Data Controllers when carrying out research within the
research environment.
TRE.
Data disseminated under this agreement for COVID-19 purposes will be restricted to the GEL Covid
research environment.
TRE.
Only COVID-19 research approved studies will be granted access to the data.
[45 words unchanged]
the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Data Processors:
DATA PROCESSORS
[6 paragraphs unchanged]
Genomics England have ceased using UKCloud (UKC) as a data processor after migrating all data to Amazon Web Services (AWS). The data in UKC
is planned to be
was
destroyed in March 2022. A detailed plan has been created which
will ensure
ensures
the deletion runbook
will be
was
followed, which is aligned to Genomics England standard operating procedures.
[1 paragraph unchanged]
Lifebit has been selected as platform partner to deliver the
Research Environment
TRE
after reviewing several proposals. The UK-based SME (Small Medium Enterprise) offered a
[55 words unchanged]
range of academic and biotech/pharma researchers with varying levels of technical competency.
Data Minimisation:
DATA MINIMISATION
This will be limited to the selected cohorts and additions and deletions will be updated regularly.
Genomics England's Trusted Research Environment (TRE) aligns to the current NHS guidance and is currently cited as an example of best practice. The detail to support this is below.
Cohort Size:
In the HDR UK TRE Principles and Best Practices paper from December 2021, the Genomics England model is explicitly described (page 17) and this document has a foreword authored by the Director of Data Policy, NHSX and Director of Tech Policy, NHSX. https://www.hdruk.ac.uk/news/new-principles-published-to-improve-public-confidence-in-access-and-use-of-data-for-health-research-through-trusted-research-environments
Conceptually, in terms of the 5 safes framework, the safes work together to protect the privacy of the individual. The TRE model makes 4 of the safes: safe-setting, safe-people, safe-projects and safe-outputs very strong, meaning that it’s possible to allow the criteria for safe-data to be relaxed (i.e., de-identification only) while maintaining the overall level of privacy protection. The TRE model effectively moves data minimisation to the output stage - ’safe outputs’, i.e., the airlock, where the airlock managers and the airlock committee do review requests to ensure summary data is minimised to only what is necessary to demonstrate externally a particular research result. The data released by the Airlock project is aggregated summary data. Given these measures, there has even been a discussion about data within TREs being regarded as "functionally anonymous” because of these safeguards, which would put them outside the constraints of GDPR.
The risk is further minimised as all the researchers are under contractual obligation not to re-identify individuals from the data. By allowing researchers access to all the data it allows "hypothesis free research" and means they may discover things they didn't suspect.
The full dataset is “adequate and relevant” for all researchers to access and the TRE framework provides the required limitation. Research in the RE is around discovering genome phenome relationships and given the complexity of the genome and how little we understand about it, it’s impossible to predict what phenotype terms may be involved, so researchers need to be able to explore all of them, rather than repeatedly requesting different datasets.
Each individual participant has agreed and is aware that their data is in the research environment and accessed by researchers.
Access to the data is necessary for this purpose to ensure the research can be carried out and the most value extracted from the data. We do have a wide range of controls outlined above to make sure the risk is minimised.
Both Article 89 and Article 25(1) of the UK GDPR support this approach.
A Data Protection Impact Assessment was conducted in November 2021 prior to the transfer of the data to the new environment and is being continually assessed and updated by the data protection team.
COHORT SIZE
[2 paragraphs unchanged]
Expected output
Access to the data will enable the research to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
Combining genomic sequence data from COVID participants together with their medical records has created a ground-breaking research resource. Researchers are currently studying this data to be able to deliver outputs that will have tangible benefits to healthcare and patients. Understanding the symptoms and impact of COVID on people with differing genetic make-ups is hoped to lead to successful treatment of the disease being investigated.
Specifically the short term (6 months) and medium term benefits and outcomes from this programme of research anticipated are;
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the TRE to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. Specific examples of journal outputs are as follows:
• Variants enable Polygenic Risk Score to predict greatest risk and avoid ITU
• A first update on mapping the human genetic architecture of COVID-19 (2022) Nature 608(7921):E1-E10. doi: 10.1038/s41586-022-04826-7
• Pre-morbid clinical conditions or biomarkers of risk and rapid NHS uptake to avoid ITU
• Mapping the human genetic architecture of COVID-19 (2021) Nature 600(7889):472-477. doi: 10.1038/s41586-021-03767-x
• Identify novel therapies or precision interventions for rapid national trials
• Whole-genome sequencing reveals host factors underlying critical COVID-19 (2022) Nature 607(7917):97-103. doi: 10.1038/s41586-022-04576-6
• Longitudinal life course sequel of COVID-19 for pandemic planning
• Genetic mechanisms of critical illness in COVID-19 (2021) Nature 591(7848):92-98. doi: 10.1038/s41586-020-03065-y
• Patient benefit:
o Providing improved clinical understanding of disease progression in COVID-19
o Correlation to disease progression and pre-morbid status
o Identification of susceptibility genes
o Develop a biomarker test(s) to predict an individual’s response to SARS-CoV-2 exposure, considering both COVID-19 severity and vulnerability to infection.
o Identify targets that can be used in to inform development of new treatments
• New scientific insights and discovery:
o with the consent of patients, creating a database of 35,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers.
o Correlation of host and viral genomic data
o Potential to provide improved testing for future pandemics
o Aide researchers to identify novel targets for vaccines and therapy
o Identification of highly penetrant rare variants in genes and pathways relating to viral susceptibility or immunodeficiency.
o Genome wide association studies (GWAS) using common variants to identify genes and pathways associated with viral response. These analyses will be aligned with other COVID-19 research consortia.
o Rare variant burden analysis to identify genes and pathways enriched in rare variants associated with viral response
• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. WGS could potentially provide the most accurate diagnostic test for COVID 19.
• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices.
• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics.
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England Research Environment where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw.
“This is a ground-breaking and far-reaching study which will harness the UK’s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
The Research environment also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public Health England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the Research Environment monthly in order to continue support for, and to further develop, this ground-breaking resource.
Outputs published to date are;
Publications:
• Pairo-Castineira E., Clohisey S., Klaric L., Bretherick A.D., Rawlik K., Pasko D., Walker S., Parkinson N., Fourman M.H., Russell C.D., Furniss J., Richmond A., Gountouna E., Wrobel N., Harrison D., Wang B., Wu Y., Meynert A., Griffiths F., Oosthuyzen W., Kousathanas A., Moutsianas L., Yang Z., Zhai R., Zheng C., Grimes G., Beale R., Millar J., Shih B., Keating S., Zechner M., Haley C., Porteous D.J., Hayward C., Yang J., Knight J., Summers C., Shankar-Hari M., Klenerman P., Turtle L., Ho A., Moore S.C., Hinds C., Horby P., Nichol A., Maslove D., Ling L., McAuley D., Montgomery H., Walsh T., Pereira A.C., Renieri A., GenOMICC I., Investigators, COVID- Human Genetics I., Investigators, BRACOVID I., Gen-COVID I., Shen X., Ponting C.P., Fawkes A., Tenesa A., Caulfield M., Scott R., Rowan K., Murphy L., Openshaw P., Semple M.G., Law A., Vitart V., Wilson J.F., Baillie J.K. Genetic mechanisms of critical illness in COVID-19. Nature. 2020;591(7848):92–98
• Kosmicki JA, Horowitz JE, Banerjee N, Lanche R, Marcketta A, Maxwell E, Bai X, Sun D, Backman JD, Sharma D, Kang HM, O'Dushlaine C, Yadav A, Mansfield AJ, Li AH, Watanabe K, Gurski L, McCarthy SE, Locke AE, Khalid S, O'Keeffe S, Mbatchou J, Chazara O, Huang Y, Kvikstad E, O'Neill A, Nioi P, Parker MM, Petrovski S, Runz H, Szustakowski JD, Wang Q, Wong E, Cordova-Palomera A, Smith EN, Szalma S, Zheng X, Esmaeeli S, Davis JW, Lai YP, Chen X, Justice AE, Leader JB, Mirshahi T, Carey DJ, Verma A, Sirugo G, Ritchie MD, Rader DJ, Povysil G, Goldstein DB, Kiryluk K, Pairo-Castineira E, Rawlik K, Pasko D, Walker S, Meynert A, Kousathanas A, Moutsianas L, Tenesa A, Caulfield M, Scott R, Wilson JF, Baillie JK, Butler-Laporte G, Nakanishi T, Lathrop M, Richards JB, Jones M, Balasubramanian S, Salerno W, Shuldiner AR, Marchini J, Overton JD, Habegger L, Cantor MN, Reid JG, Baras A, Abecasis GR, Ferreira MA. A catalog of associations between rare coding variants and COVID-19 outcomes. medRxiv [Preprint]. 2021 Feb 27:2020.10.28.20221804. doi: 10.1101/2020.10.28.20221804. PMID: 33655273; PMCID: PMC7924298.
• RECOVERY Collaborative Group. Tocilizumab in patients admitted to hospital with COVID-19 (RECOVERY): a randomised, controlled, open-label, platform trial. Lancet. 2021 May 1;397(10285):1637-1645. doi: 10.1016/S0140-6736(21)00676-0. PMID: 33933206; PMCID: PMC8084355.
Expected measurable benefits
Not stated in the previous version; added here.
Access to the data will enable the research to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
Specifically the short term (6 months) and medium term benefits and outcomes from this programme of research anticipated are;
• Variants enable Polygenic Risk Score to predict greatest risk and avoid ITU
• Pre-morbid clinical conditions or biomarkers of risk and rapid NHS uptake to avoid ITU
• Identify novel therapies or precision interventions for rapid national trials
• Longitudinal life course sequel of COVID-19 for pandemic planning
PATIENT BENEFIT
o Providing improved clinical understanding of disease progression in COVID-19
o Correlation to disease progression and pre-morbid status
o Identification of susceptibility genes
o Develop a biomarker test(s) to predict an individual͛s response to SARS-CoV-2 exposure, considering both COVID-19 severity and vulnerability to infection.
o Identify targets that can be used in to inform development of new treatments
• New scientific insights and discovery:
o with the consent of patients, creating a database of 35,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers.
o Correlation of host and viral genomic data
o Potential to provide improved testing for future pandemics
o Aide researchers to identify novel targets for vaccines and therapy
o Identification of highly penetrant rare variants in genes and pathways relating to viral susceptibility or immunodeficiency.
o Genome wide association studies (GWAS) using common variants to identify genes and pathways associated with viral response. These analyses will be aligned with other COVID-19 research consortia.
o Rare variant burden analysis to identify genes and pathways enriched in rare variants associated with viral response
• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. WGS could potentially provide the most accurate diagnostic test for COVID 19.
• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices.
• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics.
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Trusted Research Environment (TRE).
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England TRE where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw. This is a ground-breaking and far-reaching study which will harness the UK͛s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
The TRE also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public Health England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the TRE monthly in order to continue support for, and to further develop, this ground-breaking resource.
Benefits reported
[2 paragraphs unchanged]
Summary of publications and value of study updated
December 2021:
August 2022:
[1 paragraph unchanged]
Identified the following genes of interest:
May 2021:
• chromosome 12q24.13 in a gene cluster that encodes antiviral restriction enzyme activators (OAS1, OAS2 and OAS3)
• Identified 4 targets that can utilise and repurpose currently used medications for treatment and management of severe covid.
• chromosome 19p13.2 near the gene that encodes tyrosine kinase 2 (TYK2);
• Found evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19. This mechanism may also be amenable to targeted treatment with existing drugs.
• chromosome 19p13.3 within the gene that encodes dipeptidyl peptidase 9 (DPP9);
These results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
• chromosome 21q22.1 (rs2236757, P = 4.99 × 10-8) in the interferon receptor gene IFNAR2.
Mapping the human genetic architecture of COVID-19
• These targets can utilise and repurpose currently used medications for treatment and management of severe covid.
Dec 2021
• evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19.
• Contributed to the meta-analyses that consist of up to 49,562 patients with COVID-19 from 46 studies across 19 countries.
• Our results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
• Identified 13 genome-wide significant loci that are associated with SARS-CoV-2 infection or severe manifestations of COVID-19. Several of these loci correspond to previously documented associations to lung or autoimmune and inflammatory diseases.
• Both mechanisms may be amenable to targeted treatment with existing drugs.
• They represent potentially actionable mechanisms in response to infection.
In addition, two therapeutic drug trials have also commenced.
• There is evidence of a causal role for smoking and body-mass index for severe COVID-19 although not for type II diabetes.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy. Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
This working model of international collaboration underscores what is possible for future genetic discoveries in emerging pandemics, or indeed for any complex human disease.
May 2022
• Identified TMPRSS2 variant has a protective effect against severe COVID-19
• This is a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19
Whole genome sequencing of 7,491 GenOMICC participants admitted with severe COVID
Jul 2022
• Identified 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2)
• Found evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease.
• Found evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets.
The results show that comparisons between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Aug 2022
• Contributed to the meta-analyses bringing together 60 studies from 25 countries for 3 COVID-19 phenotypes.
• Between-study heterogeneity rather than differences across ancestries are a more likely explanation for the observed heterogeneity in the effect sizes across studies
• Identified that the variant rs35705950:G>T located in the promoter of MUC5B (11p15.5) is protective against hospitalization
• Identified that rs190509934:T>C, which is upstream of ACE2, is associated with decreased susceptibility risk. Recent results have shown that the rs190509934:T>C variant lowers ACE2 expression, which in turn confers protection against SARS-CoV-2 infection12.
The biological insights gained by this expansion of the COVID-19 Host Genetic Initiative showed that increasing sample size and diversity remain a fruitful activity to better understand the human genetic architecture of COVID-19.
The first genomes from severe volunteers came through in June 2020.
By June 2021, interim findings from the study were published. These findings have already helped doctors make better decisions when treating patients with COVID-19, resulting in better outcomes. In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy.
Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
Objective for processing
This Agreement is seeking approval to request continued supply of data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS Digital and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England Trusted Research Environment (TRE) to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and its outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel.
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By utilising the GenOMICC consortium’s prospective study design that leverages existing recruitment infrastructure in critical care, Genomics England will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
Genomics England Background: (For Context)
Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.
Data will be released into the Genomics England TRE where it will be linked with associated clinical data as well as additional data sources, which will include NHS Digital data requested through this agreement, COVID-19 testing feeds and viral genomics from Public Health England and data feeds from the Intensive Care National Audit Registry (ICNARC).
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable TRE for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new TRE (COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables effective COVID-19 research outcomes across a wide range of academic and Biotech/Pharma researchers with varying levels of technical competency. This offers a major upgrade to Genomics England’s current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling, and enable shared workspaces between Genomics England and Partners. COVID-RE must serve the immediate COVID-19 research effort and may also advance the transformation of Genomics England’s platform infrastructure, which is a key enabler to the research community.
There are two key value streams provided by the COVID-RE. Firstly ‘raw data to analytics ready data’ stream which must permit the collection of data from multiple locations from unstructured, through semi and fully structured forms and transforms them the appropriate data model, data store based on their ongoing use and availability. Secondly the ‘discovery to insight’ stream that supports researchers by providing the capability and the framework to support their User Journey from understanding the data available to them through to providing the necessary analytics and publishing tools.
The newly established COVID-RE will provide the following:
• Seamless interface with cloud storage and compute capabilities.
• A unified data platform containing datastores appropriate for all the required clinical and genomic data types.
• Data integration capability to deliver analytics-ready datasets into domain-specific data stores across batch and streaming integration patterns.
• Standards-based access services providing secure, fast, flexible, robust and auditable access to TRE data assets.
• Applications to support the research user journey from an exploration of data sets, through cohort building, analysis and publication - providing tools appropriate to a variety of user requirements (ways of working) and programming competency.
• Applications and workflows can be delivered natively within the platform, however COVID-RE must enable access to container-based applications and, via a set of hardened APIs, to a relevant external service (as long as security is maintained).
Details of the sublicense model via the Genomics England TRE are supplied within the processing activities section of this agreement.
Regarding genome data sharing with sublicensees - As an example, if an individual’s whole-exome or whole-genome sequence forms part of a database but all other records and identifiers have been fully deleted so that all that remains is sequence data, this data could easily be ‘individuated’ or singled out. However, it can be argued that it cannot be connected to a natural person without something further which can relate it to them.
In this perspective, the Global Alliance for Genomics and Health (GA4GH) has grouped together to provide unified strategies for addressing the major challenges of this data revolution. Genomics England are key members in this alliance and a recent publication (https://www.cell.com/cell-genomics/fulltext/S2666-979X(21)00036-7) presents the GA4GH suite of secure, interoperable technical standards and policy frameworks
• GA4GH and the PHG foundation both suggest that genetic, and particularly genomic information should not be viewed as inherently or directly identifying without some further link to or impact on an individual
• To that end, Genomics' TRE is a leading light in security and standards as per HDR UK (https://zenodo.org/record/5767586#.YcPE9hPP27M) and an author is an employee of Genomics England.
• Furthermore, Genomics' Airlock policy is clear and explicit about the level of data that can be extracted and does not permit row level/ individual data extraction
The additional clinical data is key for researchers to be able to understand and infer clinically relevant and actionable findings. The aim is to provide a high quality, diverse clinical dataset, detailing each participant’s journey and to understand pre-existing conditions and also the early behaviours in the disease course. Genomics England currently have an agreement to receive NHS Digital data for the extant 100,000 Genomes Project participants (DARS-NIC-12784) and have seen first-hand the depth and quality of the data and how it has aided researchers.
To that aim, the significant gap in the current data collection for the COVID-19 project can be addressed by the non-standard NHS Digital data feed.
The Secondary Use Services Admitted Patient Care Feed (SUS APC) would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
The associated COVID-19 datasets (including NHS111 and CV19 Testing Data in particular which will be requested under a future version of this agreement) will add to the detail in patient journey course, flagging for instance how participants were monitoring their symptoms and was there an associated poor prognosis.
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to approx. 36,000 participants.
The data would be available for analysis alongside the extant Genomics England data set of 100,000 Genomes Project participants and would be made available to approved researchers worldwide as per existing governance procedures. This 100,000 Genomes Project data set will act as an appropriate matched control group. Data on the 100,000 Genomes Cohort will be provided to NHS Digital for linkage - (as they will act as a control cohort), as well as new additions for the GenOMICC Study.
The below detail pertains to the background of the study and susceptibility to infection sets out why Genomics are investigating. The origin of the GenOMICC study was focused on smaller cohorts of participants. However, in collaboration with the COG-UK group, this is now expanded to whole genome sequence of approx. 36,000 affected individuals as set out above.
GenOMICC Study Background:
Susceptibility to infection is profoundly heritable (Sorensen et al. 1988). Patients who develop life-threatening illness following infection with usually innocuous pathogens, such as influenza (Miller et al. 2010), are genetically different from the rest of the population (Albright et al. 2008). Understanding the genetic mechanisms of susceptibility may yield new therapeutic targets (Baillie 2014) that can be used to make susceptible patients more similar to individuals who are resistant to, or tolerant of, specific pathogens.
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV [Human Immunodeficiency Virus] (Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions (e.g., small molecules to inhibit an enzyme or receptor that is dysfunctional in resistant individuals) would therefore be protective to the host in a similar way to antibiotics, with the advantage that it is conceptually more difficult for any one pathogen to evolve resistance to such a therapy.
A second, more challenging problem arises in patients who become critically ill following infection. The patterns of immune-mediated organ dysfunction, immunoparesis, and death are very similar in severe infections and sterile systemic injuries (such as burns, haemorrhage, pancreatitis and trauma). Ultimately, death is a consequence of the host response to injury (Angus and Poll 2013), through final common pathways of organ failure that are clinically and biochemically evident, and unrelated to the original precipitant.
Broadly, the severity of critical illness follows directly from the severity and duration of the initial insult. In bacterial sepsis, early antibiotics are the mainstay of therapy; in influenza, early antivirals; in haemorrhage, early resuscitation; in trauma, urgent action to prevent secondary injury. There are no therapies with which to modulate the host response to systemic injury.
There is a lack of direct evidence of heritability for outcomes of critical illness, due in part to difficulties in defining and quantifying the heterogeneous multi-organ dysfunction syndrome (MODS), and in part due to the rapid pace of change in critical care medicine, making it impossible to tackle this question in long term outcome studies. However, clinical and biological evidence support the hypothesis that the pathogenesis of MODS is immune in origin (Angus and Poll 2013). Hence, predictions can be made from the extensive knowledge of other immune conditions. Whether or not MODS is considered to be an autoimmune or infectious condition is moot: these conditions share a great deal of similarity in genetic predispositions, cell types and mechanisms of pathogenesis. It is therefore very likely that propensity to survive MODS has a heritable component, and there is some direct evidence in support of this hypothesis (Rautanen et al. 2015). If this is the case, then the identity of the specific variants that contribute to outcome could potentially be utilised to design therapies to promote survival after the onset of MODS.
This study aims to identify genetic predisposition to specific syndromes of critical illness. Specifically, susceptibility to life-threatening infections caused by an identified pathogen, and susceptibility to death following the onset of organ failure due to sepsis or sterile injury. In order to maximise the probability of identifying host genetic loci associated with susceptibility, Genomics England will restrict some analyses to younger individuals in good general health and lacking in known predisposing factors.
The same principle was used to determine an upper age limit for inclusion for some analyses. With advancing age, there is an increase in undiagnosed co-morbidity, frailty, and susceptibility to serious complications of infection or critical injury. There is therefore an increase in the probability of susceptibility to, and mortality from, critical illness that is consequent upon non-genetic factors.
Participation Group:
Patients will be identified and recruited in hospital during acute illness. Potential participants will be identified through hospital workers upon presentation at recruiting sites. The disease processes under study have a high mortality, so it is desirable to recruit patients as early as possible in the disease process.
Participants or an appropriate parent/guardian/consultee will be approached by staff trained in consent procedures that protect the rights of the patient and adhere to the ethical principles within the Declaration of Helsinki. Staff will explain the details of the study to the participant or parent/guardian/consultee and allow them time to discuss and ask questions. The staff will review the informed consent form with the person giving consent (or assent) and endeavour to ensure understanding of the contents, including study procedures, risks, benefits, and the right to withdraw. Participants who agree to participate (or their parent/guardian or consultee who declares their wishes to do so) will be asked to sign and date an informed consent form.
In view of the importance of early sampling, participants or their parent/guardian/consultee will be permitted to consent and begin to participate in the study immediately if they wish to do so. Those who prefer more time to consider participation will be approached again after an agreed time, normally one day, to discuss further.
Patients who meet the inclusion/exclusion criteria and who have given informed consent to participate directly, or have been consented by a parent/guardian or whose wishes have been declared by a consultee, will be enrolled to the study.
Samples and data will be collected according to available resources and the weight of the patient will be measured for children under 12 in order to prevent excessive volume sampling. Samples required for medical management will at all times have priority over samples taken for research tests. Aliquots or samples for research purposes should never compromise the quality or quantity of samples required for medical management. Wherever practical, taking research samples should be timed to coincide with clinical sampling. The research team will be responsible for sharing the sampling protocol with health care workers supporting patient management in order to minimise disruption to routine care and avoid unnecessary procedures.
Consent will be sought from patients who survive critical illness and regain capacity to give consent. At each of the follow-up sessions, the investigator gathering data will determine whether the patient has regained capacity. In the event that a patient continues to be incapacitated beyond the follow-up period, the local investigator will plan a subsequent capacity check at a specific date, after an interval to be determined by the nature of the incapacity. The planned dates of capacity checks on incapacitated survivors will be stored locally in the site file, together with a record of the outcome of each check.
Patients who decline to participate at this stage will be removed from the study. Where the patient cannot be recontacted despite best endeavours, they will remain in the study.
Withdrawal:
Participants are freely able to decline participation in this study or to withdraw from participation at any point without suffering any implied or explicit disadvantage. All patients will be treated according to standard practice regardless of whether they participate.
The following options of withdrawal will be made available to participants:
1. Partial withdrawal. Data WILL continue to be updated and used for research, but no further contact will be made with the participant
2. Full withdrawal.
• no further contact will be made with the participant;
• data will not be updated from health records;
• data will not be removed from research that is underway or has already been done, and an audit record will be maintained to confirm participation.
Consent version and life course follow-up
There are a series of consent versions as documented below:
• V1.08 (March 2020) - there are circa 1500 participants on this consent, allows for COVID research but not for longitudinal follow-up. Genomics are exploring reconsenting and the Confidentiality Advisory Group (CAG) as options for this cohort. Currently, these individuals will NOT be submitted to NHS Digital
• V2.1 (April 2020) - Included retention of NHS number and linkage to longitudinal follow -up, NHS Digital are named on the Participant Information Sheet (PIS). These individuals will be submitted to NHS Digital for data linkage
• V2.4 (July 2020) - In addition to already stated in v2.1, this covers community recruitment of participants and movement of identifiers. Sharing data with NHS Digital is again stated. These individuals will be submitted to NHS Digital for data linkage
PIS and Consent Version 2.1 and 2.4 are identical in their reference to accessing data from secondary data suppliers and include retention and collection of identifiable information and both name NHS Digital as a source of linkage. All individuals consented are collected with information including version of consent recruited under, thus genomics England can distinguish individuals based on consent. The most up to date protocol is listed on the website (https://genomicc.org/protocol/)
Datasets requested from NHS Digital:
This agreement permits access to the data sets requested, in line with COVID-19 related purpose. For participants recruited under the consultee process, the restrictions set out in Reg 3( 1) COPI will also apply.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data for Pandemic Planning and Research (GDPPR) extracts to assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. Genomics England understand this dataset will be used only for COVID research going forward.
Hospital Episode Statistics (HES): Outpatients (OP), Admitted Patient Care (APC), Critical Care (CC) and Accident & Emergency (AE)/ Emergency Care Data Set (ECDS). These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes, e.g., tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Secondary Uses Service datasets. The minimal latency in availability of these datasets is highly desirable for the research objectives set out in this project.
Mental Health Data sets: The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
Cancer Registration Data sets: To see the incidence of cancer within the cohort
Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
COVID datasets: These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
Assessment of the datasets has been undertaken and NHS Digital are satisfied that they are necessary for the COVID-19 work being undertaken. Genomics have confirmed that all research which is approved from the GenoMICC study using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Genomics England Industry access:
Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help turn research findings into treatments, diagnostics and benefits for patients as soon as possible.
As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations to access the COVID data, however they are charged based on storage and compute, such as running their own bioinformatics pipeline. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected.
The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a 'critical friend' and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's landmark data set.
The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
Additionally, as Genomics England will be processing health data - a special category of personal data, they will also be processing data under Article 9 (2)(j) as processing is necessary for archiving purposes in the public interest. Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
Expected output
Combining genomic sequence data from COVID participants together with their medical records has created a ground-breaking research resource. Researchers are currently studying this data to be able to deliver outputs that will have tangible benefits to healthcare and patients. Understanding the symptoms and impact of COVID on people with differing genetic make-ups is hoped to lead to successful treatment of the disease being investigated.
Specific outputs for Genomics England are to continue to release updated genomic and clinical data into the TRE to support this ongoing research. Each research group who accesses this data will have their own defined output strategy and expectations which they will aim to deliver against. Specific examples of journal outputs are as follows:
• A first update on mapping the human genetic architecture of COVID-19 (2022) Nature 608(7921):E1-E10. doi: 10.1038/s41586-022-04826-7
• Mapping the human genetic architecture of COVID-19 (2021) Nature 600(7889):472-477. doi: 10.1038/s41586-021-03767-x
• Whole-genome sequencing reveals host factors underlying critical COVID-19 (2022) Nature 607(7917):97-103. doi: 10.1038/s41586-022-04576-6
• Genetic mechanisms of critical illness in COVID-19 (2021) Nature 591(7848):92-98. doi: 10.1038/s41586-020-03065-y
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated August 2022:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
May 2021:
• Identified 4 targets that can utilise and repurpose currently used medications for treatment and management of severe covid.
• Found evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19. This mechanism may also be amenable to targeted treatment with existing drugs.
These results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
Mapping the human genetic architecture of COVID-19
Dec 2021
• Contributed to the meta-analyses that consist of up to 49,562 patients with COVID-19 from 46 studies across 19 countries.
• Identified 13 genome-wide significant loci that are associated with SARS-CoV-2 infection or severe manifestations of COVID-19. Several of these loci correspond to previously documented associations to lung or autoimmune and inflammatory diseases.
• They represent potentially actionable mechanisms in response to infection.
• There is evidence of a causal role for smoking and body-mass index for severe COVID-19 although not for type II diabetes.
This working model of international collaboration underscores what is possible for future genetic discoveries in emerging pandemics, or indeed for any complex human disease.
May 2022
• Identified TMPRSS2 variant has a protective effect against severe COVID-19
• This is a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19
Whole genome sequencing of 7,491 GenOMICC participants admitted with severe COVID
Jul 2022
• Identified 16 new independent associations, including variants within genes that are involved in interferon signalling (IL10RB and PLSCR1), leucocyte differentiation (BCL11A) and blood-type antigen secretor status (FUT2)
• Found evidence that implicates multiple genes-including reduced expression of a membrane flippase (ATP11A), and increased expression of a mucin (MUC1)-in critical disease.
• Found evidence in support of causal roles for myeloid cell adhesion molecules (SELE, ICAM5 and CD209) and the coagulation factor F8, all of which are potentially druggable targets.
The results show that comparisons between cases of critical illness and population controls is highly efficient for the detection of therapeutically relevant mechanisms of disease.
Aug 2022
• Contributed to the meta-analyses bringing together 60 studies from 25 countries for 3 COVID-19 phenotypes.
• Between-study heterogeneity rather than differences across ancestries are a more likely explanation for the observed heterogeneity in the effect sizes across studies
• Identified that the variant rs35705950:G>T located in the promoter of MUC5B (11p15.5) is protective against hospitalization
• Identified that rs190509934:T>C, which is upstream of ACE2, is associated with decreased susceptibility risk. Recent results have shown that the rs190509934:T>C variant lowers ACE2 expression, which in turn confers protection against SARS-CoV-2 infection12.
The biological insights gained by this expansion of the COVID-19 Host Genetic Initiative showed that increasing sample size and diversity remain a fruitful activity to better understand the human genetic architecture of COVID-19.
The first genomes from severe volunteers came through in June 2020.
By June 2021, interim findings from the study were published. These findings have already helped doctors make better decisions when treating patients with COVID-19, resulting in better outcomes. In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy.
Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
DARS-NIC-374190-D0N1M-v4.1 13 April 2022 to 30 September 2022
- Title
- R26 - GENOMICS ENGLAND: GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 21
- Files released
- 151
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); COVID-19 Vaccination Status; Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v3.2
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2022-04-13 | |
| Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset: legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR): legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| COVID-19 Hospitalization in England Surveillance System: legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| COVID-19 Hospitalization in England Surveillance System: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| COVID-19 SGSS First Positives (Second Generation Surveillance System): legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| COVID-19 SGSS First Positives (Second Generation Surveillance System): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Cancer Registration Data: legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Cancer Registration Data: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Civil Registrations of Death: legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Civil Registrations of Death: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Community Services Data Set (CSDS): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Demographics: legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Demographics: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Diagnostic Imaging Data Set (DID): legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Diagnostic Imaging Data Set (DID): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Emergency Care Data Set (ECDS): legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Emergency Care Data Set (ECDS): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| HES-ID to MPS-ID HES Accident and Emergency: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| HES-ID to MPS-ID HES Admitted Patient Care: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| HES-ID to MPS-ID HES Outpatients: common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Hospital Episode Statistics Accident and Emergency (HES A and E): legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Hospital Episode Statistics Accident and Emergency (HES A and E): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Hospital Episode Statistics Admitted Patient Care (HES APC): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Hospital Episode Statistics Critical Care (HES Critical Care): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Hospital Episode Statistics Outpatients (HES OP): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Mental Health Services Data Set (MHSDS): legal basis | Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c) | |
| Mental Health Services Data Set (MHSDS): common law duty of confidentiality | Consent (Reasonable Expectation) | |
| Secondary Uses Service Payment By Results Spells: common law duty of confidentiality | Consent (Reasonable Expectation) |
Datasets: + COVID-19 Vaccination Status
Objective for processing
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
December 2021 Amendment- Version 4.1 of the agreement seeks to change the way in which the common law duty of confidentiality for the cohort is being addressed moving from COPI to consent and in addition requests access to vaccination data.
_____________________________________________________________________________________
This agreement is seeking approval to request continued supply of data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
This agreement is seeking approval to request data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
[3 paragraphs unchanged]
3. To collect longitudinal life course datasets from primary care, hospital episodes,
[38 words unchanged]
upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
and UK Biobank data sets (500,000 people) with WGS where more cases may be identified, including those with milder disease and unaffected people.
This mention of linkage of the NHS Digital data to UK Biobank data is purely aspirational at present. This has not been approved under previous iterations of the agreement, and is not covered under v1.2 of this agreement. Genomics England can confirm that if such linkage were to occur - then appropriate documentation (including transparency, patient, consent and ethics materials) would be provided to NHS Digital and a subsequent amended application put in place.
[21 paragraphs unchanged]
Details of the
sublicence
sublicense
model via the Genomics England Research Environment are supplied within the processing activities section of this agreement.
Regarding genome data sharing with sublicensees - As an example, if an individual’s whole-exome or whole-genome sequence forms part of a database but all other records and identifiers have been fully deleted so that all that remains is sequence data, this data could easily be ‘individuated’ or singled out. However, it can be argued that it cannot be connected to a natural person without something further which can relate it to them.
In this perspective, the Global Alliance for Genomics and Health (GA4GH) has grouped together to provide unified strategies for addressing the major challenges of this data revolution. Genomics England are key members in this alliance and a recent publication (https://www.cell.com/cell-genomics/fulltext/S2666-979X(21)00036-7) presents the GA4GH suite of secure, interoperable technical standards and policy frameworks
• GA4GH and the PHG foundation both suggest that genetic, and particularly genomic information should not be viewed as inherently or directly identifying without some further link to or impact on an individual
• To that end, our research environment is a leading light in security and standards as per HDR UK (https://zenodo.org/record/5767586#.YcPE9hPP27M) and an author is an employee of Genomics England.
• Furthermore, our Airlock policy is clear and explicit about the level of data that can be extracted and does not permit row level/ individual data extraction
[4 paragraphs unchanged]
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to
approx
approx.
36,000 participants.
[26 paragraphs unchanged]
Consent version and
lifecourse
life course
follow-up
The first COVID positive patient was recruited to the GenOMICC study in March 2020. All patients (c.1500) currently recruited to the GenOMICC study are on consent and Protocol version 1.08 (which allows for COVID research but not longitudinal life course follow up). This Protocol v2.1 (submitted as part of this DARS request) was Research Ethics Committe (REC) approved on 23rd March 2020, IRAS [Integrated Research Application System] IDs are: 269326 & 189676 (https://www.hra.nhs.uk/covid-19-research/approved-covid-19-research/269326/). Genomics England will attempt to reconsent all patients recruited under version 1.08 onto the newly amended materials (version 2.5 of the patient information sheet – submitted as part of this DARS application) which allows for longitudinal lifecourse follow-up. For any patients on the v1.08 Protocol for which Genomics England are not able to seek consent, Genomics England will not be requesting their data for linkage. For patients prospectively recruited onto the v2.1 Protocol and consent materials, Genomics England will be requesting their data for linkage and follow-up .
There are a series of consent versions as documented below:
• V1.08 (March 2020) - there are circa 1500 participants on this consent , allows for COVID research but not for longitudinal follow-up. Genomics are exploring reconsenting and CAG as options for this cohort. Currently, these individuals will NOT be submitted to NHS Digital
• V2.1 (April 2020) - Included retention of NHS number and linkage to longitudinal follow -up, NHS Digital are named on the PIS. These individuals will be submitted to NHS Digital for data linkage
• V2.4 (July 2020) - In addition to already stated in v2.1, this covers community recruitment of participants and movement of identifiers. Sharing data with NHS Digital is again stated. These individuals will be submitted to NHS Digital for data linkage
PIS and Consent Version 2.1 and 2.4 are identical in their reference to accessing data from secondary data suppliers and include retention and collection of identifiable information and both name NHS Digital as a source of linkage. All individuals consented are collected with information including version of consent recruited under, thus genomics England can distinguish individuals based on consent. The most up to date protocol is listed on the website (https://genomicc.org/protocol/)
[1 paragraph unchanged]
Consideration has been be given
The current application
to
assess and ensure that
access to the data sets requested are
inline
in line
with COVID-19 related purpose, in terms of the restrictions set out in Reg 3(1)
COPI, below
COPI. However, based on the now REC approved protocol, we
are
details of what each of the data sets will at a high level provide insight to.
looking to move this application to Consent basis.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data
[22 words unchanged]
comorbidities and underlying medical conditions that were not captured in acute settings.
We understand this dataset remains under COVID regulations and will remain only for COVID research going forward.
[6 paragraphs unchanged]
COVID datasets: These datasets, including
vaccination,
SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus.
These will remain only for COVID based research only.
[6 paragraphs unchanged]
The lawful basis for the release and use of the confidential data being shared under this version of the agreement is Regulation 3(4) of the National Health Service (Control of Patient Information Regulations) 2002 (COPI) to require NHS Digital to share confidential patient information with organisations entitled to process this under COPI for COVID-19 purposes. The application of this has been based on the information provided in the Whole genome sequencing of patients severely affected by COVID-19 funding proposal from The GenOMICC - COVID Genomics UK (CoG-UK) partnership which was supported by the CMO of England and the CFO of the DHSC.
[7 paragraphs unchanged]
Processing activities
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
[24 paragraphs unchanged]
Following approval, GeCIP researchers must sign a specific agreement
('GeCIP rules' which is attached)
covering their
behavior
behaviour
and working practice within the data infrastructure. Data access will not then be granted until a researcher has successfully passed mandatory information governance training.
[47 paragraphs unchanged]
Genomics England have ceased using UKCloud (UKC) as a data processor after migrating all data to Amazon Web Services (AWS). The data in UKC is planned to be destroyed in March 2022. A detailed plan has been created which will ensure the deletion runbook will be followed, which is aligned to Genomics England standard operating procedures.
[7 paragraphs unchanged]
Expected output
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
Access to the data will enable the research to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
_____________________________________________________________________________________
Specifically the short term (6 months) and medium term benefits and outcomes from this programme of research anticipated are;
Genomics England completed sequencing 100,000 genomes at the end of 2018 (https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000- genomes-sequenced/). During 2018, the Genomics England Research Environment was established to allow research access to de-identified genomic and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with now over 3000 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
• Variants enable Polygenic Risk Score to predict greatest risk and avoid ITU
• Pre-morbid clinical conditions or biomarkers of risk and rapid NHS uptake to avoid ITU
• Identify novel therapies or precision interventions for rapid national trials
• Longitudinal life course sequel of COVID-19 for pandemic planning
• Patient benefit:
o Providing improved clinical understanding of disease progression in COVID-19
o Correlation to disease progression and pre-morbid status
o Identification of susceptibility genes
o Develop a biomarker test(s) to predict an individual’s response to SARS-CoV-2 exposure, considering both COVID-19 severity and vulnerability to infection.
o Identify targets that can be used in to inform development of new treatments
• New scientific insights and discovery:
o with the consent of patients, creating a database of 35,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers.
o Correlation of host and viral genomic data
o Potential to provide improved testing for future pandemics
o Aide researchers to identify novel targets for vaccines and therapy
o Identification of highly penetrant rare variants in genes and pathways relating to viral susceptibility or immunodeficiency.
o Genome wide association studies (GWAS) using common variants to identify genes and pathways associated with viral response. These analyses will be aligned with other COVID-19 research consortia.
o Rare variant burden analysis to identify genes and pathways enriched in rare variants associated with viral response
• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. WGS could potentially provide the most accurate diagnostic test for COVID 19.
• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices.
• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics.
[6 paragraphs unchanged]
As of March 2020, the Genomics England Research Environment contained 107,694 genomes, of which, 33,461 were cancer and 74,233 were rare diseases.
[2 paragraphs unchanged]
Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants and GenOMICC participants into the Research Environment on the dates shown above.
Outputs published to date are;
Publications:
• Pairo-Castineira E., Clohisey S., Klaric L., Bretherick A.D., Rawlik K., Pasko D., Walker S., Parkinson N., Fourman M.H., Russell C.D., Furniss J., Richmond A., Gountouna E., Wrobel N., Harrison D., Wang B., Wu Y., Meynert A., Griffiths F., Oosthuyzen W., Kousathanas A., Moutsianas L., Yang Z., Zhai R., Zheng C., Grimes G., Beale R., Millar J., Shih B., Keating S., Zechner M., Haley C., Porteous D.J., Hayward C., Yang J., Knight J., Summers C., Shankar-Hari M., Klenerman P., Turtle L., Ho A., Moore S.C., Hinds C., Horby P., Nichol A., Maslove D., Ling L., McAuley D., Montgomery H., Walsh T., Pereira A.C., Renieri A., GenOMICC I., Investigators, COVID- Human Genetics I., Investigators, BRACOVID I., Gen-COVID I., Shen X., Ponting C.P., Fawkes A., Tenesa A., Caulfield M., Scott R., Rowan K., Murphy L., Openshaw P., Semple M.G., Law A., Vitart V., Wilson J.F., Baillie J.K. Genetic mechanisms of critical illness in COVID-19. Nature. 2020;591(7848):92–98
• Kosmicki JA, Horowitz JE, Banerjee N, Lanche R, Marcketta A, Maxwell E, Bai X, Sun D, Backman JD, Sharma D, Kang HM, O'Dushlaine C, Yadav A, Mansfield AJ, Li AH, Watanabe K, Gurski L, McCarthy SE, Locke AE, Khalid S, O'Keeffe S, Mbatchou J, Chazara O, Huang Y, Kvikstad E, O'Neill A, Nioi P, Parker MM, Petrovski S, Runz H, Szustakowski JD, Wang Q, Wong E, Cordova-Palomera A, Smith EN, Szalma S, Zheng X, Esmaeeli S, Davis JW, Lai YP, Chen X, Justice AE, Leader JB, Mirshahi T, Carey DJ, Verma A, Sirugo G, Ritchie MD, Rader DJ, Povysil G, Goldstein DB, Kiryluk K, Pairo-Castineira E, Rawlik K, Pasko D, Walker S, Meynert A, Kousathanas A, Moutsianas L, Tenesa A, Caulfield M, Scott R, Wilson JF, Baillie JK, Butler-Laporte G, Nakanishi T, Lathrop M, Richards JB, Jones M, Balasubramanian S, Salerno W, Shuldiner AR, Marchini J, Overton JD, Habegger L, Cantor MN, Reid JG, Baras A, Abecasis GR, Ferreira MA. A catalog of associations between rare coding variants and COVID-19 outcomes. medRxiv [Preprint]. 2021 Feb 27:2020.10.28.20221804. doi: 10.1101/2020.10.28.20221804. PMID: 33655273; PMCID: PMC7924298.
• RECOVERY Collaborative Group. Tocilizumab in patients admitted to hospital with COVID-19 (RECOVERY): a randomised, controlled, open-label, platform trial. Lancet. 2021 May 1;397(10285):1637-1645. doi: 10.1016/S0140-6736(21)00676-0. PMID: 33933206; PMCID: PMC8084355.
Expected measurable benefits
Stated in the previous version and removed here.
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
Access to the data will enable the research to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
Specifically the short term (6 months) and medium term benefits and outcomes from this programme of research anticipated are;
• Variants enable Polygenic Risk Score to predict greatest risk and avoid ITU
• Pre-morbid clinical conditions or biomarkers of risk and rapid NHS uptake to avoid ITU
• Identify novel therapies or precision interventions for rapid national trials
• Longitudinal life course sequel of COVID-19 for pandemic planning
• Patient benefit:
o Providing improved clinical understanding of disease progression in COVID-19
o Correlation to disease progression and pre-morbid status
o Identification of susceptibility genes
o Develop a biomarker test(s) to predict an individual’s response to SARS-CoV-2 exposure, considering both COVID-19 severity and vulnerability to infection.
o Identify targets that can be used in to inform development of new treatments
• New scientific insights and discovery:
o with the consent of patients, creating a database of 35,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers.
o Correlation of host and viral genomic data
o Potential to provide improved testing for future pandemics
o Aide researchers to identify novel targets for vaccines and therapy
o Identification of highly penetrant rare variants in genes and pathways relating to viral susceptibility or immunodeficiency.
o Genome wide association studies (GWAS) using common variants to identify genes and pathways associated with viral response. These analyses will be aligned with other COVID-19 research consortia.
o Rare variant burden analysis to identify genes and pathways enriched in rare variants associated with viral response
• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. WGS could potentially provide the most accurate diagnostic test for COVID 19.
• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices.
• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics.
Yielded Benefits:
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
The GenOMICC study is very much in its infancy and whole genome sequencing has only begun in the last month. It is thus too early to demonstrate any significant outcomes. These outcomes will only gain power and relevance with more prospectively recruited patients and a breadth and depth of clinical data. By providing this to researchers, Genomics England can provide them the necessary tools to explore the genomic data.
This proposal could enable Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
Benefits reported
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
[1 paragraph unchanged]
The GenOMICC study is very much in its infancy and whole genome sequencing has only begun in the last month. It is thus too early to demonstrate any significant outcomes. These outcomes will only gain power and relevance with more prospectively recruited patients and a breadth and depth of clinical data. By providing this to researchers, Genomics England can provide them the necessary tools to explore the genomic data.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated December 2021:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
Identified the following genes of interest:
• chromosome 12q24.13 in a gene cluster that encodes antiviral restriction enzyme activators (OAS1, OAS2 and OAS3)
• chromosome 19p13.2 near the gene that encodes tyrosine kinase 2 (TYK2);
• chromosome 19p13.3 within the gene that encodes dipeptidyl peptidase 9 (DPP9);
• chromosome 21q22.1 (rs2236757, P = 4.99 × 10-8) in the interferon receptor gene IFNAR2.
• These targets can utilise and repurpose currently used medications for treatment and management of severe covid.
• evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19.
• Our results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
• Both mechanisms may be amenable to targeted treatment with existing drugs.
In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy. Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
Objective for processing
December 2021 Amendment- Version 4.1 of the agreement seeks to change the way in which the common law duty of confidentiality for the cohort is being addressed moving from COPI to consent and in addition requests access to vaccination data.
This agreement is seeking approval to request continued supply of data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS Digital and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people)
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England Research Environment to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and it’s outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel.
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By undertaking a prospective study design that leverages existing recruitment infrastructure in critical care, together with Genomics England, NHS England, Devolved Nations and Public Health England (PHE) infrastructure this research will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
Genomics England Background: (For Context)
Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.
Data will be released into the Genomics England Research Environment where it will be linked with associated clinical data as well as additional data sources, which will include NHS Digital data requested through this agreement, COVID-19 testing feeds and viral genomics from Public Health England and data feeds from the Intensive Care National Audit Registry (ICNARC).
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable Research Environment for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new Research Environment (COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables effective COVID-19 research outcomes across a wide range of academic and Biotech/Pharma researchers with varying levels of technical competency. This offers a major upgrade to our current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling, and enable shared workspaces between Genomics England and Partners. COVID-RE must serve the immediate COVID-19 research effort and may also advance the transformation of Genomics England’s platform infrastructure, which is a key enabler to the research community.
There are two key value streams provided by the COVID-RE. Firstly ‘raw data to analytics ready data’ stream which must permit the collection of data from multiple locations from unstructured, through semi and fully structured forms and transforms them the appropriate data model, data store based on their ongoing use and availability. Secondly the ‘discovery to insight’ stream that supports researchers by providing the capability and the framework to support their User Journey from understanding the data available to them through to providing the necessary analytics and publishing tools.
The newly established COVID-RE will provide the following:
• Seamless interface with cloud storage and compute capabilities.
• A unified data platform containing datastores appropriate for all the required clinical and genomic data types.
• Data integration capability to deliver analytics-ready datasets into domain-specific data stores across batch and streaming integration patterns.
• Standards-based access services providing secure, fast, flexible, robust and auditable access to TRE data assets.
• Applications to support the research user journey from an exploration of data sets, through cohort building, analysis and publication - providing tools appropriate to a variety of user requirements (ways of working) and programming competency.
• Applications and workflows can be delivered natively within the platform, however COVID-RE must enable access to container-based applications and, via a set of hardened APIs, to a relevant external service (as long as security is maintained).
Details of the sublicense model via the Genomics England Research Environment are supplied within the processing activities section of this agreement.
Regarding genome data sharing with sublicensees - As an example, if an individual’s whole-exome or whole-genome sequence forms part of a database but all other records and identifiers have been fully deleted so that all that remains is sequence data, this data could easily be ‘individuated’ or singled out. However, it can be argued that it cannot be connected to a natural person without something further which can relate it to them.
In this perspective, the Global Alliance for Genomics and Health (GA4GH) has grouped together to provide unified strategies for addressing the major challenges of this data revolution. Genomics England are key members in this alliance and a recent publication (https://www.cell.com/cell-genomics/fulltext/S2666-979X(21)00036-7) presents the GA4GH suite of secure, interoperable technical standards and policy frameworks
• GA4GH and the PHG foundation both suggest that genetic, and particularly genomic information should not be viewed as inherently or directly identifying without some further link to or impact on an individual
• To that end, our research environment is a leading light in security and standards as per HDR UK (https://zenodo.org/record/5767586#.YcPE9hPP27M) and an author is an employee of Genomics England.
• Furthermore, our Airlock policy is clear and explicit about the level of data that can be extracted and does not permit row level/ individual data extraction
The additional clinical data is key for researchers to be able to understand and infer clinically relevant and actionable findings. The aim is to provide a high quality, diverse clinical dataset, detailing each participant’s journey and to understand pre-existing conditions and also the early behaviours in the disease course. Genomics England currently have an agreement to receive NHS Digital data for the extant 100,000 Genomes Project participants (DARS-NIC-12784) and have seen first-hand the depth and quality of the data and how it has aided researchers.
To that aim, the significant gap in the current data collection for the COVID-19 project can be addressed by the non-standard NHS Digital data feed.
The Secondary Use Services Admitted Patient Care Feed (SUS APC) would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
The associated COVID-19 data-sets (including NHS111 and CV19 Testing Data in particular which will be requested under a future version of this agreement) will add to the detail in patient journey course, flagging for instance how participants were monitoring their symptoms and was there an associated poor prognosis.
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to approx. 36,000 participants.
The data would be available for analysis alongside the extant Genomics England data set of 100,000 Genomes Project participants and would be made available to approved researchers worldwide as per existing governance procedures. This 100,000 Genomes Project data set will act as an appropriate matched control group. Data on the 100,000 Genomes Cohort will be provided to NHS Digital for linkage - (as they will act as a control cohort), as well as new additions for the GenOMICC Study.
The below detail pertains to the background of the study and susceptibility to infection sets out why Genomics are investigating. The origin of the GenOMICC study was focused on smaller cohorts of participants. However, in collaboration with the COG-UK group, this is now expanded to whole genome sequence of approx. 36,000 affected individuals as set out above.
GenOMICC Study Background:
Susceptibility to infection is profoundly heritable (Sorensen et al. 1988). Patients who develop life-threatening illness following infection with usually innocuous pathogens, such as influenza (Miller et al. 2010), are genetically different from the rest of the population (Albright et al. 2008). Understanding the genetic mechanisms of susceptibility may yield new therapeutic targets (Baillie 2014) that can be used to make susceptible patients more similar to individuals who are resistant to, or tolerant of, specific pathogens.
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV [Human Immunodeficiency Virus] (Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions (e.g. small molecules to inhibit an enzyme or receptor that is dysfunctional in resistant individuals) would therefore be protective to the host in a similar way to antibiotics, with the advantage that it is conceptually more difficult for any one pathogen to evolve resistance to such a therapy.
A second, more challenging problem arises in patients who become critically ill following infection. The patterns of immune-mediated organ dysfunction, immunoparesis, and death are very similar in severe infections and sterile systemic injuries (such as burns, haemorrhage, pancreatitis and trauma). Ultimately, death is a consequence of the host response to injury (Angus and Poll 2013), through final common pathways of organ failure that are clinically and biochemically evident, and unrelated to the original precipitant.
Broadly, the severity of critical illness follows directly from the severity and duration of the initial insult. In bacterial sepsis, early antibiotics are the mainstay of therapy; in influenza, early antivirals; in haemorrhage, early resuscitation; in trauma, urgent action to prevent secondary injury. There are no therapies with which to modulate the host response to systemic injury.
There is a lack of direct evidence of heritability for outcomes of critical illness, due in part to difficulties in defining and quantifying the heterogeneous multi-organ dysfunction syndrome (MODS), and in part due to the rapid pace of change in critical care medicine, making it impossible to tackle this question in long term outcome studies. However, clinical and biological evidence support the hypothesis that the pathogenesis of MODS is immune in origin (Angus and Poll 2013). Hence, predictions can be made from the extensive knowledge of other immune conditions. Whether or not MODS is considered to be an autoimmune or infectious condition is moot: these conditions share a great deal of similarity in genetic predispositions, cell types and mechanisms of pathogenesis. It is therefore very likely that propensity to survive MODS has a heritable component, and there is some direct evidence in support of this hypothesis (Rautanen et al. 2015). If this is the case, then the identity of the specific variants that contribute to outcome could potentially be utilised to design therapies to promote survival after the onset of MODS.
This study aims to identify genetic predisposition to specific syndromes of critical illness. Specifically, susceptibility to life-threatening infections caused by an identified pathogen, and susceptibility to death following the onset of organ failure due to sepsis or sterile injury. In order to maximise the probability of identifying host genetic loci associated with susceptibility, Genomics England will restrict some analyses to younger individuals in good general health and lacking in known predisposing factors.
The same principle was used to determine an upper age limit for inclusion for some analyses. With advancing age, there is an increase in undiagnosed co-morbidity, frailty, and susceptibility to serious complications of infection or critical injury. There is therefore an increase in the probability of susceptibility to, and mortality from, critical illness that is consequent upon non-genetic factors.
Participation Group:
Patients will be identified and recruited in hospital during acute illness. Potential participants will be identified through hospital workers upon presentation at recruiting sites. The disease processes under study have a high mortality, so it is desirable to recruit patients as early as possible in the disease process.
Participants or an appropriate parent/guardian/consultee will be approached by staff trained in consent procedures that protect the rights of the patient and adhere to the ethical principles within the Declaration of Helsinki. Staff will explain the details of the study to the participant or parent/guardian/consultee and allow them time to discuss and ask questions. The staff will review the informed consent form with the person giving consent (or assent) and endeavour to ensure understanding of the contents, including study procedures, risks, benefits, and the right to withdraw. Participants who agree to participate (or their parent/guardian or consultee who declares their wishes to do so) will be asked to sign and date an informed consent form.
In view of the importance of early sampling, participants or their parent/guardian/consultee will be permitted to consent and begin to participate in the study immediately if they wish to do so. Those who prefer more time to consider participation will be approached again after an agreed time, normally one day, to discuss further.
Patients who meet the inclusion/exclusion criteria and who have given informed consent to participate directly, or have been consented by a parent/guardian or whose wishes have been declared by a consultee, will be enrolled to the study.
Samples and data will be collected according to available resources and the weight of the patient will be measured for children under 12 in order to prevent excessive volume sampling. Samples required for medical management will at all times have priority over samples taken for research tests. Aliquots or samples for research purposes should never compromise the quality or quantity of samples required for medical management. Wherever practical, taking research samples should be timed to coincide with clinical sampling. The research team will be responsible for sharing the sampling protocol with health care workers supporting patient management in order to minimise disruption to routine care and avoid unnecessary procedures.
Consent will be sought from patients who survive critical illness and regain capacity to give consent. At each of the follow-up sessions, the investigator gathering data will determine whether the patient has regained capacity. In the event that a patient continues to be incapacitated beyond the follow-up period, the local investigator will plan a subsequent capacity check at a specific date, after an interval to be determined by the nature of the incapacity. The planned dates of capacity checks on incapacitated survivors will be stored locally in the site file, together with a record of the outcome of each check.
Patients who decline to participate at this stage will be removed from the study. Where the patient cannot be recontacted despite best endeavours, they will remain in the study.
Withdrawal:
Participants are freely able to decline participation in this study or to withdraw from participation at any point without suffering any implied or explicit disadvantage. All patients will be treated according to standard practice regardless of whether they participate.
The following options of withdrawal will be made available to participants:
1. Partial withdrawal. Data WILL continue to be updated and used for research, but no further contact will be made with the participant
2. Full withdrawal.
• no further contact will be made with the participant;
• data will not be updated from health records;
• data will not be removed from research that is underway or has already been done, and an audit record will be maintained to confirm participation.
Consent version and life course follow-up
There are a series of consent versions as documented below:
• V1.08 (March 2020) - there are circa 1500 participants on this consent , allows for COVID research but not for longitudinal follow-up. Genomics are exploring reconsenting and CAG as options for this cohort. Currently, these individuals will NOT be submitted to NHS Digital
• V2.1 (April 2020) - Included retention of NHS number and linkage to longitudinal follow -up, NHS Digital are named on the PIS. These individuals will be submitted to NHS Digital for data linkage
• V2.4 (July 2020) - In addition to already stated in v2.1, this covers community recruitment of participants and movement of identifiers. Sharing data with NHS Digital is again stated. These individuals will be submitted to NHS Digital for data linkage
PIS and Consent Version 2.1 and 2.4 are identical in their reference to accessing data from secondary data suppliers and include retention and collection of identifiable information and both name NHS Digital as a source of linkage. All individuals consented are collected with information including version of consent recruited under, thus genomics England can distinguish individuals based on consent. The most up to date protocol is listed on the website (https://genomicc.org/protocol/)
Datasets requested from NHS Digital:
The current application to access to the data sets requested are in line with COVID-19 related purpose, in terms of the restrictions set out in Reg 3(1) COPI. However, based on the now REC approved protocol, we are looking to move this application to Consent basis.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data for Pandemic Planning and Research (GDPPR) extracts to assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings. We understand this dataset remains under COVID regulations and will remain only for COVID research going forward.
Hospital Episode Statistics (HES): Outpatients (OP), Admitted Patient Care (APC), Critical Care (CC) and Accident & Emergency (AE)/ Emergency Care Data Set (ECDS). These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Secondary Uses Service datasets. The minimal latency in availability of these datasets is highly desirable for the research objectives set out in this project.
Mental Health Data sets: The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
Cancer Registration Data sets: To see the incidence of cancer within the cohort
Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
COVID datasets: These datasets, including vaccination, SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus. These will remain only for COVID based research only.
Assessment of the datasets has been undertaken and NHS Digital are satisfied that they are necessary for the COVID-19 work being undertaken. Genomics have confirmed that all research which is approved from the GenoMICC study using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Genomics England Industry access:
Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help turn research findings into treatments, diagnostics and benefits for patients as soon as possible.
As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations to access the COVID data, however they are charged based on storage and compute, such as running their own bioinformatics pipeline. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected.
The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a 'critical friend' and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's landmark data set.
The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
Additionally, as Genomics England will be processing health data - a special category of personal data, they will also be processing data under Article 9 (2)(j) as processing is necessary for archiving purposes in the public interest. Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
Expected output
Access to the data will enable the research to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
Specifically the short term (6 months) and medium term benefits and outcomes from this programme of research anticipated are;
• Variants enable Polygenic Risk Score to predict greatest risk and avoid ITU
• Pre-morbid clinical conditions or biomarkers of risk and rapid NHS uptake to avoid ITU
• Identify novel therapies or precision interventions for rapid national trials
• Longitudinal life course sequel of COVID-19 for pandemic planning
• Patient benefit:
o Providing improved clinical understanding of disease progression in COVID-19
o Correlation to disease progression and pre-morbid status
o Identification of susceptibility genes
o Develop a biomarker test(s) to predict an individual’s response to SARS-CoV-2 exposure, considering both COVID-19 severity and vulnerability to infection.
o Identify targets that can be used in to inform development of new treatments
• New scientific insights and discovery:
o with the consent of patients, creating a database of 35,000 whole genome sequences linked to continually updated long term patient health and personal information for analysis by researchers.
o Correlation of host and viral genomic data
o Potential to provide improved testing for future pandemics
o Aide researchers to identify novel targets for vaccines and therapy
o Identification of highly penetrant rare variants in genes and pathways relating to viral susceptibility or immunodeficiency.
o Genome wide association studies (GWAS) using common variants to identify genes and pathways associated with viral response. These analyses will be aligned with other COVID-19 research consortia.
o Rare variant burden analysis to identify genes and pathways enriched in rare variants associated with viral response
• Accelerating the uptake of genomic medicine in the NHS: working with NHSE and other partners to deliver a scale-able WGS and informatics platform to enable these services to be made widely available for NHS patients. WGS could potentially provide the most accurate diagnostic test for COVID 19.
• Stimulating and enhancing UK industry and investment: by providing access to this unique data resource by industry for the purpose of developing new knowledge, methods of analysis, medicines, diagnostics and devices.
• Increasing public knowledge and support for genomic medicine: delivering an ethical and transparent programme which has public trust and confidence and working with a range of partners to increase knowledge of genomics.
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England Research Environment where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw.
“This is a ground-breaking and far-reaching study which will harness the UK’s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
The Research environment also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public Health England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the Research Environment monthly in order to continue support for, and to further develop, this ground-breaking resource.
Outputs published to date are;
Publications:
• Pairo-Castineira E., Clohisey S., Klaric L., Bretherick A.D., Rawlik K., Pasko D., Walker S., Parkinson N., Fourman M.H., Russell C.D., Furniss J., Richmond A., Gountouna E., Wrobel N., Harrison D., Wang B., Wu Y., Meynert A., Griffiths F., Oosthuyzen W., Kousathanas A., Moutsianas L., Yang Z., Zhai R., Zheng C., Grimes G., Beale R., Millar J., Shih B., Keating S., Zechner M., Haley C., Porteous D.J., Hayward C., Yang J., Knight J., Summers C., Shankar-Hari M., Klenerman P., Turtle L., Ho A., Moore S.C., Hinds C., Horby P., Nichol A., Maslove D., Ling L., McAuley D., Montgomery H., Walsh T., Pereira A.C., Renieri A., GenOMICC I., Investigators, COVID- Human Genetics I., Investigators, BRACOVID I., Gen-COVID I., Shen X., Ponting C.P., Fawkes A., Tenesa A., Caulfield M., Scott R., Rowan K., Murphy L., Openshaw P., Semple M.G., Law A., Vitart V., Wilson J.F., Baillie J.K. Genetic mechanisms of critical illness in COVID-19. Nature. 2020;591(7848):92–98
• Kosmicki JA, Horowitz JE, Banerjee N, Lanche R, Marcketta A, Maxwell E, Bai X, Sun D, Backman JD, Sharma D, Kang HM, O'Dushlaine C, Yadav A, Mansfield AJ, Li AH, Watanabe K, Gurski L, McCarthy SE, Locke AE, Khalid S, O'Keeffe S, Mbatchou J, Chazara O, Huang Y, Kvikstad E, O'Neill A, Nioi P, Parker MM, Petrovski S, Runz H, Szustakowski JD, Wang Q, Wong E, Cordova-Palomera A, Smith EN, Szalma S, Zheng X, Esmaeeli S, Davis JW, Lai YP, Chen X, Justice AE, Leader JB, Mirshahi T, Carey DJ, Verma A, Sirugo G, Ritchie MD, Rader DJ, Povysil G, Goldstein DB, Kiryluk K, Pairo-Castineira E, Rawlik K, Pasko D, Walker S, Meynert A, Kousathanas A, Moutsianas L, Tenesa A, Caulfield M, Scott R, Wilson JF, Baillie JK, Butler-Laporte G, Nakanishi T, Lathrop M, Richards JB, Jones M, Balasubramanian S, Salerno W, Shuldiner AR, Marchini J, Overton JD, Habegger L, Cantor MN, Reid JG, Baras A, Abecasis GR, Ferreira MA. A catalog of associations between rare coding variants and COVID-19 outcomes. medRxiv [Preprint]. 2021 Feb 27:2020.10.28.20221804. doi: 10.1101/2020.10.28.20221804. PMID: 33655273; PMCID: PMC7924298.
• RECOVERY Collaborative Group. Tocilizumab in patients admitted to hospital with COVID-19 (RECOVERY): a randomised, controlled, open-label, platform trial. Lancet. 2021 May 1;397(10285):1637-1645. doi: 10.1016/S0140-6736(21)00676-0. PMID: 33933206; PMCID: PMC8084355.
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
This proposal has already enabled Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allowing investigation of the impact of viral genomic features on outcomes. The ongoing prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Benefits thus far stated below.
Summary of publications and value of study updated December 2021:
Genome wide studies (GWAS) of first 2244 GenOMICC participants admitted with severe COVID
Identified the following genes of interest:
• chromosome 12q24.13 in a gene cluster that encodes antiviral restriction enzyme activators (OAS1, OAS2 and OAS3)
• chromosome 19p13.2 near the gene that encodes tyrosine kinase 2 (TYK2);
• chromosome 19p13.3 within the gene that encodes dipeptidyl peptidase 9 (DPP9);
• chromosome 21q22.1 (rs2236757, P = 4.99 × 10-8) in the interferon receptor gene IFNAR2.
• These targets can utilise and repurpose currently used medications for treatment and management of severe covid.
• evidence that low expression of IFNAR2, or high expression of TYK2, are associated with life-threatening disease; and transcriptome-wide association in lung tissue revealed that high expression of the monocyte-macrophage chemotactic receptor CCR2 is associated with severe COVID-19.
• Our results identify robust genetic signals relating to key host antiviral defence mechanisms and mediators of inflammatory organ damage in COVID-19.
• Both mechanisms may be amenable to targeted treatment with existing drugs.
In addition, two therapeutic drug trials have also commenced.
By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy. Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
DARS-NIC-374190-D0N1M-v3.2 1 October 2021 to 30 September 2022
- Title
- R26 - GENOMICS ENGLAND: GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 20
- Files released
- 0
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); HES-ID to MPS-ID HES Accident and Emergency; HES-ID to MPS-ID HES Admitted Patient Care; HES-ID to MPS-ID HES Outpatients; Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v2.3
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2021-10-01 | |
| End date | 2022-09-30 |
Datasets: + HES-ID to MPS-ID HES Accident and Emergency; + HES-ID to MPS-ID HES Admitted Patient Care; + HES-ID to MPS-ID HES Outpatients
Objective for processing
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
[8 paragraphs unchanged]
6. To provide access to these data sets via the Genomics England
Trusted
Research Environment to international and national academia and industry and facilitate international collaboration on COVID-19.
[8 paragraphs unchanged]
Data will be released into the Genomics England
Trusted
Research Environment where it will be linked with associated clinical data as
[24 words unchanged]
England and data feeds from the Intensive Care National Audit Registry (ICNARC).
[9 paragraphs unchanged]
Details of the sublicence model via the Genomics England
Trusted
Research Environment are supplied within the processing activities section of this agreement.
[51 paragraphs unchanged]
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
[1 paragraph unchanged]
Additionally, as Genomics England will be processing health data - a special category of personal data, they will also be processing data under Article 9 (2)(j) as processing is necessary for archiving purposes in the public interest.
Patients and the public will be at the heart of this programme.
[38 words unchanged]
represented on all committees and working groups and will also meet separately.
[4 paragraphs unchanged]
The only permitted activities under this Data Sharing Agreement (DSA) are for COVID-19 purposes and within bounds of Reg 3(2) COPI. Reg 3 (2) COPI states that: "2) For the purposes of this regulation, “processing” includes any operations, or set of operations set out in regulation 2(2) which are undertaken for the purposes set out in paragraph (1)." The research relates to the monitoring and managing of COVID-19 and would therefore be covered by Reg 3(1)(d) of COPI
Processing activities
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
[8 paragraphs unchanged]
All research analysis on the Genomics England dataset will only be carried out via a secure analysis environment hosted within the Genomics England data
center
centre
- the Genomics England Research Environment. Analytical tools and applications are available
[39 words unchanged]
and out of the Research Environment is governed via an 'Airlock' Policy.
[18 paragraphs unchanged]
Each Discovery Forum member signs a Data Access Agreement with Genomics England.
[16 words unchanged]
number of genomes sequences that can be accessed. It covers the Company's
behavior
behaviour
and working practices: in particular it binds users to Genomics England's Airlock
[40 words unchanged]
by the Company within the Research Environment must receive prior ARC approval.
[19 paragraphs unchanged]
The Research Environment contains External Data (for
example
example,
Hospital Episodes Statistics [HES]) which is subject to data sharing framework contracts
[58 words unchanged]
will be more conservative than those applied to 100,000 Genomes Data alone.
[12 paragraphs unchanged]
Genomics England provides NHS Digital with linking data in order to receive
longtitudinal
longitudinal
data sets. These data sets are delivered to Genomics England by NHS
[16 words unchanged]
the linking data and agrees with NHS Digital the scope of the
longtitunidal
longitudinal
data being provided. Genomics England determines the method of de-identification and storage
[9 words unchanged]
use by approved researchers only. Genomics England determines who these researchers are.
[18 paragraphs unchanged]
Expected output
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
[3 paragraphs unchanged]
Genomic data will be released into the Genomics England
Trusted
Research Environment where it will be linked with associated clinical data.
[5 paragraphs unchanged]
As the GenOMICC study prospectively recruits participants, the aim will be to
[41 words unchanged]
NHS Digital and viral and host genomic data, into the Research Environment
montly
monthly
in order to continue support for, and to further develop, this ground-breaking resource.
[1 paragraph unchanged]
Expected measurable benefits
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension. _____________________________________________________________________________________ [27 paragraphs unchanged]
Benefits reported
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension. _____________________________________________________________________________________ [2 paragraphs unchanged]
Objective for processing
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
This agreement is seeking approval to request data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole Genome Sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS Digital and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people) and UK Biobank data sets (500,000 people) with WGS where more cases may be identified, including those with milder disease and unaffected people.
This mention of linkage of the NHS Digital data to UK Biobank data is purely aspirational at present. This has not been approved under previous iterations of the agreement, and is not covered under v1.2 of this agreement. Genomics England can confirm that if such linkage were to occur - then appropriate documentation (including transparency, patient, consent and ethics materials) would be provided to NHS Digital and a subsequent amended application put in place.
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England Research Environment to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and it’s outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel.
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By undertaking a prospective study design that leverages existing recruitment infrastructure in critical care, together with Genomics England, NHS England, Devolved Nations and Public Health England (PHE) infrastructure this research will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
Genomics England Background: (For Context)
Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.
Data will be released into the Genomics England Research Environment where it will be linked with associated clinical data as well as additional data sources, which will include NHS Digital data requested through this agreement, COVID-19 testing feeds and viral genomics from Public Health England and data feeds from the Intensive Care National Audit Registry (ICNARC).
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable Research Environment for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new Research Environment (COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables effective COVID-19 research outcomes across a wide range of academic and Biotech/Pharma researchers with varying levels of technical competency. This offers a major upgrade to our current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling, and enable shared workspaces between Genomics England and Partners. COVID-RE must serve the immediate COVID-19 research effort and may also advance the transformation of Genomics England’s platform infrastructure, which is a key enabler to the research community.
There are two key value streams provided by the COVID-RE. Firstly ‘raw data to analytics ready data’ stream which must permit the collection of data from multiple locations from unstructured, through semi and fully structured forms and transforms them the appropriate data model, data store based on their ongoing use and availability. Secondly the ‘discovery to insight’ stream that supports researchers by providing the capability and the framework to support their User Journey from understanding the data available to them through to providing the necessary analytics and publishing tools.
The newly established COVID-RE will provide the following:
• Seamless interface with cloud storage and compute capabilities.
• A unified data platform containing datastores appropriate for all the required clinical and genomic data types.
• Data integration capability to deliver analytics-ready datasets into domain-specific data stores across batch and streaming integration patterns.
• Standards-based access services providing secure, fast, flexible, robust and auditable access to TRE data assets.
• Applications to support the research user journey from an exploration of data sets, through cohort building, analysis and publication - providing tools appropriate to a variety of user requirements (ways of working) and programming competency.
• Applications and workflows can be delivered natively within the platform, however COVID-RE must enable access to container-based applications and, via a set of hardened APIs, to a relevant external service (as long as security is maintained).
Details of the sublicence model via the Genomics England Research Environment are supplied within the processing activities section of this agreement.
The additional clinical data is key for researchers to be able to understand and infer clinically relevant and actionable findings. The aim is to provide a high quality, diverse clinical dataset, detailing each participant’s journey and to understand pre-existing conditions and also the early behaviours in the disease course. Genomics England currently have an agreement to receive NHS Digital data for the extant 100,000 Genomes Project participants (DARS-NIC-12784) and have seen first-hand the depth and quality of the data and how it has aided researchers.
To that aim, the significant gap in the current data collection for the COVID-19 project can be addressed by the non-standard NHS Digital data feed.
The Secondary Use Services Admitted Patient Care Feed (SUS APC) would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
The associated COVID-19 data-sets (including NHS111 and CV19 Testing Data in particular which will be requested under a future version of this agreement) will add to the detail in patient journey course, flagging for instance how participants were monitoring their symptoms and was there an associated poor prognosis.
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to approx 36,000 participants.
The data would be available for analysis alongside the extant Genomics England data set of 100,000 Genomes Project participants and would be made available to approved researchers worldwide as per existing governance procedures. This 100,000 Genomes Project data set will act as an appropriate matched control group. Data on the 100,000 Genomes Cohort will be provided to NHS Digital for linkage - (as they will act as a control cohort), as well as new additions for the GenOMICC Study.
The below detail pertains to the background of the study and susceptibility to infection sets out why Genomics are investigating. The origin of the GenOMICC study was focused on smaller cohorts of participants. However, in collaboration with the COG-UK group, this is now expanded to whole genome sequence of approx. 36,000 affected individuals as set out above.
GenOMICC Study Background:
Susceptibility to infection is profoundly heritable (Sorensen et al. 1988). Patients who develop life-threatening illness following infection with usually innocuous pathogens, such as influenza (Miller et al. 2010), are genetically different from the rest of the population (Albright et al. 2008). Understanding the genetic mechanisms of susceptibility may yield new therapeutic targets (Baillie 2014) that can be used to make susceptible patients more similar to individuals who are resistant to, or tolerant of, specific pathogens.
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV [Human Immunodeficiency Virus] (Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions (e.g. small molecules to inhibit an enzyme or receptor that is dysfunctional in resistant individuals) would therefore be protective to the host in a similar way to antibiotics, with the advantage that it is conceptually more difficult for any one pathogen to evolve resistance to such a therapy.
A second, more challenging problem arises in patients who become critically ill following infection. The patterns of immune-mediated organ dysfunction, immunoparesis, and death are very similar in severe infections and sterile systemic injuries (such as burns, haemorrhage, pancreatitis and trauma). Ultimately, death is a consequence of the host response to injury (Angus and Poll 2013), through final common pathways of organ failure that are clinically and biochemically evident, and unrelated to the original precipitant.
Broadly, the severity of critical illness follows directly from the severity and duration of the initial insult. In bacterial sepsis, early antibiotics are the mainstay of therapy; in influenza, early antivirals; in haemorrhage, early resuscitation; in trauma, urgent action to prevent secondary injury. There are no therapies with which to modulate the host response to systemic injury.
There is a lack of direct evidence of heritability for outcomes of critical illness, due in part to difficulties in defining and quantifying the heterogeneous multi-organ dysfunction syndrome (MODS), and in part due to the rapid pace of change in critical care medicine, making it impossible to tackle this question in long term outcome studies. However, clinical and biological evidence support the hypothesis that the pathogenesis of MODS is immune in origin (Angus and Poll 2013). Hence, predictions can be made from the extensive knowledge of other immune conditions. Whether or not MODS is considered to be an autoimmune or infectious condition is moot: these conditions share a great deal of similarity in genetic predispositions, cell types and mechanisms of pathogenesis. It is therefore very likely that propensity to survive MODS has a heritable component, and there is some direct evidence in support of this hypothesis (Rautanen et al. 2015). If this is the case, then the identity of the specific variants that contribute to outcome could potentially be utilised to design therapies to promote survival after the onset of MODS.
This study aims to identify genetic predisposition to specific syndromes of critical illness. Specifically, susceptibility to life-threatening infections caused by an identified pathogen, and susceptibility to death following the onset of organ failure due to sepsis or sterile injury. In order to maximise the probability of identifying host genetic loci associated with susceptibility, Genomics England will restrict some analyses to younger individuals in good general health and lacking in known predisposing factors.
The same principle was used to determine an upper age limit for inclusion for some analyses. With advancing age, there is an increase in undiagnosed co-morbidity, frailty, and susceptibility to serious complications of infection or critical injury. There is therefore an increase in the probability of susceptibility to, and mortality from, critical illness that is consequent upon non-genetic factors.
Participation Group:
Patients will be identified and recruited in hospital during acute illness. Potential participants will be identified through hospital workers upon presentation at recruiting sites. The disease processes under study have a high mortality, so it is desirable to recruit patients as early as possible in the disease process.
Participants or an appropriate parent/guardian/consultee will be approached by staff trained in consent procedures that protect the rights of the patient and adhere to the ethical principles within the Declaration of Helsinki. Staff will explain the details of the study to the participant or parent/guardian/consultee and allow them time to discuss and ask questions. The staff will review the informed consent form with the person giving consent (or assent) and endeavour to ensure understanding of the contents, including study procedures, risks, benefits, and the right to withdraw. Participants who agree to participate (or their parent/guardian or consultee who declares their wishes to do so) will be asked to sign and date an informed consent form.
In view of the importance of early sampling, participants or their parent/guardian/consultee will be permitted to consent and begin to participate in the study immediately if they wish to do so. Those who prefer more time to consider participation will be approached again after an agreed time, normally one day, to discuss further.
Patients who meet the inclusion/exclusion criteria and who have given informed consent to participate directly, or have been consented by a parent/guardian or whose wishes have been declared by a consultee, will be enrolled to the study.
Samples and data will be collected according to available resources and the weight of the patient will be measured for children under 12 in order to prevent excessive volume sampling. Samples required for medical management will at all times have priority over samples taken for research tests. Aliquots or samples for research purposes should never compromise the quality or quantity of samples required for medical management. Wherever practical, taking research samples should be timed to coincide with clinical sampling. The research team will be responsible for sharing the sampling protocol with health care workers supporting patient management in order to minimise disruption to routine care and avoid unnecessary procedures.
Consent will be sought from patients who survive critical illness and regain capacity to give consent. At each of the follow-up sessions, the investigator gathering data will determine whether the patient has regained capacity. In the event that a patient continues to be incapacitated beyond the follow-up period, the local investigator will plan a subsequent capacity check at a specific date, after an interval to be determined by the nature of the incapacity. The planned dates of capacity checks on incapacitated survivors will be stored locally in the site file, together with a record of the outcome of each check.
Patients who decline to participate at this stage will be removed from the study. Where the patient cannot be recontacted despite best endeavours, they will remain in the study.
Withdrawal:
Participants are freely able to decline participation in this study or to withdraw from participation at any point without suffering any implied or explicit disadvantage. All patients will be treated according to standard practice regardless of whether they participate.
The following options of withdrawal will be made available to participants:
1. Partial withdrawal. Data WILL continue to be updated and used for research, but no further contact will be made with the participant
2. Full withdrawal.
• no further contact will be made with the participant;
• data will not be updated from health records;
• data will not be removed from research that is underway or has already been done, and an audit record will be maintained to confirm participation.
Consent version and lifecourse follow-up
The first COVID positive patient was recruited to the GenOMICC study in March 2020. All patients (c.1500) currently recruited to the GenOMICC study are on consent and Protocol version 1.08 (which allows for COVID research but not longitudinal life course follow up). This Protocol v2.1 (submitted as part of this DARS request) was Research Ethics Committe (REC) approved on 23rd March 2020, IRAS [Integrated Research Application System] IDs are: 269326 & 189676 (https://www.hra.nhs.uk/covid-19-research/approved-covid-19-research/269326/). Genomics England will attempt to reconsent all patients recruited under version 1.08 onto the newly amended materials (version 2.5 of the patient information sheet – submitted as part of this DARS application) which allows for longitudinal lifecourse follow-up. For any patients on the v1.08 Protocol for which Genomics England are not able to seek consent, Genomics England will not be requesting their data for linkage. For patients prospectively recruited onto the v2.1 Protocol and consent materials, Genomics England will be requesting their data for linkage and follow-up .
Datasets requested from NHS Digital:
Consideration has been be given to assess and ensure that access to the data sets requested are inline with COVID-19 related purpose, in terms of the restrictions set out in Reg 3(1) COPI, below are details of what each of the data sets will at a high level provide insight to.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data for Pandemic Planning and Research (GDPPR) extracts to assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings.
Hospital Episode Statistics (HES): Outpatients (OP), Admitted Patient Care (APC), Critical Care (CC) and Accident & Emergency (AE)/ Emergency Care Data Set (ECDS). These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Secondary Uses Service datasets. The minimal latency in availability of these datasets is highly desirable for the research objectives set out in this project.
Mental Health Data sets: The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
Cancer Registration Data sets: To see the incidence of cancer within the cohort
Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
COVID datasets: These datasets, including SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus.
Assessment of the datasets has been undertaken and NHS Digital are satisfied that they are necessary for the COVID-19 work being undertaken. Genomics have confirmed that all research which is approved from the GenoMICC study using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Genomics England Industry access:
Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help turn research findings into treatments, diagnostics and benefits for patients as soon as possible.
As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations to access the COVID data, however they are charged based on storage and compute, such as running their own bioinformatics pipeline. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected.
The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a 'critical friend' and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's landmark data set.
The lawful basis for the release and use of the confidential data being shared under this version of the agreement is Regulation 3(4) of the National Health Service (Control of Patient Information Regulations) 2002 (COPI) to require NHS Digital to share confidential patient information with organisations entitled to process this under COPI for COVID-19 purposes. The application of this has been based on the information provided in the Whole genome sequencing of patients severely affected by COVID-19 funding proposal from The GenOMICC - COVID Genomics UK (CoG-UK) partnership which was supported by the CMO of England and the CFO of the DHSC.
The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
Additionally, as Genomics England will be processing health data - a special category of personal data, they will also be processing data under Article 9 (2)(j) as processing is necessary for archiving purposes in the public interest. Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
Expected output
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
Genomics England completed sequencing 100,000 genomes at the end of 2018 (https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000- genomes-sequenced/). During 2018, the Genomics England Research Environment was established to allow research access to de-identified genomic and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with now over 3000 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England Research Environment where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw.
“This is a ground-breaking and far-reaching study which will harness the UK’s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
As of March 2020, the Genomics England Research Environment contained 107,694 genomes, of which, 33,461 were cancer and 74,233 were rare diseases.
The Research environment also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public Health England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the Research Environment monthly in order to continue support for, and to further develop, this ground-breaking resource.
Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants and GenOMICC participants into the Research Environment on the dates shown above.
Benefits reported
SEPTEMBER 2021 - v3 is a extension at no cost to keep customer in agreement ONLY. No data is to flow under this Extension.
_____________________________________________________________________________________
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
The GenOMICC study is very much in its infancy and whole genome sequencing has only begun in the last month. It is thus too early to demonstrate any significant outcomes. These outcomes will only gain power and relevance with more prospectively recruited patients and a breadth and depth of clinical data. By providing this to researchers, Genomics England can provide them the necessary tools to explore the genomic data.
DARS-NIC-374190-D0N1M-v2.3 1 April 2021 to 30 September 2021
- Title
- R26 - GENOMICS ENGLAND: GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 17
- Files released
- 461
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v1.3
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2021-04-01 | |
| End date | 2021-09-30 | |
| Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset: legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set: legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| COVID-19 Hospitalization in England Surveillance System: legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| COVID-19 SGSS First Positives (Second Generation Surveillance System): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Cancer Registration Data: legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Civil Registrations of Death: legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Community Services Data Set (CSDS): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Demographics: legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Diagnostic Imaging Data Set (DID): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Emergency Care Data Set (ECDS): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Hospital Episode Statistics Accident and Emergency (HES A and E): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Hospital Episode Statistics Admitted Patient Care (HES APC): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Hospital Episode Statistics Critical Care (HES Critical Care): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Hospital Episode Statistics Outpatients (HES OP): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Mental Health Services Data Set (MHSDS): legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) | |
| Secondary Uses Service Payment By Results Spells: legal basis | CV19: Regulation 3 (4) of the Health Service (Control of Patient Information) Regulations 2002; Health and Social Care Act 2012 - s261(5)(d) |
Objective for processing
[60 paragraphs unchanged]
The first COVID positive patient was recruited to the GenOMICC study in
[67 words unchanged]
all patients recruited under version 1.08 onto the newly amended materials (version
2.1
2.5 of the patient information sheet
– submitted as part of this DARS application) which allows for longitudinal
[40 words unchanged]
Genomics England will be requesting their data for linkage and follow-up .
[16 paragraphs unchanged]
The lawful basis for the release and use of the confidential data being shared under this version of the agreement is Regulation 3(4) of the National Health Service (Control of Patient Information Regulations) 2002 (COPI) to require NHS Digital to share confidential patient information with organisations entitled to process this under COPI for COVID-19 purposes. The application of this has been based on the information provided in the Whole genome sequencing of patients severely affected by COVID-19 funding proposal from The GenOMICC - COVID Genomics UK (CoG-UK) partnership which was supported by the CMO of England and the CFO of the DHSC.
[8 paragraphs unchanged]
The lawful basis for the release and use of the confidential data being shared under this version of the agreement is Regulation 3(4) of the National Health Service (Control of Patient Information Regulations) 2002 (COPI) to require NHS Digital to share confidential patient information with organisations entitled to process this under COPI for COVID-19 purposes. The application of this has been based on the information provided in the Whole genome sequencing of patients severely affected by COVID-19 funding proposal from The GenOMICC - COVID Genomics UK (CoG-UK) partnership which was supported by the CMO of England and the CFO of the DHSC.
[1 paragraph unchanged]
Expected output
[9 paragraphs unchanged]
As the GenOMICC study prospectively recruits participants, the aim will be to
[41 words unchanged]
NHS Digital and viral and host genomic data, into the Research Environment
on the following dates,
montly
in order to continue support for, and to further develop, this ground-breaking
resource:
resource.
• 3rd August 2020
• 7th September 2020
• 5th October 2020
• 2nd November 2020
[1 paragraph unchanged]
Benefits reported
Not stated in the previous version; added here.
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
The GenOMICC study is very much in its infancy and whole genome sequencing has only begun in the last month. It is thus too early to demonstrate any significant outcomes. These outcomes will only gain power and relevance with more prospectively recruited patients and a breadth and depth of clinical data. By providing this to researchers, Genomics England can provide them the necessary tools to explore the genomic data.
Unchanged: Processing activities, Expected measurable benefits.
Objective for processing
This agreement is seeking approval to request data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole genome sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS Digital and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people) and UK Biobank data sets (500,000 people) with WGS where more cases may be identified, including those with milder disease and unaffected people.
This mention of linkage of the NHS Digital data to UK Biobank data is purely aspirational at present. This has not been approved under previous iterations of the agreement, and is not covered under v1.2 of this agreement. Genomics England can confirm that if such linkage were to occur - then appropriate documentation (including transparency, patient, consent and ethics materials) would be provided to NHS Digital and a subsequent amended application put in place.
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England Trusted Research Environment to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and it’s outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel.
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By undertaking a prospective study design that leverages existing recruitment infrastructure in critical care, together with Genomics England, NHS England, Devolved Nations and Public Health England (PHE) infrastructure this research will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
Genomics England Background: (For Context)
Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.
Data will be released into the Genomics England Trusted Research Environment where it will be linked with associated clinical data as well as additional data sources, which will include NHS Digital data requested through this agreement, COVID-19 testing feeds and viral genomics from Public Health England and data feeds from the Intensive Care National Audit Registry (ICNARC).
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable Research Environment for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new Research Environment (COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables effective COVID-19 research outcomes across a wide range of academic and Biotech/Pharma researchers with varying levels of technical competency. This offers a major upgrade to our current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling, and enable shared workspaces between Genomics England and Partners. COVID-RE must serve the immediate COVID-19 research effort and may also advance the transformation of Genomics England’s platform infrastructure, which is a key enabler to the research community.
There are two key value streams provided by the COVID-RE. Firstly ‘raw data to analytics ready data’ stream which must permit the collection of data from multiple locations from unstructured, through semi and fully structured forms and transforms them the appropriate data model, data store based on their ongoing use and availability. Secondly the ‘discovery to insight’ stream that supports researchers by providing the capability and the framework to support their User Journey from understanding the data available to them through to providing the necessary analytics and publishing tools.
The newly established COVID-RE will provide the following:
• Seamless interface with cloud storage and compute capabilities.
• A unified data platform containing datastores appropriate for all the required clinical and genomic data types.
• Data integration capability to deliver analytics-ready datasets into domain-specific data stores across batch and streaming integration patterns.
• Standards-based access services providing secure, fast, flexible, robust and auditable access to TRE data assets.
• Applications to support the research user journey from an exploration of data sets, through cohort building, analysis and publication - providing tools appropriate to a variety of user requirements (ways of working) and programming competency.
• Applications and workflows can be delivered natively within the platform, however COVID-RE must enable access to container-based applications and, via a set of hardened APIs, to a relevant external service (as long as security is maintained).
Details of the sublicence model via the Genomics England Trusted Research Environment are supplied within the processing activities section of this agreement.
The additional clinical data is key for researchers to be able to understand and infer clinically relevant and actionable findings. The aim is to provide a high quality, diverse clinical dataset, detailing each participant’s journey and to understand pre-existing conditions and also the early behaviours in the disease course. Genomics England currently have an agreement to receive NHS Digital data for the extant 100,000 Genomes Project participants (DARS-NIC-12784) and have seen first-hand the depth and quality of the data and how it has aided researchers.
To that aim, the significant gap in the current data collection for the COVID-19 project can be addressed by the non-standard NHS Digital data feed.
The Secondary Use Services Admitted Patient Care Feed (SUS APC) would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
The associated COVID-19 data-sets (including NHS111 and CV19 Testing Data in particular which will be requested under a future version of this agreement) will add to the detail in patient journey course, flagging for instance how participants were monitoring their symptoms and was there an associated poor prognosis.
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to approx 36,000 participants.
The data would be available for analysis alongside the extant Genomics England data set of 100,000 Genomes Project participants and would be made available to approved researchers worldwide as per existing governance procedures. This 100,000 Genomes Project data set will act as an appropriate matched control group. Data on the 100,000 Genomes Cohort will be provided to NHS Digital for linkage - (as they will act as a control cohort), as well as new additions for the GenOMICC Study.
The below detail pertains to the background of the study and susceptibility to infection sets out why Genomics are investigating. The origin of the GenOMICC study was focused on smaller cohorts of participants. However, in collaboration with the COG-UK group, this is now expanded to whole genome sequence of approx. 36,000 affected individuals as set out above.
GenOMICC Study Background:
Susceptibility to infection is profoundly heritable (Sorensen et al. 1988). Patients who develop life-threatening illness following infection with usually innocuous pathogens, such as influenza (Miller et al. 2010), are genetically different from the rest of the population (Albright et al. 2008). Understanding the genetic mechanisms of susceptibility may yield new therapeutic targets (Baillie 2014) that can be used to make susceptible patients more similar to individuals who are resistant to, or tolerant of, specific pathogens.
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV [Human Immunodeficiency Virus] (Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions (e.g. small molecules to inhibit an enzyme or receptor that is dysfunctional in resistant individuals) would therefore be protective to the host in a similar way to antibiotics, with the advantage that it is conceptually more difficult for any one pathogen to evolve resistance to such a therapy.
A second, more challenging problem arises in patients who become critically ill following infection. The patterns of immune-mediated organ dysfunction, immunoparesis, and death are very similar in severe infections and sterile systemic injuries (such as burns, haemorrhage, pancreatitis and trauma). Ultimately, death is a consequence of the host response to injury (Angus and Poll 2013), through final common pathways of organ failure that are clinically and biochemically evident, and unrelated to the original precipitant.
Broadly, the severity of critical illness follows directly from the severity and duration of the initial insult. In bacterial sepsis, early antibiotics are the mainstay of therapy; in influenza, early antivirals; in haemorrhage, early resuscitation; in trauma, urgent action to prevent secondary injury. There are no therapies with which to modulate the host response to systemic injury.
There is a lack of direct evidence of heritability for outcomes of critical illness, due in part to difficulties in defining and quantifying the heterogeneous multi-organ dysfunction syndrome (MODS), and in part due to the rapid pace of change in critical care medicine, making it impossible to tackle this question in long term outcome studies. However, clinical and biological evidence support the hypothesis that the pathogenesis of MODS is immune in origin (Angus and Poll 2013). Hence, predictions can be made from the extensive knowledge of other immune conditions. Whether or not MODS is considered to be an autoimmune or infectious condition is moot: these conditions share a great deal of similarity in genetic predispositions, cell types and mechanisms of pathogenesis. It is therefore very likely that propensity to survive MODS has a heritable component, and there is some direct evidence in support of this hypothesis (Rautanen et al. 2015). If this is the case, then the identity of the specific variants that contribute to outcome could potentially be utilised to design therapies to promote survival after the onset of MODS.
This study aims to identify genetic predisposition to specific syndromes of critical illness. Specifically, susceptibility to life-threatening infections caused by an identified pathogen, and susceptibility to death following the onset of organ failure due to sepsis or sterile injury. In order to maximise the probability of identifying host genetic loci associated with susceptibility, Genomics England will restrict some analyses to younger individuals in good general health and lacking in known predisposing factors.
The same principle was used to determine an upper age limit for inclusion for some analyses. With advancing age, there is an increase in undiagnosed co-morbidity, frailty, and susceptibility to serious complications of infection or critical injury. There is therefore an increase in the probability of susceptibility to, and mortality from, critical illness that is consequent upon non-genetic factors.
Participation Group:
Patients will be identified and recruited in hospital during acute illness. Potential participants will be identified through hospital workers upon presentation at recruiting sites. The disease processes under study have a high mortality, so it is desirable to recruit patients as early as possible in the disease process.
Participants or an appropriate parent/guardian/consultee will be approached by staff trained in consent procedures that protect the rights of the patient and adhere to the ethical principles within the Declaration of Helsinki. Staff will explain the details of the study to the participant or parent/guardian/consultee and allow them time to discuss and ask questions. The staff will review the informed consent form with the person giving consent (or assent) and endeavour to ensure understanding of the contents, including study procedures, risks, benefits, and the right to withdraw. Participants who agree to participate (or their parent/guardian or consultee who declares their wishes to do so) will be asked to sign and date an informed consent form.
In view of the importance of early sampling, participants or their parent/guardian/consultee will be permitted to consent and begin to participate in the study immediately if they wish to do so. Those who prefer more time to consider participation will be approached again after an agreed time, normally one day, to discuss further.
Patients who meet the inclusion/exclusion criteria and who have given informed consent to participate directly, or have been consented by a parent/guardian or whose wishes have been declared by a consultee, will be enrolled to the study.
Samples and data will be collected according to available resources and the weight of the patient will be measured for children under 12 in order to prevent excessive volume sampling. Samples required for medical management will at all times have priority over samples taken for research tests. Aliquots or samples for research purposes should never compromise the quality or quantity of samples required for medical management. Wherever practical, taking research samples should be timed to coincide with clinical sampling. The research team will be responsible for sharing the sampling protocol with health care workers supporting patient management in order to minimise disruption to routine care and avoid unnecessary procedures.
Consent will be sought from patients who survive critical illness and regain capacity to give consent. At each of the follow-up sessions, the investigator gathering data will determine whether the patient has regained capacity. In the event that a patient continues to be incapacitated beyond the follow-up period, the local investigator will plan a subsequent capacity check at a specific date, after an interval to be determined by the nature of the incapacity. The planned dates of capacity checks on incapacitated survivors will be stored locally in the site file, together with a record of the outcome of each check.
Patients who decline to participate at this stage will be removed from the study. Where the patient cannot be recontacted despite best endeavours, they will remain in the study.
Withdrawal:
Participants are freely able to decline participation in this study or to withdraw from participation at any point without suffering any implied or explicit disadvantage. All patients will be treated according to standard practice regardless of whether they participate.
The following options of withdrawal will be made available to participants:
1. Partial withdrawal. Data WILL continue to be updated and used for research, but no further contact will be made with the participant
2. Full withdrawal.
• no further contact will be made with the participant;
• data will not be updated from health records;
• data will not be removed from research that is underway or has already been done, and an audit record will be maintained to confirm participation.
Consent version and lifecourse follow-up
The first COVID positive patient was recruited to the GenOMICC study in March 2020. All patients (c.1500) currently recruited to the GenOMICC study are on consent and Protocol version 1.08 (which allows for COVID research but not longitudinal life course follow up). This Protocol v2.1 (submitted as part of this DARS request) was Research Ethics Committe (REC) approved on 23rd March 2020, IRAS [Integrated Research Application System] IDs are: 269326 & 189676 (https://www.hra.nhs.uk/covid-19-research/approved-covid-19-research/269326/). Genomics England will attempt to reconsent all patients recruited under version 1.08 onto the newly amended materials (version 2.5 of the patient information sheet – submitted as part of this DARS application) which allows for longitudinal lifecourse follow-up. For any patients on the v1.08 Protocol for which Genomics England are not able to seek consent, Genomics England will not be requesting their data for linkage. For patients prospectively recruited onto the v2.1 Protocol and consent materials, Genomics England will be requesting their data for linkage and follow-up .
Datasets requested from NHS Digital:
Consideration has been be given to assess and ensure that access to the data sets requested are inline with COVID-19 related purpose, in terms of the restrictions set out in Reg 3(1) COPI, below are details of what each of the data sets will at a high level provide insight to.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data for Pandemic Planning and Research (GDPPR) extracts to assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings.
Hospital Episode Statistics (HES): Outpatients (OP), Admitted Patient Care (APC), Critical Care (CC) and Accident & Emergency (AE)/ Emergency Care Data Set (ECDS). These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Secondary Uses Service datasets. The minimal latency in availability of these datasets is highly desirable for the research objectives set out in this project.
Mental Health Data sets: The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
Cancer Registration Data sets: To see the incidence of cancer within the cohort
Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
COVID datasets: These datasets, including SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus.
Assessment of the datasets has been undertaken and NHS Digital are satisfied that they are necessary for the COVID-19 work being undertaken. Genomics have confirmed that all research which is approved from the GenoMICC study using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Genomics England Industry access:
Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help turn research findings into treatments, diagnostics and benefits for patients as soon as possible.
As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations to access the COVID data, however they are charged based on storage and compute, such as running their own bioinformatics pipeline. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected.
The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a 'critical friend' and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's landmark data set.
The lawful basis for the release and use of the confidential data being shared under this version of the agreement is Regulation 3(4) of the National Health Service (Control of Patient Information Regulations) 2002 (COPI) to require NHS Digital to share confidential patient information with organisations entitled to process this under COPI for COVID-19 purposes. The application of this has been based on the information provided in the Whole genome sequencing of patients severely affected by COVID-19 funding proposal from The GenOMICC - COVID Genomics UK (CoG-UK) partnership which was supported by the CMO of England and the CFO of the DHSC.
The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The only permitted activities under this Data Sharing Agreement (DSA) are for COVID-19 purposes and within bounds of Reg 3(2) COPI. Reg 3 (2) COPI states that: "2) For the purposes of this regulation, “processing” includes any operations, or set of operations set out in regulation 2(2) which are undertaken for the purposes set out in paragraph (1)." The research relates to the monitoring and managing of COVID-19 and would therefore be covered by Reg 3(1)(d) of COPI
Expected output
Genomics England completed sequencing 100,000 genomes at the end of 2018 (https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000- genomes-sequenced/). During 2018, the Genomics England Research Environment was established to allow research access to de-identified genomic and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with now over 3000 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England Trusted Research Environment where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw.
“This is a ground-breaking and far-reaching study which will harness the UK’s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
As of March 2020, the Genomics England Research Environment contained 107,694 genomes, of which, 33,461 were cancer and 74,233 were rare diseases.
The Research environment also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public HeaLth England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the Research Environment montly in order to continue support for, and to further develop, this ground-breaking resource.
Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants and GenOMICC participants into the Research Environment on the dates shown above.
Benefits reported
The 100,000 genomes project has been hugely successful and provided numerous academic and clinical publications and discoveries. The success of the project has been based on the strength of the clinical data provided by NHS Digital. Understanding this significant value is why Genomics England are so keen to add NHS Digital data to its clinical data source for the GenOMICC study.
The GenOMICC study is very much in its infancy and whole genome sequencing has only begun in the last month. It is thus too early to demonstrate any significant outcomes. These outcomes will only gain power and relevance with more prospectively recruited patients and a breadth and depth of clinical data. By providing this to researchers, Genomics England can provide them the necessary tools to explore the genomic data.
DARS-NIC-374190-D0N1M-v1.3 1 October 2020 to 31 March 2021
- Title
- R26 - GENOMICS ENGLAND: GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 17
- Files released
- 437
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
What changed from DARS-NIC-374190-D0N1M-v0.4
Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.
| Field | Was | Became |
|---|---|---|
| Start date | 2020-10-01 | |
| End date | 2021-03-31 |
Datasets: + COVID-19 General Practice Extraction Service (GPES) Data for Pandemic Planning and Research (GDPPR)
Objective for processing
This agreement is seeking approval to request data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole genome sequencing
(WGS)
of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England
(CMO)
[3 paragraphs unchanged]
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS Digital and Health Data Research
UK. We
UK (HDRUK). Genomics England
will include deep immune “omic” datasets on a subset of
patients 9.
patients.
This will allow case-control studies that capitalise upon unique UK assets, such
[18 words unchanged]
cases may be identified, including those with milder disease and unaffected people.
This mention of linkage of the NHS Digital data to UK Biobank data is purely aspirational at present. This has not been approved under previous iterations of the agreement, and is not covered under v1.2 of this agreement. Genomics England can confirm that if such linkage were to occur - then appropriate documentation (including transparency, patient, consent and ethics materials) would be provided to NHS Digital and a subsequent amended application put in place.
[6 paragraphs unchanged]
The variable response to COVID-19 suggests that, as with susceptibility to other
[45 words unchanged]
in critical care, together with Genomics England, NHS England, Devolved Nations and
PHE
Public Health England (PHE)
infrastructure this research will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
[1 paragraph unchanged]
The retrospective cohorts offer control arms for the study but also new
[82 words unchanged]
co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000
WGS
whole genome sequences (WGS)
this year, building to 500,000 WGS over the next 18 months from
[16 words unchanged]
COVID-19 but will provide additional cases and controls, including mildly affected individuals.
Over the next five years these datasets will be enriched further by consented individuals from the NHS Genomic Medicine Service, and the new Genomics England programmes proposed in their strategic plan under consideration by Government. Furthermore, the Accelerating Detection of Disease Cohort will enrol up to 5 million people with genome-wide variants available with longitudinal life course data sets that could particularly add value to host response-variant associations and polygenic risk scores.
Genomics England Background: (For Context)
Genomics England Background:
[14 paragraphs unchanged]
The
SUS APC feed
Secondary Use Services Admitted Patient Care Feed (SUS APC)
would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
[6 paragraphs unchanged]
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV
[Human Immunodeficiency Virus]
(Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific
[37 words unchanged]
difficult for any one pathogen to evolve resistance to such a therapy.
[5 paragraphs unchanged]
Participation Group:
Patients will be identified and recruited in hospital during acute illness. Potential participants will be identified through hospital workers upon presentation at recruiting sites. The disease processes under study have a high mortality, so it is desirable to recruit patients as early as possible in the disease process.
Participants or an appropriate parent/guardian/consultee will be approached by staff trained in consent procedures that protect the rights of the patient and adhere to the ethical principles within the Declaration of Helsinki. Staff will explain the details of the study to the participant or parent/guardian/consultee and allow them time to discuss and ask questions. The staff will review the informed consent form with the person giving consent (or assent) and endeavour to ensure understanding of the contents, including study procedures, risks, benefits, and the right to withdraw. Participants who agree to participate (or their parent/guardian or consultee who declares their wishes to do so) will be asked to sign and date an informed consent form.
In view of the importance of early sampling, participants or their parent/guardian/consultee will be permitted to consent and begin to participate in the study immediately if they wish to do so. Those who prefer more time to consider participation will be approached again after an agreed time, normally one day, to discuss further.
Patients who meet the inclusion/exclusion criteria and who have given informed consent to participate directly, or have been consented by a parent/guardian or whose wishes have been declared by a consultee, will be enrolled to the study.
Samples and data will be collected according to available resources and the weight of the patient will be measured for children under 12 in order to prevent excessive volume sampling. Samples required for medical management will at all times have priority over samples taken for research tests. Aliquots or samples for research purposes should never compromise the quality or quantity of samples required for medical management. Wherever practical, taking research samples should be timed to coincide with clinical sampling. The research team will be responsible for sharing the sampling protocol with health care workers supporting patient management in order to minimise disruption to routine care and avoid unnecessary procedures.
Consent will be sought from patients who survive critical illness and regain capacity to give consent. At each of the follow-up sessions, the investigator gathering data will determine whether the patient has regained capacity. In the event that a patient continues to be incapacitated beyond the follow-up period, the local investigator will plan a subsequent capacity check at a specific date, after an interval to be determined by the nature of the incapacity. The planned dates of capacity checks on incapacitated survivors will be stored locally in the site file, together with a record of the outcome of each check.
Patients who decline to participate at this stage will be removed from the study. Where the patient cannot be recontacted despite best endeavours, they will remain in the study.
[9 paragraphs unchanged]
The first COVID positive patient was recruited to the GenOMICC study in
[29 words unchanged]
up). This Protocol v2.1 (submitted as part of this DARS request) was
REC
Research Ethics Committe (REC)
approved on 23rd March 2020, IRAS
[Integrated Research Application System]
IDs are: 269326 & 189676 (https://www.hra.nhs.uk/covid-19-research/approved-covid-19-research/269326/). Genomics England will attempt to reconsent
[65 words unchanged]
Genomics England will be requesting their data for linkage and follow-up .
[2 paragraphs unchanged]
Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency/ ECDS. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data for Pandemic Planning and Research (GDPPR) extracts to assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings.
Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants' phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Hospital Episode Statistics (HES): Outpatients (OP), Admitted Patient Care (APC), Critical Care (CC) and Accident & Emergency (AE)/ Emergency Care Data Set (ECDS). These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
[9 paragraphs unchanged]
As the Discovery Forum is a collaborative venture, no fees are levied
[37 words unchanged]
at the point at which intellectual property for any product is protected.
Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their de-identified genome and health data.
[1 paragraph unchanged]
The lawful basis for processing Participant Data under the General Data Protection:
The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
[2 paragraphs unchanged]
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then
we
others
will
add others
be added
who have been affected by COVID-19 at a later point. These participants
[7 words unchanged]
represented on all committees and working groups and will also meet separately.
[5 paragraphs unchanged]
The only permitted activities under this Data Sharing Agreement (DSA) are for COVID-19 purposes and within bounds of Reg 3(2) COPI. Reg 3 (2) COPI states that: "2) For the purposes of this regulation, “processing” includes any operations, or set of operations set out in regulation 2(2) which are undertaken for the purposes set out in paragraph (1)." The research relates to the monitoring and managing of COVID-19 and would therefore be covered by Reg 3(1)(d) of COPI
Processing activities
[15 paragraphs unchanged]
o UK and foreign governmental departments that carry out significant research activity (e.g.
MRC, NIH, PHE)
Medical Research Council, National Institute for Health, Public Health England)
[6 paragraphs unchanged]
o The applicant's GeCIP domain has submitted a detailed research plan and it has been approved by the Genomics England Access Review Committee
[ARC]
(see below).
[44 paragraphs unchanged]
o Approved researchers will only be able to access Lifebit’s PaaS
[Platform-as-a-service]
CloudOS
[Operating System]
through a virtual desktop.
[3 paragraphs unchanged]
o A security and DPIA
assessment
(Data Protection Impact Assessment)
will be conducted prior to loading live
[1 paragraph unchanged]
Lifebit has been selected as platform partner to deliver the Research Environment after reviewing several proposals. The UK-based SME
(Small Medium Enterprise)
offered a proven and innovative technology solution offering a blend of robustness and ease of use. Lifebit CloudOS
(Operating System)
provides a secure and collaborative workspace to enable researchers to easily perform
[25 words unchanged]
range of academic and biotech/pharma researchers with varying levels of technical competency.
[5 paragraphs unchanged]
Expected measurable benefits
[26 paragraphs unchanged] This proposal could enable Genomics England to discover new rare and common variants alongside new multi-omic biomarkers that underpin host response to infection, allow investigation of the impact of viral genomic features on outcomes and allow creation of a polygenic risk score, which may detect risk of severe response to similar viruses. The prospective component could allow nested clinical trials or case-control resources to add value to this study by detecting variants, which stratify response or predict outcomes. Although the 100,000 Genomes Project and the Genomic Medicine Service may include participants biased to specific disease ascertainment, the scale of these resources and the presence of parents helps compensate for this problem.
Benefits reported
Stated in the previous version and removed here.
Yielded Benefits is not a requirement for new applications.
Unchanged: Expected output.
Objective for processing
This agreement is seeking approval to request data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole genome sequencing (WGS) of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England (CMO)
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS Digital and Health Data Research UK (HDRUK). Genomics England will include deep immune “omic” datasets on a subset of patients. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people) and UK Biobank data sets (500,000 people) with WGS where more cases may be identified, including those with milder disease and unaffected people.
This mention of linkage of the NHS Digital data to UK Biobank data is purely aspirational at present. This has not been approved under previous iterations of the agreement, and is not covered under v1.2 of this agreement. Genomics England can confirm that if such linkage were to occur - then appropriate documentation (including transparency, patient, consent and ethics materials) would be provided to NHS Digital and a subsequent amended application put in place.
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England Trusted Research Environment to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and it’s outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel.
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By undertaking a prospective study design that leverages existing recruitment infrastructure in critical care, together with Genomics England, NHS England, Devolved Nations and Public Health England (PHE) infrastructure this research will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 whole genome sequences (WGS) this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
Genomics England Background: (For Context)
Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.
Data will be released into the Genomics England Trusted Research Environment where it will be linked with associated clinical data as well as additional data sources, which will include NHS Digital data requested through this agreement, COVID-19 testing feeds and viral genomics from Public Health England and data feeds from the Intensive Care National Audit Registry (ICNARC).
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable Research Environment for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new Research Environment (COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables effective COVID-19 research outcomes across a wide range of academic and Biotech/Pharma researchers with varying levels of technical competency. This offers a major upgrade to our current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling, and enable shared workspaces between Genomics England and Partners. COVID-RE must serve the immediate COVID-19 research effort and may also advance the transformation of Genomics England’s platform infrastructure, which is a key enabler to the research community.
There are two key value streams provided by the COVID-RE. Firstly ‘raw data to analytics ready data’ stream which must permit the collection of data from multiple locations from unstructured, through semi and fully structured forms and transforms them the appropriate data model, data store based on their ongoing use and availability. Secondly the ‘discovery to insight’ stream that supports researchers by providing the capability and the framework to support their User Journey from understanding the data available to them through to providing the necessary analytics and publishing tools.
The newly established COVID-RE will provide the following:
• Seamless interface with cloud storage and compute capabilities.
• A unified data platform containing datastores appropriate for all the required clinical and genomic data types.
• Data integration capability to deliver analytics-ready datasets into domain-specific data stores across batch and streaming integration patterns.
• Standards-based access services providing secure, fast, flexible, robust and auditable access to TRE data assets.
• Applications to support the research user journey from an exploration of data sets, through cohort building, analysis and publication - providing tools appropriate to a variety of user requirements (ways of working) and programming competency.
• Applications and workflows can be delivered natively within the platform, however COVID-RE must enable access to container-based applications and, via a set of hardened APIs, to a relevant external service (as long as security is maintained).
Details of the sublicence model via the Genomics England Trusted Research Environment are supplied within the processing activities section of this agreement.
The additional clinical data is key for researchers to be able to understand and infer clinically relevant and actionable findings. The aim is to provide a high quality, diverse clinical dataset, detailing each participant’s journey and to understand pre-existing conditions and also the early behaviours in the disease course. Genomics England currently have an agreement to receive NHS Digital data for the extant 100,000 Genomes Project participants (DARS-NIC-12784) and have seen first-hand the depth and quality of the data and how it has aided researchers.
To that aim, the significant gap in the current data collection for the COVID-19 project can be addressed by the non-standard NHS Digital data feed.
The Secondary Use Services Admitted Patient Care Feed (SUS APC) would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
The associated COVID-19 data-sets (including NHS111 and CV19 Testing Data in particular which will be requested under a future version of this agreement) will add to the detail in patient journey course, flagging for instance how participants were monitoring their symptoms and was there an associated poor prognosis.
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to approx 36,000 participants.
The data would be available for analysis alongside the extant Genomics England data set of 100,000 Genomes Project participants and would be made available to approved researchers worldwide as per existing governance procedures. This 100,000 Genomes Project data set will act as an appropriate matched control group. Data on the 100,000 Genomes Cohort will be provided to NHS Digital for linkage - (as they will act as a control cohort), as well as new additions for the GenOMICC Study.
The below detail pertains to the background of the study and susceptibility to infection sets out why Genomics are investigating. The origin of the GenOMICC study was focused on smaller cohorts of participants. However, in collaboration with the COG-UK group, this is now expanded to whole genome sequence of approx. 36,000 affected individuals as set out above.
GenOMICC Study Background:
Susceptibility to infection is profoundly heritable (Sorensen et al. 1988). Patients who develop life-threatening illness following infection with usually innocuous pathogens, such as influenza (Miller et al. 2010), are genetically different from the rest of the population (Albright et al. 2008). Understanding the genetic mechanisms of susceptibility may yield new therapeutic targets (Baillie 2014) that can be used to make susceptible patients more similar to individuals who are resistant to, or tolerant of, specific pathogens.
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV [Human Immunodeficiency Virus] (Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions (e.g. small molecules to inhibit an enzyme or receptor that is dysfunctional in resistant individuals) would therefore be protective to the host in a similar way to antibiotics, with the advantage that it is conceptually more difficult for any one pathogen to evolve resistance to such a therapy.
A second, more challenging problem arises in patients who become critically ill following infection. The patterns of immune-mediated organ dysfunction, immunoparesis, and death are very similar in severe infections and sterile systemic injuries (such as burns, haemorrhage, pancreatitis and trauma). Ultimately, death is a consequence of the host response to injury (Angus and Poll 2013), through final common pathways of organ failure that are clinically and biochemically evident, and unrelated to the original precipitant.
Broadly, the severity of critical illness follows directly from the severity and duration of the initial insult. In bacterial sepsis, early antibiotics are the mainstay of therapy; in influenza, early antivirals; in haemorrhage, early resuscitation; in trauma, urgent action to prevent secondary injury. There are no therapies with which to modulate the host response to systemic injury.
There is a lack of direct evidence of heritability for outcomes of critical illness, due in part to difficulties in defining and quantifying the heterogeneous multi-organ dysfunction syndrome (MODS), and in part due to the rapid pace of change in critical care medicine, making it impossible to tackle this question in long term outcome studies. However, clinical and biological evidence support the hypothesis that the pathogenesis of MODS is immune in origin (Angus and Poll 2013). Hence, predictions can be made from the extensive knowledge of other immune conditions. Whether or not MODS is considered to be an autoimmune or infectious condition is moot: these conditions share a great deal of similarity in genetic predispositions, cell types and mechanisms of pathogenesis. It is therefore very likely that propensity to survive MODS has a heritable component, and there is some direct evidence in support of this hypothesis (Rautanen et al. 2015). If this is the case, then the identity of the specific variants that contribute to outcome could potentially be utilised to design therapies to promote survival after the onset of MODS.
This study aims to identify genetic predisposition to specific syndromes of critical illness. Specifically, susceptibility to life-threatening infections caused by an identified pathogen, and susceptibility to death following the onset of organ failure due to sepsis or sterile injury. In order to maximise the probability of identifying host genetic loci associated with susceptibility, Genomics England will restrict some analyses to younger individuals in good general health and lacking in known predisposing factors.
The same principle was used to determine an upper age limit for inclusion for some analyses. With advancing age, there is an increase in undiagnosed co-morbidity, frailty, and susceptibility to serious complications of infection or critical injury. There is therefore an increase in the probability of susceptibility to, and mortality from, critical illness that is consequent upon non-genetic factors.
Participation Group:
Patients will be identified and recruited in hospital during acute illness. Potential participants will be identified through hospital workers upon presentation at recruiting sites. The disease processes under study have a high mortality, so it is desirable to recruit patients as early as possible in the disease process.
Participants or an appropriate parent/guardian/consultee will be approached by staff trained in consent procedures that protect the rights of the patient and adhere to the ethical principles within the Declaration of Helsinki. Staff will explain the details of the study to the participant or parent/guardian/consultee and allow them time to discuss and ask questions. The staff will review the informed consent form with the person giving consent (or assent) and endeavour to ensure understanding of the contents, including study procedures, risks, benefits, and the right to withdraw. Participants who agree to participate (or their parent/guardian or consultee who declares their wishes to do so) will be asked to sign and date an informed consent form.
In view of the importance of early sampling, participants or their parent/guardian/consultee will be permitted to consent and begin to participate in the study immediately if they wish to do so. Those who prefer more time to consider participation will be approached again after an agreed time, normally one day, to discuss further.
Patients who meet the inclusion/exclusion criteria and who have given informed consent to participate directly, or have been consented by a parent/guardian or whose wishes have been declared by a consultee, will be enrolled to the study.
Samples and data will be collected according to available resources and the weight of the patient will be measured for children under 12 in order to prevent excessive volume sampling. Samples required for medical management will at all times have priority over samples taken for research tests. Aliquots or samples for research purposes should never compromise the quality or quantity of samples required for medical management. Wherever practical, taking research samples should be timed to coincide with clinical sampling. The research team will be responsible for sharing the sampling protocol with health care workers supporting patient management in order to minimise disruption to routine care and avoid unnecessary procedures.
Consent will be sought from patients who survive critical illness and regain capacity to give consent. At each of the follow-up sessions, the investigator gathering data will determine whether the patient has regained capacity. In the event that a patient continues to be incapacitated beyond the follow-up period, the local investigator will plan a subsequent capacity check at a specific date, after an interval to be determined by the nature of the incapacity. The planned dates of capacity checks on incapacitated survivors will be stored locally in the site file, together with a record of the outcome of each check.
Patients who decline to participate at this stage will be removed from the study. Where the patient cannot be recontacted despite best endeavours, they will remain in the study.
Withdrawal:
Participants are freely able to decline participation in this study or to withdraw from participation at any point without suffering any implied or explicit disadvantage. All patients will be treated according to standard practice regardless of whether they participate.
The following options of withdrawal will be made available to participants:
1. Partial withdrawal. Data WILL continue to be updated and used for research, but no further contact will be made with the participant
2. Full withdrawal.
• no further contact will be made with the participant;
• data will not be updated from health records;
• data will not be removed from research that is underway or has already been done, and an audit record will be maintained to confirm participation.
Consent version and lifecourse follow-up
The first COVID positive patient was recruited to the GenOMICC study in March 2020. All patients (c.1500) currently recruited to the GenOMICC study are on consent and Protocol version 1.08 (which allows for COVID research but not longitudinal life course follow up). This Protocol v2.1 (submitted as part of this DARS request) was Research Ethics Committe (REC) approved on 23rd March 2020, IRAS [Integrated Research Application System] IDs are: 269326 & 189676 (https://www.hra.nhs.uk/covid-19-research/approved-covid-19-research/269326/). Genomics England will attempt to reconsent all patients recruited under version 1.08 onto the newly amended materials (version 2.1 – submitted as part of this DARS application) which allows for longitudinal lifecourse follow-up. For any patients on the v1.08 Protocol for which Genomics England are not able to seek consent, Genomics England will not be requesting their data for linkage. For patients prospectively recruited onto the v2.1 Protocol and consent materials, Genomics England will be requesting their data for linkage and follow-up .
Datasets requested from NHS Digital:
Consideration has been be given to assess and ensure that access to the data sets requested are inline with COVID-19 related purpose, in terms of the restrictions set out in Reg 3(1) COPI, below are details of what each of the data sets will at a high level provide insight to.
Primary Care dataset: Genomics England anticipate GPES (General Practice Extraction Service) Data for Pandemic Planning and Research (GDPPR) extracts to assist in completing the clinical image of a patient particularly around the identification of comorbidities and underlying medical conditions that were not captured in acute settings.
Hospital Episode Statistics (HES): Outpatients (OP), Admitted Patient Care (APC), Critical Care (CC) and Accident & Emergency (AE)/ Emergency Care Data Set (ECDS). These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset (DIDS). This provides invaluable, detailed information to build on participants' phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Secondary Uses Service datasets. The minimal latency in availability of these datasets is highly desirable for the research objectives set out in this project.
Mental Health Data sets: The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
Cancer Registration Data sets: To see the incidence of cancer within the cohort
Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
COVID datasets: These datasets, including SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus.
Assessment of the datasets has been undertaken and NHS Digital are satisfied that they are necessary for the COVID-19 work being undertaken. Genomics have confirmed that all research which is approved from the GenoMICC study using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Genomics England Industry access:
Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help turn research findings into treatments, diagnostics and benefits for patients as soon as possible.
As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations to access the COVID data, however they are charged based on storage and compute, such as running their own bioinformatics pipeline. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected.
The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a 'critical friend' and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's landmark data set.
The lawful basis for processing Participant Data under the General Data Protection Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then others will be added who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The lawful basis for the release and use of the confidential data being shared under this version of the agreement is Regulation 3(4) of the National Health Service (Control of Patient Information Regulations) 2002 (COPI) to require NHS Digital to share confidential patient information with organisations entitled to process this under COPI for COVID-19 purposes. The application of this has been based on the information provided in the Whole genome sequencing of patients severely affected by COVID-19 funding proposal from The GenOMICC - COVID Genomics UK (CoG-UK) partnership which was supported by the CMO of England and the CFO of the DHSC.
The only permitted activities under this Data Sharing Agreement (DSA) are for COVID-19 purposes and within bounds of Reg 3(2) COPI. Reg 3 (2) COPI states that: "2) For the purposes of this regulation, “processing” includes any operations, or set of operations set out in regulation 2(2) which are undertaken for the purposes set out in paragraph (1)." The research relates to the monitoring and managing of COVID-19 and would therefore be covered by Reg 3(1)(d) of COPI
Expected output
Genomics England completed sequencing 100,000 genomes at the end of 2018 (https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000- genomes-sequenced/). During 2018, the Genomics England Research Environment was established to allow research access to de-identified genomic and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with now over 3000 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England Trusted Research Environment where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw.
“This is a ground-breaking and far-reaching study which will harness the UK’s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
As of March 2020, the Genomics England Research Environment contained 107,694 genomes, of which, 33,461 were cancer and 74,233 were rare diseases.
The Research environment also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public HeaLth England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the Research Environment on the following dates, in order to continue support for, and to further develop, this ground-breaking resource:
• 3rd August 2020
• 7th September 2020
• 5th October 2020
• 2nd November 2020
Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants and GenOMICC participants into the Research Environment on the dates shown above.
DARS-NIC-374190-D0N1M-v0.4 21 July 2020 to 30 September 2020
- Title
- R26 - GENOMICS ENGLAND: GenOMICC COVID-19 Study
- Commercial
- Yes
- Sublicensing
- Yes
- Datasets
- 16
- Files released
- 265
Datasets: Bridge file: Hospital Episode Statistics to Diagnostic Imaging Dataset; Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); COVID-19 Hospitalization in England Surveillance System; COVID-19 SGSS First Positives (Second Generation Surveillance System); Demographics; Diagnostic Imaging Data Set (DID); Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Mental Health Services Data Set (MHSDS); Secondary Uses Service Payment By Results Spells
Objective for processing
This agreement is seeking approval to request data to support The GenOMICC - COVID Genomics UK (CoG-UK) partnership in researching Whole genome sequencing of patients severely affected by COVID-19. The work programme has sign off and prioritisation from the Chief Medical Officer for England
The goals of this work programme are set out below:
1. To harness world-leading UK healthcare and genomic infrastructure and systems to undertake prospective host whole genome sequencing at scale. This will elucidate the genetic architecture of host response to SARS-CoV-2 and identify opportunities to improve outcomes in the current pandemic, via international collaboration.
2. To identify rare and common variants that may affect susceptibility to response, identify novel opportunities for intervention and accelerate recovery.
3. To collect longitudinal life course datasets from primary care, hospital episodes, intensive care registries and outcomes via an extant partnership with NHS Digital and Health Data Research UK. We will include deep immune “omic” datasets on a subset of patients 9. This will allow case-control studies that capitalise upon unique UK assets, such as the 100,000 Genomes Project (97,000 people) and UK Biobank data sets (500,000 people) with WGS where more cases may be identified, including those with milder disease and unaffected people.
4. To use these rich data sets to understand the premorbid, concurrent and consequent sequelae of COVID-19 infection.
5. In partnership with the CoG-UK Viral Programme to evaluate the combination of viral and host genomics on outcomes to give pre-emptive insights into subsequent outbreaks and potentially future pandemics.
6. To provide access to these data sets via the Genomics England Trusted Research Environment to international and national academia and industry and facilitate international collaboration on COVID-19.
7. To link this to national COVID-19 clinical trials infrastructure offering potential for genomics to add value with insights into precision medicine and building a global-leading knowledgebase to enable better UK-wide and international capacity for future pandemic preparedness.
8. To engage and involve public and patients in setting strategy and priorities that shape the programme and it’s outputs. This will initially be based upon the 100,000 Genomes Project Participant Panel.
The prospective GenOMICC CoG-UK study
The variable response to COVID-19 suggests that, as with susceptibility to other infections, critical illness and mortality from COVID-19 may be determined by host genetic factors. From the 100,000 Genomes Project and the NIHR BioResource for Rare Disease it is known that rare variants cause immunodeficiency. By undertaking a prospective study design that leverages existing recruitment infrastructure in critical care, together with Genomics England, NHS England, Devolved Nations and PHE infrastructure this research will be able to apply the most advanced genomic testing to those most severely affected people admitted to hospital or intensive care.
The retrospective GenOMICC CoG-UK study:
The retrospective cohorts offer control arms for the study but also new case finding potential, particularly for people who have a milder clinical case. The study propose to harness the potential of two key national data assets. Firstly, analysis of the 100,000 Genomes Project data set which provides the genome sequences of 97,000 participants where they can use their longitudinal life course datasets to identify those affected by COVID-19, as well as providing appropriate unaffected controls. This dataset includes 627 families with rare immunodeficiency syndromes, which may allow insights to be accelerated because of co-existence of rare variants. Secondly, the UK Biobank cohort will provide 120,000 WGS this year, building to 500,000 WGS over the next 18 months from people, currently aged circa 55-85 years old, which is skewed towards the at-risk age groups for COVID-19 but will provide additional cases and controls, including mildly affected individuals.
Over the next five years these datasets will be enriched further by consented individuals from the NHS Genomic Medicine Service, and the new Genomics England programmes proposed in their strategic plan under consideration by Government. Furthermore, the Accelerating Detection of Disease Cohort will enrol up to 5 million people with genome-wide variants available with longitudinal life course data sets that could particularly add value to host response-variant associations and polygenic risk scores.
Genomics England Background:
Genomics England was established by the Department of Health to deliver the 100,000 Genomes Project. This followed the announcement in December 2012 by the Prime Minister of a programme of whole genome sequencing (WGS) as part of the UK Government’s Life Sciences Strategy. The principal objective of the 100,000 Genomes Project was to sequence 100,000 genomes from participants with cancer and rare disorders, and to link the sequence data to a standardised, extensible account of diagnosis, treatment, and outcomes gathered at recruitment, but primarily through the ongoing collection of medical records.
Data will be released into the Genomics England Trusted Research Environment where it will be linked with associated clinical data as well as additional data sources, which will include NHS Digital data requested through this agreement, COVID-19 testing feeds and viral genomics from Public Health England and data feeds from the Intensive Care National Audit Registry (ICNARC).
The Department of Health and Social Care have granted approval for Genomics England to procure a new, rapidly deployable Research Environment for the COVID-19 programme from existing core funding, which will provide a secure and collaborative workspace that enables researchers to perform COVID-19 genomic data analysis. The new Research Environment (COVID-RE) will provide an intuitive, integrated and collaborative user experience that enables effective COVID-19 research outcomes across a wide range of academic and Biotech/Pharma researchers with varying levels of technical competency. This offers a major upgrade to our current environment and will comprise user-centric, contemporary bioinformatic workflows, support opensource tooling, and enable shared workspaces between Genomics England and Partners. COVID-RE must serve the immediate COVID-19 research effort and may also advance the transformation of Genomics England’s platform infrastructure, which is a key enabler to the research community.
There are two key value streams provided by the COVID-RE. Firstly ‘raw data to analytics ready data’ stream which must permit the collection of data from multiple locations from unstructured, through semi and fully structured forms and transforms them the appropriate data model, data store based on their ongoing use and availability. Secondly the ‘discovery to insight’ stream that supports researchers by providing the capability and the framework to support their User Journey from understanding the data available to them through to providing the necessary analytics and publishing tools.
The newly established COVID-RE will provide the following:
• Seamless interface with cloud storage and compute capabilities.
• A unified data platform containing datastores appropriate for all the required clinical and genomic data types.
• Data integration capability to deliver analytics-ready datasets into domain-specific data stores across batch and streaming integration patterns.
• Standards-based access services providing secure, fast, flexible, robust and auditable access to TRE data assets.
• Applications to support the research user journey from an exploration of data sets, through cohort building, analysis and publication - providing tools appropriate to a variety of user requirements (ways of working) and programming competency.
• Applications and workflows can be delivered natively within the platform, however COVID-RE must enable access to container-based applications and, via a set of hardened APIs, to a relevant external service (as long as security is maintained).
Details of the sublicence model via the Genomics England Trusted Research Environment are supplied within the processing activities section of this agreement.
The additional clinical data is key for researchers to be able to understand and infer clinically relevant and actionable findings. The aim is to provide a high quality, diverse clinical dataset, detailing each participant’s journey and to understand pre-existing conditions and also the early behaviours in the disease course. Genomics England currently have an agreement to receive NHS Digital data for the extant 100,000 Genomes Project participants (DARS-NIC-12784) and have seen first-hand the depth and quality of the data and how it has aided researchers.
To that aim, the significant gap in the current data collection for the COVID-19 project can be addressed by the non-standard NHS Digital data feed.
The SUS APC feed would provide as close to real-time data for researchers and in the current climate is vital to ensure no delays.
The associated COVID-19 data-sets (including NHS111 and CV19 Testing Data in particular which will be requested under a future version of this agreement) will add to the detail in patient journey course, flagging for instance how participants were monitoring their symptoms and was there an associated poor prognosis.
The group of survivors eligible for recruitment for this study are generally healthy individuals who have suffered critical illness. It is anticipated this cohort will grow to approx 36,000 participants.
The data would be available for analysis alongside the extant Genomics England data set of 100,000 Genomes Project participants and would be made available to approved researchers worldwide as per existing governance procedures. This 100,000 Genomes Project data set will act as an appropriate matched control group. Data on the 100,000 Genomes Cohort will be provided to NHS Digital for linkage - (as they will act as a control cohort), as well as new additions for the GenOMICC Study.
The below detail pertains to the background of the study and susceptibility to infection sets out why Genomics are investigating. The origin of the GenOMICC study was focused on smaller cohorts of participants. However, in collaboration with the COG-UK group, this is now expanded to whole genome sequence of approx. 36,000 affected individuals as set out above.
GenOMICC Study Background:
Susceptibility to infection is profoundly heritable (Sorensen et al. 1988). Patients who develop life-threatening illness following infection with usually innocuous pathogens, such as influenza (Miller et al. 2010), are genetically different from the rest of the population (Albright et al. 2008). Understanding the genetic mechanisms of susceptibility may yield new therapeutic targets (Baillie 2014) that can be used to make susceptible patients more similar to individuals who are resistant to, or tolerant of, specific pathogens.
The genetic mechanisms of susceptibility to infection are likely to be highly pathogen-specific and may even have opposing roles in different infections (as for CCR5 variants in HIV (Huang et al. 1996) and WNV (Glass et al. 2006) infection). Pathogen-specific interventions (e.g. small molecules to inhibit an enzyme or receptor that is dysfunctional in resistant individuals) would therefore be protective to the host in a similar way to antibiotics, with the advantage that it is conceptually more difficult for any one pathogen to evolve resistance to such a therapy.
A second, more challenging problem arises in patients who become critically ill following infection. The patterns of immune-mediated organ dysfunction, immunoparesis, and death are very similar in severe infections and sterile systemic injuries (such as burns, haemorrhage, pancreatitis and trauma). Ultimately, death is a consequence of the host response to injury (Angus and Poll 2013), through final common pathways of organ failure that are clinically and biochemically evident, and unrelated to the original precipitant.
Broadly, the severity of critical illness follows directly from the severity and duration of the initial insult. In bacterial sepsis, early antibiotics are the mainstay of therapy; in influenza, early antivirals; in haemorrhage, early resuscitation; in trauma, urgent action to prevent secondary injury. There are no therapies with which to modulate the host response to systemic injury.
There is a lack of direct evidence of heritability for outcomes of critical illness, due in part to difficulties in defining and quantifying the heterogeneous multi-organ dysfunction syndrome (MODS), and in part due to the rapid pace of change in critical care medicine, making it impossible to tackle this question in long term outcome studies. However, clinical and biological evidence support the hypothesis that the pathogenesis of MODS is immune in origin (Angus and Poll 2013). Hence, predictions can be made from the extensive knowledge of other immune conditions. Whether or not MODS is considered to be an autoimmune or infectious condition is moot: these conditions share a great deal of similarity in genetic predispositions, cell types and mechanisms of pathogenesis. It is therefore very likely that propensity to survive MODS has a heritable component, and there is some direct evidence in support of this hypothesis (Rautanen et al. 2015). If this is the case, then the identity of the specific variants that contribute to outcome could potentially be utilised to design therapies to promote survival after the onset of MODS.
This study aims to identify genetic predisposition to specific syndromes of critical illness. Specifically, susceptibility to life-threatening infections caused by an identified pathogen, and susceptibility to death following the onset of organ failure due to sepsis or sterile injury. In order to maximise the probability of identifying host genetic loci associated with susceptibility, Genomics England will restrict some analyses to younger individuals in good general health and lacking in known predisposing factors.
The same principle was used to determine an upper age limit for inclusion for some analyses. With advancing age, there is an increase in undiagnosed co-morbidity, frailty, and susceptibility to serious complications of infection or critical injury. There is therefore an increase in the probability of susceptibility to, and mortality from, critical illness that is consequent upon non-genetic factors.
Withdrawal:
Participants are freely able to decline participation in this study or to withdraw from participation at any point without suffering any implied or explicit disadvantage. All patients will be treated according to standard practice regardless of whether they participate.
The following options of withdrawal will be made available to participants:
1. Partial withdrawal. Data WILL continue to be updated and used for research, but no further contact will be made with the participant
2. Full withdrawal.
• no further contact will be made with the participant;
• data will not be updated from health records;
• data will not be removed from research that is underway or has already been done, and an audit record will be maintained to confirm participation.
Consent version and lifecourse follow-up
The first COVID positive patient was recruited to the GenOMICC study in March 2020. All patients (c.1500) currently recruited to the GenOMICC study are on consent and Protocol version 1.08 (which allows for COVID research but not longitudinal life course follow up). This Protocol v2.1 (submitted as part of this DARS request) was REC approved on 23rd March 2020, IRAS IDs are: 269326 & 189676 (https://www.hra.nhs.uk/covid-19-research/approved-covid-19-research/269326/). Genomics England will attempt to reconsent all patients recruited under version 1.08 onto the newly amended materials (version 2.1 – submitted as part of this DARS application) which allows for longitudinal lifecourse follow-up. For any patients on the v1.08 Protocol for which Genomics England are not able to seek consent, Genomics England will not be requesting their data for linkage. For patients prospectively recruited onto the v2.1 Protocol and consent materials, Genomics England will be requesting their data for linkage and follow-up .
Datasets requested from NHS Digital:
Consideration has been be given to assess and ensure that access to the data sets requested are inline with COVID-19 related purpose, in terms of the restrictions set out in Reg 3(1) COPI, below are details of what each of the data sets will at a high level provide insight to.
Hospital Episode Statistics: Outpatients, Admitted Patient Care, Critical Care and Accident & Emergency/ ECDS. These datasets provide the core clinical data for participants and are vital to the provision of a detailed medical history for participants.
Diagnostic Imaging Dataset. This provides invaluable, detailed information to build on participants' phenotypes, e.g. tumour size and spread in cancer, adding to the understanding of patients' histories on individual and cohort level and their relationship with genomic alterations.
Secondary Uses Service datasets. The minimal latency in availability of these datasets is highly desirable for the research objectives set out in this project.
Mental Health Data sets: The 100,000 Genomes Project includes recruitment of psychiatric diseases and others with mental health phenotypes: intellectual disability and seizures are some of the most prevalent conditions within the Project. To-date nearly 10% of project participants have a mental health record. Mental health data are therefore vital in ensuring that a complete and relevant medical history is available for all participants.
Cancer Registration Data sets: To see the incidence of cancer within the cohort
Mortality data are essential for performing survival analyses and as a metric for success of medical care: this is crucial information for research in combination with other medical history. Cause of death information is vital in order to determine if mortality is related to the primary disease of a participant or to highlight unforeseen trends. Knowledge of participant death is also vital for the correct analysis of medical timeline data and for the management of participant cohorts.
COVID datasets: These datasets, including SGSS (Second Generation Surveillance System Data Set) and CHESS (COVID-19 Hospitalization in England Surveillance System) will be crucial to identify early Prognostic features in those affected with Coronavirus.
Assessment of the datasets has been undertaken and NHS Digital are satisfied that they are necessary for the COVID-19 work being undertaken. Genomics have confirmed that all research which is approved from the GenoMICC study using the data for the COVID-19 specific purposes will be published here https://www.genomicsengland.co.uk/about-gecip/research-2/
Genomics England Industry access:
Genomics England works with industry through its Discovery Forum. The Forum provides a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Industry partners comprise pharmaceutical, biotech and diagnostic companies, and those specialising in laboratory and data analysis. These companies have joined the Forum to work in a pre-competitive environment with access to a selection of genomic and associated clinical data. Ultimately, the Discovery Forum aims to help turn research findings into treatments, diagnostics and benefits for patients as soon as possible.
As the Discovery Forum is a collaborative venture, no fees are levied on participating organisations to access the COVID data, however they are charged based on storage and compute, such as running their own bioinformatics pipeline. All members of the Forum are obliged to publish all findings and research at the point at which intellectual property for any product is protected. Participants in the 100,000 Genomes Project have been asked explicitly to give consent for commercial companies to access their de-identified genome and health data.
The Forum was created in July 2017 and allows industrial partners to report back to Genomics England on what aspects of the data are proving to be most useful to their research studies, what data is missing and how the data should be collected and developed further so it is captures what industry needs, in a format that is compatible with their research and data systems. These partners act as a 'critical friend' and have made many helpful suggestions to increase the likelihood of successful research in the future for all those using Genomics England's landmark data set.
The lawful basis for processing Participant Data under the General Data Protection:
Regulation (GDPR) used by Genomics England is legitimate interests as set out under Article 6(1)(f) of the GDPR. It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
It is necessary for Genomics England to process Participant Data for its legitimate interests in carrying out medical research and in providing reports used by clinicians in their care of Participants.
The processing is necessary to support and enable Genomics England's legitimate interests in enabling new medical research on using genomics in health care, and on the causes, diagnosis and treatment of COVID-19.
Patients and the public will be at the heart of this programme. Initially the researchers will involve the extant 35 strong Genomics England Participant Panel and then we will add others who have been affected by COVID-19 at a later point. These participants and members of the public will be represented on all committees and working groups and will also meet separately.
The beneficiaries are:
o Participants - through the work Genomics England do will ultimately influence their care;
o researchers and industry - by giving them access to a unique ground-breaking resource of genomic data combined with life-course clinical data;
o and the wider public - by accelerating the uptake of genomic medicine making it available to patients in the UK.
The lawful basis for the release and use of the confidential data being shared under this version of the agreement is Regulation 3(4) of the National Health Service (Control of Patient Information Regulations) 2002 (COPI) to require NHS Digital to share confidential patient information with organisations entitled to process this under COPI for COVID-19 purposes. The application of this has been based on the information provided in the Whole genome sequencing of patients severely affected by COVID-19 funding proposal from The GenOMICC - COVID Genomics UK (CoG-UK) partnership which was supported by the CMO of England and the CFO of the DHSC.
Expected output
Genomics England completed sequencing 100,000 genomes at the end of 2018 (https://www.newscientist.com/article/2187499-uk-dna-project-hits-major-milestone-with-100000- genomes-sequenced/). During 2018, the Genomics England Research Environment was established to allow research access to de-identified genomic and clinical data received from NHS Digital. Thirty disease and cross-cutting GeCIP research domains were requested and approved, with now over 3000 GeCIP members given access to the Research Environment. Genomics England had also created the industry Discovery Forum to provide a platform for collaboration and engagement between Genomics England, industry partners, academia, the NHS and the wider UK genomics landscape.
Although the 100,000 genomes project has completed recruitment, Genomics England is committed to continue gathering life-long clinical data from the participants and making these available in the Research Environment.
Genomics England will be responsible for the onward workflow, in partnership with Illumina for the delivery of 30X whole genome sequences, subject to passing appropriate sequence QC, into the Genomics England data centre. Alignment and variant calling will be performed alongside the potential application of bespoke immunodeficiency panels as part of the Genomics England bioinformatics pipeline analysis.
Genomic data will be released into the Genomics England Trusted Research Environment where it will be linked with associated clinical data.
The GenOMICC study is backed by £28 million from Genomics England, UK Research and Innovation, the Department of Health and Social Care and the National Institute for Health Research. Illumina will sequence all 35,000 genomes and share some of the cost via an in-kind contribution.
A press release on 13/05/20 included a comment from Health and Social Care Secretary Matt Hancock: “As each day passes, we are learning more about this virus, and understanding how genetic makeup may influence how people react to it is a critical piece of the jigsaw.
“This is a ground-breaking and far-reaching study which will harness the UK’s world-leading genomics science to improve treatments and ultimately save lives across the world.” To date, nearly 3000 patients have been recruited into the project.
As of March 2020, the Genomics England Research Environment contained 107,694 genomes, of which, 33,461 were cancer and 74,233 were rare diseases.
The Research environment also contained clinical data on 89,157 participants (this is because cancer participants have two genomes submitted). The clinical data for 17,246 cancer participants includes clinical data from NHS Digital (HES OP/APC/ CC and AE) but also cancer specific data from Public HeaLth England Cancer Registry (NCRAS). The combination of clinical data for all 100,000 participants totals about 5m records.
As the GenOMICC study prospectively recruits participants, the aim will be to use the existing 100,000 participants and age and match-ranked controls for those entered into the study. As Genomics England prospectively enrolls more participants into the study, Genomics England plans further releases of genomic and clinical data, including clinical data received from NHS Digital and viral and host genomic data, into the Research Environment on the following dates, in order to continue support for, and to further develop, this ground-breaking resource:
• 3rd August 2020
• 7th September 2020
• 5th October 2020
• 2nd November 2020
Specific outputs over the period of this agreement are therefore to release updated genomic and clinical data for the 100,000 genomes participants and GenOMICC participants into the Research Environment on the dates shown above.
Benefits reported
Yielded Benefits is not a requirement for new applications.
Register history
When this agreement appeared in, or was edited in, each monthly edition of the register. Built by comparing every edition this site holds, the earliest of which is July 2021.
-
July 2021 —
already listed in the earliest edition this site holds, so it may be older. 3 versions: DARS-NIC-374190-D0N1M-v0.4, DARS-NIC-374190-D0N1M-v1.3, DARS-NIC-374190-D0N1M-v2.3
-
October 2021
1 version added: DARS-NIC-374190-D0N1M-v3.2
-
June 2022
1 version added: DARS-NIC-374190-D0N1M-v4.1
-
September 2022
Amended DARS-NIC-374190-D0N1M-v4.1
- Benefits reported:
reworded
Show the change
[13 paragraphs unchanged] In addition, two therapeutic drug trials have also commenced. By March 2022 some 16 new genetic variants associated with severe Covid-19, including some related to blood clotting, immune response and intensity of inflammation, have been identified. These findings will act as a roadmap for future efforts, opening new fields of research focused on potential new therapies and diagnostics with pinpoint accuracy. Determining the whole genome sequence for all participants in the study allowed the team to create a precise map and identify genetic variation linked to severity of Covid-19. The team found key differences in 16 genes in the ICU patients when compared with the DNA of the other groups. They also confirmed the involvement of seven other genetic variations already associated with severe Covid-19 discovered in earlier studies from the same team. The findings included how a single gene variant that disrupts a key messenger molecule in immune system signalling – called interferon alpha-10 – was enough to increase a patient’s risk of severe disease. This highlights the gene’s key role in the immune system and suggests that treating patients with interferon – proteins released by immune cells to defend against viruses – may help manage disease in the early stages. The study also found that variations in genes that control the levels of a central component of blood clotting – known as Factor 8 – were associated with critical illness in Covid-19. This may explain some of the clotting abnormalities that are seen in severe cases of Covid-19. Factor 8 is the gene underlying the most common type of haemophilia.
- Benefits reported:
reworded
-
November 2022
1 version added: DARS-NIC-374190-D0N1M-v5.5
-
January 2023
Amended DARS-NIC-374190-D0N1M-v0.4
- Datasets:
+ COVID-19 SGSS First Positives (Second Generation Surveillance System) ·
− COVID-19 Second Generation Surveillance System (SGSS)
Amended DARS-NIC-374190-D0N1M-v1.3- Datasets:
+ COVID-19 SGSS First Positives (Second Generation Surveillance System) ·
− COVID-19 Second Generation Surveillance System (SGSS)
Amended DARS-NIC-374190-D0N1M-v2.3- Datasets:
+ COVID-19 SGSS First Positives (Second Generation Surveillance System) ·
− COVID-19 Second Generation Surveillance System (SGSS)
Amended DARS-NIC-374190-D0N1M-v3.2- Datasets:
+ COVID-19 SGSS First Positives (Second Generation Surveillance System) ·
− COVID-19 Second Generation Surveillance System (SGSS)
Amended DARS-NIC-374190-D0N1M-v4.1- Datasets:
+ COVID-19 SGSS First Positives (Second Generation Surveillance System) ·
− COVID-19 Second Generation Surveillance System (SGSS)
Amended DARS-NIC-374190-D0N1M-v5.5- Datasets:
+ COVID-19 SGSS First Positives (Second Generation Surveillance System) ·
− COVID-19 Second Generation Surveillance System (SGSS)
- Datasets:
+ COVID-19 SGSS First Positives (Second Generation Surveillance System) ·
-
September 2023
Amended DARS-NIC-374190-D0N1M-v5.5
- Demographics: legal basis:
Consent (Reasonable Expectation); Health and Social Care Act 2012 – s261(2)(c)→ Not stated
- Demographics: legal basis:
-
March 2024
1 version added: DARS-NIC-374190-D0N1M-v6.4
-
November 2024
1 version added: DARS-NIC-374190-D0N1M-v7.2
-
February 2025
1 version added: DARS-NIC-374190-D0N1M-v8.2
-
February 2026
2 versions added: DARS-NIC-374190-D0N1M-v10.2, DARS-NIC-374190-D0N1M-v9.3
"Amended in place" means NHS England changed the record without issuing a new version number. The register publishes no changelog for those edits; this site infers them by comparing editions. An edit is attributed to the edition it first appears in, not to the date it was made.
Cite this page
NHS England (2026) Data Uses Register, September 2026 edition, agreement DARS-NIC-374190-D0N1M, “GenOMICC COVID-19 Study”. Read via NHS Data Access Explorer (unofficial), https://healthdatauses.uk/agreements/dars-nic-374190-d0n1m/ (accessed [date]).
This address stays the same, but the page is rebuilt with each monthly edition, so the citation names the edition it shows. Every edition's data is kept in the facts store.
Source: datausesregister_september2026.xlsx, September 2026 edition of the NHS England Data Uses Register. Search that workbook for DARS-NIC-374190-D0N1M to see the original rows.