Unofficial. This site is an experimental reformatting of data published by NHS England. It is not endorsed by NHS England. Always check the official Data Uses Register before relying on anything here.

Genes and Health

Barts and the London School of Medicine and Dentistry · Academic

Listed under Queen Mary University of London.

In term In term in the September 2026 edition: the latest version runs to 23 July 2028.

Reference
DARS-NIC-338864-B3Z3J
Current version
v6.2
Term of current version
24 February 2026 to 23 July 2028
Start date
9 July 2021
Data controller
Sole Data Controller
Commercial purposes
Yes
Sublicensing
Yes
Files released to date
617

Data controllers

Why the data was released

Objective for processing

The Data will be used for the purposes of the research project: Genes & Health (www.genesandhealth.org).

Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis. Genes & Health is currently used by over 140 groups of researchers worldwide, with an outstanding set of outputs and publications (https://www.genesandhealth.org/about-study/scientific-publications). Researchers access Genes & Health data via the Genes & Health Trusted Research Environment.

Genes & Health seeks to combine data from NHS England with existing data from genetics, other biological data from volunteer recall, health data from multiple other NHS health providers, and volunteer questionnaires data from the Genes & Health.

The strategic objective of Genes & Health is to develop and maintain a bioresource of genetic and health record data available to the worldwide research community (academic and life sciences industry partners) to improve the health of people of British Bangladeshi and British Pakistani origin people and their communities worldwide through high quality research. Non-white European ethnicities are profoundly underrepresented in large human genetic studies, biobanks, and clinical trials. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at Queen Mary University of London have built a Trusted Research Environment to investigate genetic data and health data from Genes & Health volunteers, which includes data from NHS England under this Data Sharing Agreement (DSA). This database will be used by researchers at Queen Mary University of London, as well as other researchers in other universities, academic institutions, and commercial life sciences industry companies.

Genes & Health is an open access resource, available to any bona fide scientific researchers worldwide who makes an application that is approved by the Genes Executive.

This database is eligible to be accessed by researchers at Queen Mary University of London (QMUL), as well as researchers at other external universities, academic institutions, and commercial life sciences industry companies

The Principal Investigator of the external organisation applying to access the Genes & Health resource, will initiate a request to access NHS England data for the purpose of a specific research project.

The organisation must sign a sublicence with QMUL relating to this Data Sharing Agreement (DSA) and must also sign the Genes & Health Data Access Agreement or the Data, Samples, Recall Agreement.

Access is granted to approved research projects and is limited to the Principal Investigator and employees (or those with honorary contracts) of the requesting organisation(s). An institutional email is required for each user to access the data, and this email must remain active for continued access. Using this email address, users are mandated to complete annual NHS Information Governance training and other access procedures. .

Genes & Health research database spans health and biomedical research, using human genetics/genomics/multiomics and longitudinal health care data to investigate any area of health or disease.

Research in four areas has been prioritised by local communities: 1) Diabetes 2) Cardiovascular Disease 3) Mental Health 4) Cancer, but Genes & Health do not work exclusively in these areas.

The Genes & Health Executive oversees all research activity and research project applications to use the data, operating under the Terms of Reference.

Applicants are asked to provide a lay summary, detailing project aims, how the research will be undertaken, and how the results will improve health for British Bangladeshi and British Pakistani, or other South Asian populations.

Any projects that do not have the primary purpose of researching and improving the health of British Bangladeshi, British Pakistani and other South Asian Populations are excluded. Projects without clear health related purpose (e.g. LLM processing of the entire dataset without specific aims) are also excluded.

Approved organisations accessing the Genes & Health database (including commercial organisations) will not be permitted to use the data for marketing purposes.

Although not forbidden, research in any potentially sensitive or difficult area is also subject to separate approval from the Genes & Health Community Advisory Board (e.g. relating to topics of mental health, sexual health, consanguinity).

Ultimate decision-making responsibility for approving requests is with the Genes & Health Executive. This performs Data Access Committee equivalent functions at its monthly meetings.

The Genes & Health Executive judges, through detailed review and discussion of each application (including of any formal peer-review e.g. from funders) assess whether the application meets the primary objective of the Genes & Health resource.

The Executive reviews all applications to use Genes & Health data to ensure that each one:

• Is submitted by a reputable institution and investigator

• Aligns with the goals of Genes & Health, namely improving health outcomes.

• Is sensitive to the needs and interests of participant communities (British Bangladeshi, British Pakistani and other South Asian Populations).

Where necessary Genes & Health Executive judges will prioritise research studies that meet any of the following criteria:

• Research described in core awards funding Genes & Health

• Research in four areas prioritised by our local communities: 1) Diabetes 2) Cardiovascular Disease 3) Mental Health 4) Cancer.

• Other research that is i) exceptional science; and/or ii) will bring substantial resource to Genes & Health; and/or iii) of strategic importance for Genes & Health

• Research that has been independently and expertly peer-reviewed to the highest standards (e.g. Wellcome, MRC, NIH)

The end-to-end process for receiving approval to access the Genes & Health resource is the following:

1. Applicant completes the Genes & Health application form

2. Genes & Health Executive review (monthly)

3. Genes & Health Community Advisory Board review if required

4. A Decision on Rejection, or request for further information/resubmit (in which case the application goes back to step 1). If a decision for approval is provided, proceed to step 5.

5. (a) Sign Genes & Health Data Access Agreement or Data, Samples, Recall Agreement and (b) sign a sublicence agreement with QMUL (if requesting NHS England data)

6. Pay QMUL admin and cloud compute fees.

7. Complete user Trusted Research Environment (TRE) onboarding including Information Governance training. Continue to complete annual Information Governance training.

8. Access granted to the TRE.

9. Data export controls apply within the TRE.

10. Publication in accordance with our Genes & Health authorship policy, available here https://www.genesandhealth.org/researchers/publications/

All Genes & Health approved research studies are recorded on the Genes & Health website. Furthermore, those with NHS England access, who have signed the sublicence, are listed on the website: These registers are updated at least every 3 months.

Genes & Health will charge commercial partners and academic partners for use of the bioresource, but this will cover infrastructure and sustainability costs only and will not be profit-making.

The following NHS England data is requested.

A full historic refresh for all NHS numbers each time an annual DSA is approved. This is because there will be new volunteers and new NHS numbers and their historic data is required. Not just a data refresh for the current year.

• Hospital Episode Statistics (HES) Admitted Patient Care – necessary because the data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. This dataset will support research into ill health and disease identified, diagnosed, and treated during hospital admissions, as well as episodes of pregnancy care.

• (HES) Accident & Emergency / Emergency Care Data Set (ECDS) – necessary because this dataset will support research into acute presentations of ill health, e.g., heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

• HES) Critical Care – necessary because the dataset will support research into severe ill health/disease, specifically related to critical care in this context. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g., in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity.

• (HES) Outpatients – necessary because the dataset will support research into ill health and disease, e.g., depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care.

• Mental Health Service Dataset (including its predecessors the Mental Health Minimum Data Set and Mental Health and Learning Disabilities Data Set) – necessary because these datasets will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multi morbidity, specifically to characterise states of mental ill health and disease, the care and treatment received.

• Civil Registration Mortality – necessary because this dataset will support research into ill health and disease by identifying death and cause of death, specifically to characterise states of ill health, disease, and their long-term outcomes, in this case death and cause of death.

• Cancer Registration – necessary because this will supplement existing sources and aid the analyses of the links between genetics and cancer.

• Demographics – necessary to mirror those collected independently by Genes & Health during follow-up contacts with volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

• National Diabetes Audit – necessary because this dataset will support extensive and high-quality research into diabetes and its related co-morbidities and multi morbidity. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment, and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

• Improving Access to Psychological Therapies (IAPT), Mental Health and Learning Disabilities and Community Services – necessary because these datasets are hoped to improve the completeness of data that is often under reported from other sources and allow a more accurate picture of community service use, particularly in areas like mental health.

• Maternity Services Data Set (MSDS) – necessary because this dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC. The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies.

• Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

• Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information.

Data is only accessed within the Trusted Research Environment. Download of individual level data is not permitted by researchers. Inference Control is applied to statistics involving small numbers of individuals as recommended by the UK Information Commissioners Office. No immediately identifiable Demographic Data (names, addresses, phone numbers, dates of birth, contact details) are available within the TRE. Month and year of birth; first part of postcode; geographic region (LSOA) may be available.

Data returned by NHSEngland is subsequently linked with other health data directly provided by NHS Primary Care providers, other data directly provided by NHS Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data.

The data will be minimised as follows:

• Limited to the cohort of ~73,000 volunteers with valid NHS numbers (as at time of renewal ~12/2025) who gave individual written informed consent to participate.

• Full data from earliest available to latest available on all volunteers.

Genes & Health volunteers have given written informed consent for their NHS health data to be accessed stored and processed in line with the study objectives. Volunteers also gave consent for linkage to their medical records, including from their GP and hospital records as well as from nationally held datasets such as those included in this agreement, for research purposes.

Genes & Health recruitment started in 2015 in East London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited approx. 73,000 volunteers (at December 2025) and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on East London, Luton, Bradford, West Midlands, and Greater Manchester). The cohort is expected to reach over 100,000 volunteers. Recruitment is ongoing.

Queen Mary University of London is the research Sponsor and the Data Controller as the organisation responsible for ensuring that the data will only be processed for the purpose described above.

The lawful basis for processing personal data under the UK GDPR is:

• Article 6(1)(e) - processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller.

The lawful basis for processing special category data under the UK GDPR is:

• Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

This processing is in the public interest as British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, and this resource looks to maximise its potential to deliver benefit to population health to these two groups.

The Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of QMUL. Wellcome Sanger Institute’s role is limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data.

Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data.

Genes & Health has developed an industry consortium with partners across a range of life sciences industry companies to support research that maximises the translation from basic science to therapeutics and clinical care. Consortium partners are Astra Zeneca, Pfizer, GSK, Merck, Novo Nordisk, Maze Theraputics, BMS, and Takeda. Consortium partners are working with the Genes & Health bioresource, under rigorous regulations set out in a Consortium Agreement. At present consortium partners are using genetic and non-NHS England health datasets. If partners use individual level raw NHS England data in the future, QMUL will require these partners to sign up to the sublicense agreement. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented.

QMUL research users access NHS England data under this DSA.

A Public and Patient Involvement and Engagement group helped refine the purpose of the research. The group strongly supported the collection of the data for the purposes described above. The Genes & Health Community Advisory Board meets 3-4 times a year and provides advice by email as required for other issues.

Volunteers are regularly invited to participate in workshops, focus groups, and informal information sessions. A newsletter is distributed 4x per year by Mailchimp email. These activities are designed to educate volunteers about health and genetics, and to gather public opinions on Genes & Health research.

Processing activities

Queen Mary University of London will transfer data to NHS England. The data will consist of research volunteer identifying details (specifically NHS Number, Date of Birth, Address, Postcode, First Name, Surname, Gender, and a unique person ID) for the cohort to be linked with NHS England data.

NHS England data will provide the relevant records from the datasets listed in Annex A Section 2 to Queen Mary University of London. NHS England provided data will contain no identifiable Demographic Data (name, address, date of birth, contact details and NHS number). NHS England provided data may contain month and year of birth, and first part of postcode.

Queen Mary University of London already has identifiable Demographic Data on research volunteers through their direct research participation. This will not be linked with NHS England Data.

Volunteer level health data from a variety of datasets from a variety of providers, including NHS England, along with the unique cohort study ID supplied by Genes & Health, will be transferred to Genes & Health Trusted Research Environment.

The data will be stored at Queen Mary University of London including Queen Mary University of London's Genes & Health Trusted Research Environment, using only servers and storage located in England (currently, Google Cloud Platform's eu-west-2 London datacentre).

Researchers from external institutions, including universities and commercial companies, will, with the agreement and oversight of QMUL, analyse Genes & Health data in a secure data environment: the Genes & Health Trusted Research Environment. This is administered by the Welcome Sanger Institute, as part of the Queen Mary University of London collaboration with Wellcome Sanger Institute on the Genes & Health programme.

Researchers using the TRE will only view and analyse data that is held in the London datacentre. No actual datafiles of health and genetic information will not be transferred to the users' compute systems in their country of use.

Only QMUL personnel will be allowed to download NHS England datafiles. No other organisations will have access to download NHS England datafiles held by QMUL.

Analysts and researchers from Wellcome Sanger Institute, Queen Mary University of London, and the Institutions of approved researchers via sublicensing will process/analyse the NHS England data for the purposes described above.

All personnel accessing the data have been appropriately trained in data protection and confidentiality. All personnel will have up to date NHS information governance training, annually refreshed.

The data will be linked at person record level with NHS Primary Care data, NHS Secondary Care data direct from Secondary Care providers, volunteer questionnaire data, volunteer survey data, volunteer genetic data, and other bio sample data obtained at volunteer recall.

There will be no requirement and no attempt to reidentify individuals by using NHS England provided data.

The identifying details will be stored in a separate Genes & Health database (which does not have NHS England Data) to the linked dataset used for analysis. All analyses will use the pseudonymised Trusted Research Environment dataset using the cohort unique person ID. There will be no requirement and no attempt to reidentify individuals when using the Trusted Research Environment dataset.

Expected output

The expected outputs of the processing will be:

• Submissions of scientific biomedical publications to peer reviewed journals – ongoing can be found under the following link https://www.genesandhealth.org/about-study/scientific-publications

• Presentations at appropriate conferences - national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications.

The outputs will not contain NHS England data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the dataset(s) from which the information was derived.

The outputs will be communicated to relevant recipients through the following dissemination channels:

- Journals

- Educational workshops and interactive web- and app involving the wider scientific community supported by the QMUL Life Sciences Initiative

- Conferences

- Social media: instagram, X, linkedin, bluesky.

- Public events the study team are guided by NIHR INVOLVE (https://www.invo.org.uk/) in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g., Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families.

- Patient Information leaflets available at www.genesandhealth.org

- Press/media engagement

- Participant newsletters

Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are design to educate participants about health and genetics, and to gather public opinions on Genes & Health research.

Expected targets for outputs are, (a) short-term, e.g., publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g., building and expanding the bioresource within 5 years and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years).

Outputs have already been generated (see Yielded Benefits, below) and further outputs are to be expected to be generated continuously until the end of the study.

Expected measurable benefits

The findings of this research study are expected to contribute to evidence-based decision-making for policymakers, local decision-makers such as doctors, and patients to inform best practice to improve the care, treatment and experience of health care users relevant to the subject matter of the study.

The use of the data could:

- Help the system to better understand the health and care needs of populations.

- Lead to the identification or improvement of treatments or interventions, or health and care system design to improve health and care outcomes or experience.

- Advance understanding of the need for, or effectiveness of, preventative health and care measures for particular populations or conditions such as obesity and diabetes.

- Provide a mechanism for checking the quality of care. This could include identifying areas of good practice to learn from, or areas of poorer practice which need to be addressed.

- Support knowledge creation or exploratory research (and the innovations and developments that might result from that exploratory work).

It is anticipated that the potential impact of this research will be broad, and cover a range of domains, including academic, health policy and economic impact, as well as direct impact on the individuals contributing to the research through public engagement. It is hoped dissemination will largely bring benefit through the acquisition of new knowledge relating to health and disease in a previously under-represented and understudied population. New interdisciplinary knowledge will be derived from the combined analysis of health and genomic data on a population scale. Detailed health data from NHS England should allow Genes & Health users to generate this knowledge for an otherwise under-studied British South Asian population who experience high rates of, and premature, disease. This new knowledge could be highly complementary to other major bioresources, e.g., UK Biobank and the Clinical Practice Research Datalink, and the availability of linked data curated across both could support high quality replication and validation studies thereby hopefully maximising outputs. The use of routine health data in this work could support direct and feasible translation of findings back to clinical care in the NHS. Additionally, the research could deliver new methodological insights related to the integration of health data science and genomics that could have direct impact to the academic community, including through the training of junior academic researchers.

Dissemination of new knowledge arising from the Genes & Health bioresource is likely to benefit health and social care through the development of new treatments and risk prediction strategies, as well as finding new causes of disease or its complications.

Longer-term, Genes & Health anticipate that the research will support policy-level improvements in health care with potential economic impact, e.g., through targeting prevention strategies to those at the highest risk, delivering more effective treatments, and reducing health inequalities.

Sharing new knowledge arising from Genes & Health is likely to benefit the public through raised awareness and understanding of health and disease and improving access and availability to effective treatment. These benefits will relate to British Bangladeshi and British Pakistanis and may be generalisable to other south Asian groups in the UK or globally. Genes & Health community partners and advisory group will have an active role in dissemination of results to ensure that they are effectively and sensitively disseminated to the communities the study represents.

It is hoped that research outputs arising from use of the Genes & Health bioresource will have direct relevance to a UK population of nearly 2 million British Bangladeshis and British Pakistanis, who are otherwise underrepresented in medical research.

Specific outputs from the Genes & Health multipurpose bioresource will be wide-ranging in terms of their impact. For example, the study of the impact of a rare gene variant on health and disease may have a direct impact on only a small number of individuals and their families, but the magnitude of its impact could be large if, for example, it allows the development of a new treatment for a severe disease (as exemplified by previous Genes & Health work on the HAO1 gene variant). Conversely, the use of health and genetic data to build improved disease prediction models for conditions such as type 2 diabetes might have long-term benefits on a large proportion of the population as approximately 14% of British south Asians are thought to have the condition.

Genes & Health anticipate short, and medium-long term benefits to arise and be disseminated. It is hoped that short-term benefits will include outputs from research currently underway on disease prediction using polygenic risk scores and is likely to include new insights into disease associations with rare gene variants, related to type 2 diabetes, mental health and cardiovascular disease. Medium-long-term dissemination will include similar work across an expanded range of disease areas, as well as a wider portfolio of collaborative research with partners. The current research funders will not derive exploitable benefits from the Genes & Health research outputs.

It is hoped that through publication of findings in appropriate media, the findings of this research will add to the body of evidence that is considered by the bodies, organisations and individual care practitioners charged with making policy decisions for or within the NHS or treatment decisions in relation to specific patients.

Genes & Health have recently employed a Communications and Engagement Manger, who will ensure that, in addition to research and NHS groups, local government and third sector organisations are involved in the development and dissemination of the research programme.

Benefits reported so far

Genes and Health first received NHS England data in Dec 2021. Data has been integrated with other health care datasets (e.g. direct from local primary care and secondary care NHS Providers), and with genetic data and volunteer recall data. The genetic data (GSA chip genotyping, exome sequencing) has only been released in the last few years and laboratory work is ongoing.

Genes & Health has attracted 116 research applications from scientists worldwide to use the resource, in addition to 8 international life sciences companies. The Genes & Health Trusted Research Environment has 239 active users.

Genes & Health has published 44 scientific publications (https://www.genesandhealth.org/about-study/scientific-publications) , many of the very highest international excellence including 10 in Nature, 2 in Nature Medicine, 1 in Nature Genetics, 1 in Science, 1 in New England Journal of Medicine, and 2 in Cell. Many more publications will follow over the next 5 years as the exome sequencing data is fully explored.

Measurable benefits yielded include:

- 50 volunteers diagnosed with LDLR gene Familial Hypercholesterolemia, an inherited condition with very high risk of early onset heart attack. No volunteers knew their diagnosis, and almost all were either not or were undertreated, despite effective treatments being available. All have now been treated and referred to NHS Cardiology clinics for long term management and prevention. See press release https://www.qmul.ac.uk/blizard/research/featured-research/look-at-those-genes--gene-research-offers-better-health-outcomes-for-british-bangladeshi-and-pakistani-people/?ocid=BingNews%2CBingNews

- 1 of the volunteers with Familial Hypercholesterolemia took part in a novel phase 1 trial aimed at a single dose (liver gene editing) treatment to treat once and for all lifelong. The participant met the inclusion criteria due to the high quality research data from Genes & Health. Press release: https://www.qmul.ac.uk/media/news/2023/smd/a-participant-from-queen-mary-university-of-london-genes--health-study-is-the-10th-person-enrolled-in-a-gene-editing-clinical-trial-for-heart-disease-.html

- A recent discovery that ~7% of south asians in Genes & Health carry a south asian-specific genetic variant that affects the accuracy of the HbA1C test for diabetes, and leads to late diagnosis and undertreatment. Press release: https://www.qmul.ac.uk/media/news/2024/fmd/standard-diabetes-test-may-be-inaccurate-for-10000s-of-south-asian-people-in-uk.html and https://www.diabetes.org.uk/about-us/news-and-views/type-2-diabetes-test-may-be-inaccurate-thousands-south-asian-people

- The discovery of a new pathway for growth, body composition and the onset of puberty (Nature 2021, https://doi.org/10.1038/s41586-021-04088-9) through recall of a Genes & Health volunteer lacking normal melanocortin receptor MC3R function.

- The discovery of human variants influencing COVID-19 hospitalisation and severity that are more frequent in south Asian versus white European populations and explain a part of the poorer outcomes in these ethnic groups (Nature 2021, https://doi.org/10.1038/s41586-021-03767-x).

- The study of a Genes & Health volunteer lacking the HAO1 gene, leading to a new drug treatment lumasiran (targeting HAO1) for the rare disease paediatric hyperoxaluria. Genes & Health data was included in the successful FDA New Drug Application and subsequent regulatory approval (eLife 2020, https://doi.org/10.7554/eLife.54363).

- Discovery that south asians (in Genes & Health) have a high prevalence of genetic variants influencing clopidogrel metabolism (and hence effectiveness). The drug clopidogrel is used to treat heart attacks, and volunteers with the variant had a higher rate of recurrent heart attack when treated with clopidogrel ( https://doi.org/10.1016/j.jacadv.2023.100573). This is likely to change prescribing guidelines in these ethnic groups.

- Attracting £30m of international life sciences industry investment to the UK for Genes & Health Industry Consortium 1.

Datasets on the current version

Legal basis for provision: Health and Social Care Act 2012 – s261(2)(c)

Datasets approved under DARS-NIC-338864-B3Z3J-v6.2
DatasetType of dataSensitivity FrequencyConfidential data
Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Cancer Registration Data Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Civil Registrations of Death Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Community Services Data Set (CSDS) Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Demographics Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Emergency Care Data Set (ECDS) Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Hospital Episode Statistics Accident and Emergency (HES A and E) Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Hospital Episode Statistics Admitted Patient Care (HES APC) Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Hospital Episode Statistics Critical Care (HES Critical Care) Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Hospital Episode Statistics Outpatients (HES OP) Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Improving Access to Psychological Therapies (IAPT) v1.5 Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Improving Access to Psychological Therapies (IAPT) v2 Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)
Maternity Services Data Set (MSDS) v1.5 Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Mental Health and Learning Disabilities Data Set (MHLDDS) Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Mental Health Minimum Data Set (MHMDS) Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Mental Health Services Data Set (MHSDS) Identifiable Sensitive One-Off Consent (Reasonable Expectation)
National Diabetes Audit Identifiable Sensitive One-Off Consent (Reasonable Expectation)
Patient Reported Outcome Measures (Linkable to HES) Identifiable Non-Sensitive One-Off Consent (Reasonable Expectation)

Files released

Files released counts only files released externally by DARS. Access granted in NHS England's own systems, such as its Secure Data Environment, is not included.

This agreement permits sublicensing: the applicant may pass data on to others. Anything passed on is not recorded in this register.

Patient opt-outs were not applied to any of the 617 files released under this agreement, across every version. About opt-outs

No files recorded as released under the current version. 617 were released under earlier versions, shown in the version history.

Version history

The register lists each renewal of this agreement as a separate row. This site has 7 versions.

DARS-NIC-338864-B3Z3J-v6.2 24 February 2026 to 23 July 2028
Title
Genes and Health
Commercial
Yes
Sublicensing
Yes
Datasets
18
Files released
0

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); Demographics; Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Improving Access to Psychological Therapies (IAPT) v1.5; Improving Access to Psychological Therapies (IAPT) v2; Maternity Services Data Set (MSDS) v1.5; Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); National Diabetes Audit; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-338864-B3Z3J-v5.4

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-338864-B3Z3J-v5.4
FieldWasBecame
Start date2025-03-192026-02-24
End date2028-07-182028-07-23

Objective for processing

Queen Mary University of London (QMUL) requests access to NHS England data The Data will be used for the purpose purposes of the following research project: Genes & Health (www.genesandhealth.org). Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis. Genes & Health is currently used by over 120 140 groups of researchers worldwide, with an outstanding set of outputs and publications (https://www.genesandhealth.org/about-study/scientific-publications). Researchers access Genes & Health data via the Genes & Health Trusted Research Environment. [2 paragraphs unchanged] Genes & Health is an open access resource, available to any bona fide scientific researchers worldwide who makes an application that is approved by the Genes Executive. This database is eligible to be accessed by researchers at Queen Mary University of London (QMUL), as well as researchers at other external universities, academic institutions, and commercial life sciences industry companies The Principal Investigator of the external organisation applying to access the Genes & Health resource, will initiate a request to access NHS England data for the purpose of a specific research project. The organisation must sign a sublicence with QMUL relating to this Data Sharing Agreement (DSA) and must also sign the Genes & Health Data Access Agreement or the Data, Samples, Recall Agreement. Access is granted to approved research projects and is limited to the Principal Investigator and employees (or those with honorary contracts) of the requesting organisation(s). An institutional email is required for each user to access the data, and this email must remain active for continued access. Using this email address, users are mandated to complete annual NHS Information Governance training and other access procedures. . Genes & Health research database spans health and biomedical research, using human genetics/genomics/multiomics and longitudinal health care data to investigate any area of health or disease. Research in four areas has been prioritised by local communities: 1) Diabetes 2) Cardiovascular Disease 3) Mental Health 4) Cancer, but Genes & Health do not work exclusively in these areas. The Genes & Health Executive oversees all research activity and research project applications to use the data, operating under the Terms of Reference. Applicants are asked to provide a lay summary, detailing project aims, how the research will be undertaken, and how the results will improve health for British Bangladeshi and British Pakistani, or other South Asian populations. Any projects that do not have the primary purpose of researching and improving the health of British Bangladeshi, British Pakistani and other South Asian Populations are excluded. Projects without clear health related purpose (e.g. LLM processing of the entire dataset without specific aims) are also excluded. Approved organisations accessing the Genes & Health database (including commercial organisations) will not be permitted to use the data for marketing purposes. Although not forbidden, research in any potentially sensitive or difficult area is also subject to separate approval from the Genes & Health Community Advisory Board (e.g. relating to topics of mental health, sexual health, consanguinity). Ultimate decision-making responsibility for approving requests is with the Genes & Health Executive. This performs Data Access Committee equivalent functions at its monthly meetings. The Genes & Health Executive judges, through detailed review and discussion of each application (including of any formal peer-review e.g. from funders) assess whether the application meets the primary objective of the Genes & Health resource. The Executive reviews all applications to use Genes & Health data to ensure that each one: • Is submitted by a reputable institution and investigator • Aligns with the goals of Genes & Health, namely improving health outcomes. • Is sensitive to the needs and interests of participant communities (British Bangladeshi, British Pakistani and other South Asian Populations). Where necessary Genes & Health Executive judges will prioritise research studies that meet any of the following criteria: • Research described in core awards funding Genes & Health • Research in four areas prioritised by our local communities: 1) Diabetes 2) Cardiovascular Disease 3) Mental Health 4) Cancer. • Other research that is i) exceptional science; and/or ii) will bring substantial resource to Genes & Health; and/or iii) of strategic importance for Genes & Health • Research that has been independently and expertly peer-reviewed to the highest standards (e.g. Wellcome, MRC, NIH) The end-to-end process for receiving approval to access the Genes & Health resource is the following: 1. Applicant completes the Genes & Health application form 2. Genes & Health Executive review (monthly) 3. Genes & Health Community Advisory Board review if required 4. A Decision on Rejection, or request for further information/resubmit (in which case the application goes back to step 1). If a decision for approval is provided, proceed to step 5. 5. (a) Sign Genes & Health Data Access Agreement or Data, Samples, Recall Agreement and (b) sign a sublicence agreement with QMUL (if requesting NHS England data) 6. Pay QMUL admin and cloud compute fees. 7. Complete user Trusted Research Environment (TRE) onboarding including Information Governance training. Continue to complete annual Information Governance training. 8. Access granted to the TRE. 9. Data export controls apply within the TRE. 10. Publication in accordance with our Genes & Health authorship policy, available here https://www.genesandhealth.org/researchers/publications/ All Genes & Health approved research studies are recorded on the Genes & Health website. Furthermore, those with NHS England access, who have signed the sublicence, are listed on the website: These registers are updated at least every 3 months. [19 paragraphs unchanged] • Limited to the cohort of ~70,000 ~73,000 volunteers with valid NHS numbers (as at time of renewal ~7/2024) ~12/2025) who gave individual written informed consent to participate. [2 paragraphs unchanged] Genes & Health recruitment started in 2015 in East London, where limited [5 words unchanged] through linkage to local health systems. Genes & Health has recruited approx. 70,000 73,000 volunteers (at March 2024) December 2025) and is expanding nationally with recruitment taking place across multiple geographical regions [15 words unchanged] The cohort is expected to reach over 100,000 volunteers. Recruitment is ongoing. [5 paragraphs unchanged] Genes & Health also has written individual consent, which provides an additional English Common Law lawful basis. [1 paragraph unchanged] The funding comes from multiple sources. Substantial funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for England Catalyst, Barts Charity, Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames), Alnylam Pharmaceuticals, Genomics PLC, EU Horizon 2020, National Institutes of Health (USA); and a Life Sciences Industry Consortium of Astra Zeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. Funding is in place until 2030. Funding to continue the work described will be sought on an ongoing basis. The Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of QMUL. Wellcome Sanger Institute’s role is limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data. The Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of Queen Mary University of London. Wellcome Sanger Institute’s role is limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data. [2 paragraphs unchanged] QMUL research users also access NHS England data under this DSA. [1 paragraph unchanged] Volunteers are regularly invited to participate in workshops, focus groups, and informal information sessions. A newsletter is distributed 2-3x 4x per year by Mailchimp email. These activities are designed to educate volunteers about health and genetics, and to gather public opinions on Genes & Health research.

Processing activities

[5 paragraphs unchanged] Researchers from external institutions, including universities and commercial companies, will, with the [38 words unchanged] London collaboration with Wellcome Sanger Institute on the Genes & Health programme. The Wellcome Sanger Institute holds Data Security and Protection Toolkits, and full ISO27001 certification for the Genes & Health Trusted Research Environment. The Genes & Health Trusted Research Environment (TRE) is a secure platform that has been developed by QMUL and Welcome Sanger Institute specifically for Genes & Health, based on the successful platform developed by the similar Finngen consortium (https://www.finngen.fi/en). Full ISO27001 Certification was achieved in March 2023. IT administration is provided by Genes & Health funded staff at the Wellcome Sanger Institute. Researchers using the TRE will only view and analyse data that is held in the London datacentre. Actual No actual datafiles of health and genetic information will not be transferred to the users' compute systems in their country of use. [1 paragraph unchanged] Access to NHS England raw data will be restricted to users in the Territory of Use as described in 2c. The institutions of non-QMUL users will be required to sign a sublicence. Access to the Data will not be allowed to users in UK/EU/US sanctioned countries. Genes & Health is an open access resource, available to any bona fide scientific researcher worldwide who makes an application that is approved by the Genes Executive. Sensitive applications (e.g. mental health, sexual health, consanguinity) are subject to further scrutiny by the Genes & Health Community Advisory Board. Access to the TRE is subject to payment (to cover costs only), Data Access Agreement execution, and Information Governance training. QMUL may grant sublicenses to individuals from academic or industrial/commercial entities. These may include pharmaceutical and genomics companies for purposes that include but are not limited to supporting research that maximises the translation from basic science to therapeutics and clinical care. Industrial partners will be invited to work with the Genes & Health bioresource, under rigorous regulations set out in a Data Sharing Agreement and Consortium Agreement. Where these partners use NHS England datasets, QMUL will require these partners to sign up to the sublicense agreements. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented. A Data Access Agreement (DAA) will be signed by external institutions before access to Genes & Health data is granted. A sublicense will also be signed for access to NHS England data if access to NHS England Data or NHS England Manipulated Data is required. Users and institutions may choose not to have access to NHS England Data or NHS England Manipulated Data, but choose to have access to other Genes & Health data. Research users are legally prohibited, and not technically capable (via Identity and Access Management permissions), of downloading or copying data from the Trusted Research Environment to local or other devices. The data will always remain in England, on the systems Data Controlled by QMUL, and will not be downloadable. [5 paragraphs unchanged]

Unchanged: Expected output, Expected measurable benefits, Benefits reported.

DARS-NIC-338864-B3Z3J-v5.4 19 March 2025 to 18 July 2028
Title
Genes and Health
Commercial
Yes
Sublicensing
Yes
Datasets
18
Files released
84

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); Demographics; Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Improving Access to Psychological Therapies (IAPT) v1.5; Improving Access to Psychological Therapies (IAPT) v2; Maternity Services Data Set (MSDS) v1.5; Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); National Diabetes Audit; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-338864-B3Z3J-v4.2

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-338864-B3Z3J-v4.2
FieldWasBecame
Start date2024-07-192025-03-19
End date2025-07-182028-07-18

Objective for processing

[1 paragraph unchanged] The following is a summary of the aims of the research project provided by Genes & Health on behalf of Queen Mary University of London. [1 paragraph unchanged] Genes & Health seeks to combine data from NHS England with existing [14 words unchanged] NHS health providers, and volunteer questionnaires data from the Genes & Health. Access to NHS England data would be via sub-licensing arrangements. The strategic objective of Genes & Health is to develop and maintain a bioresource of genetic and health record data available to the worldwide research community (academic and industrial life sciences industry partners) to improve the health of people of British Bangladeshi and British [85 words unchanged] Health volunteers, which includes data from NHS England under this Data Sharing Agreement. Agreement (DSA). This database will be used by researchers at Queen Mary University of [5 words unchanged] researchers in other universities, academic institutions, and commercial life sciences industry companies. [1 paragraph unchanged] The following NHS England data will be accessed: is requested. A full historic refresh for all NHS numbers each time an annual DSA is approved. This is because there will be new volunteers and new NHS numbers and their historic data is required. Not just a data refresh for the current year. [13 paragraphs unchanged] Data is only accessed within the Trusted Research Environment. Download of individual [41 words unchanged] available within the TRE. Month and year of birth; first part of postcode postcode; geographic region (LSOA) may be available. Genetic data is available which makes the dataset Personal Data under GDPR. Data returned by NHSEngland is subsequently linked with other health data directly provided by NHS Primary Care providers, other data directly provided by NHS Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data. [1 paragraph unchanged] • Limited to the cohort of 70,000 ~70,000 volunteers with valid NHS numbers (as at time of renewal ~7/2024) who gave individual written informed consent to participate. • Limited to data between 1997/98 to latest available. • Full data from earliest available to latest available on all volunteers. [1 paragraph unchanged] Genes & Health recruitment started in 2015 in East London, where limited [5 words unchanged] through linkage to local health systems. Genes & Health has recruited approx. 61,000 70,000 volunteers (at March 2024) and is expanding nationally with recruitment taking place [13 words unchanged] Bradford, West Midlands, and Greater Manchester). The cohort is expected to reach over 100,000 volunteers. Recruitment is ongoing. [5 paragraphs unchanged] Genes & Health also has written individual consent, which additionally provides a common law an additional English Common Law lawful basis. [1 paragraph unchanged] The funding comes from multiple sources. Substqntial Substantial funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for [70 words unchanged] Pfizer Inc, Takeda Development Centre Americas Inc. Funding is in place until 2028. 2030. Funding to continue the work described will be sought on an ongoing basis. [1 paragraph unchanged] Several organisations have provided data or services but are not processors of the NHS England data. These are Kings College London, UK Biocentre, Broad Institute (USA) who have processed DNA samples from volunteers. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership. Genes & Health funders are not involved in the day to day [42 words unchanged] Genes & Health project and do not have access to the data. Genes & Health has developed an industry consortium with partners across a [60 words unchanged] using genetic and non-NHS England health datasets. If partners use individual level raw NHS England data in the future, QMUL will require these partners to sign up to the sublicense agreements. agreement. These agreements and procedures will ensure that research using the data in [20 words unchanged] and UK GDPR compliance are upheld, and that commercial exploitation is prevented. QMUL users also access NHS England data under this DSA. [2 paragraphs unchanged]

Processing activities

Queen Mary University of London will transfer data to NHS England. The data will consist of research volunteer identifying details (specifically NHS Number, Date of Birth, Address, Postcode, First Name, Surname, Gender, and a unique person ID) for the cohort to be linked with NHS England data. NHS England data will provide the relevant records from the HES, ECDS, Cancer Registration, Civil Registration of Death, Demographics, Maternity Services, Mental Health and Learning Disabilities, Mental Health Minimum, Community Services, Improving Access to Psychological Therapies, Patient Reported Outcome Measures and National Diabetes Audit datasets listed in Annex A Section 2 to Queen Mary University of London. NHS England provided data will contain [18 words unchanged] may contain month and year of birth, and first part of postcode. [2 paragraphs unchanged] The data will be stored at Queen Mary University of London including Queen Mary University of London's (as Data Controller) Genes & Health Trusted Research Environment, using only servers and storage located in England (currently, Google Cloud Platform's eu-west-2 London datacentre). Researchers from external institutions, including universities and commercial companies, will, with the [113 words unchanged] Certification was achieved in March 2023. IT administration is provided by Genes and & Health funded staff at the Wellcome Sanger Institute. Genes & Health is an open access resource, available to any bona fide scientific researcher worldwide who makes an application that is approved by the Genes Executive. Sensitive applications (e.g. mental health, sexual health, consanguinity) are subject to further scrunity by the Genes & Health Community Advisory Board. Access to the TRE is subject to payment (to cover costs only), Data Access Agreement execution, and Information Governance training. Researchers using the TRE will only view and analyse data that is held in the London datacentre. Actual datafiles of health and genetic will not be transferred to the users' compute systems in their country of use. Only QMUL personnel will be allowed to download NHS England datafiles. No other organisations will have access to download NHS England datafiles held by QMUL. Access to NHS England raw data will be restricted to users in the Territory of Use as described in 2c. The institutions of non-QMUL users will be required to sign a sublicence. Access to the Data will not be allowed to users in UK/EU/US sanctioned countries. Genes & Health is an open access resource, available to any bona fide scientific researcher worldwide who makes an application that is approved by the Genes Executive. Sensitive applications (e.g. mental health, sexual health, consanguinity) are subject to further scrutiny by the Genes & Health Community Advisory Board. Access to the TRE is subject to payment (to cover costs only), Data Access Agreement execution, and Information Governance training. QMUL may grant sublicenses to individuals from academic or industrial/commercial entities. These may include pharmaceutical and genomics companies for purposes that include but are not limited to supporting research that maximises the translation from basic science to therapeutics and clinical care. Industrial partners will be invited to work with the Genes & Health bioresource, under rigorous regulations set out in a Data Sharing Agreement and Consortium Agreement. Where these partners use NHS England datasets, QMUL will require these partners to sign up to the sublicense agreements. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented. [8 paragraphs unchanged] Analysts and researchers from Wellcome Sanger Institute, Queen Mary University of London and the Institutions of approved researchers via sublicensing will process/analyse the data for the purposes described above.

Expected output

[7 paragraphs unchanged] - Webinars open to members of the Queens Mary University of London. - Conferences - Social media through twitter found at (@eastlondongenes and @bradfordgenes, and @manchestergenes). - Social media: instagram, X, linkedin, bluesky. [7 paragraphs unchanged]

Benefits reported

[11 paragraphs unchanged] - Attracting £30m of international life sciences industry investment to the UK for Genes & Health Industry Consortium 1. We are now pitching for Consortium 2 which we hope will bring a further £50m and enhance the UK's reputation as a life sciences powerhouse.

Unchanged: Expected measurable benefits.

Objective for processing

Queen Mary University of London (QMUL) requests access to NHS England data for the purpose of the following research project: Genes & Health (www.genesandhealth.org).

Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis. Genes & Health is currently used by over 120 groups of researchers worldwide, with an outstanding set of outputs and publications (https://www.genesandhealth.org/about-study/scientific-publications). Researchers access Genes & Health data via the Genes & Health Trusted Research Environment.

Genes & Health seeks to combine data from NHS England with existing data from genetics, other biological data from volunteer recall, health data from multiple other NHS health providers, and volunteer questionnaires data from the Genes & Health.

The strategic objective of Genes & Health is to develop and maintain a bioresource of genetic and health record data available to the worldwide research community (academic and life sciences industry partners) to improve the health of people of British Bangladeshi and British Pakistani origin people and their communities worldwide through high quality research. Non-white European ethnicities are profoundly underrepresented in large human genetic studies, biobanks, and clinical trials. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at Queen Mary University of London have built a Trusted Research Environment to investigate genetic data and health data from Genes & Health volunteers, which includes data from NHS England under this Data Sharing Agreement (DSA). This database will be used by researchers at Queen Mary University of London, as well as other researchers in other universities, academic institutions, and commercial life sciences industry companies.

Genes & Health will charge commercial partners and academic partners for use of the bioresource, but this will cover infrastructure and sustainability costs only and will not be profit-making.

The following NHS England data is requested.

A full historic refresh for all NHS numbers each time an annual DSA is approved. This is because there will be new volunteers and new NHS numbers and their historic data is required. Not just a data refresh for the current year.

• Hospital Episode Statistics (HES) Admitted Patient Care – necessary because the data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. This dataset will support research into ill health and disease identified, diagnosed, and treated during hospital admissions, as well as episodes of pregnancy care.

• (HES) Accident & Emergency / Emergency Care Data Set (ECDS) – necessary because this dataset will support research into acute presentations of ill health, e.g., heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

• HES) Critical Care – necessary because the dataset will support research into severe ill health/disease, specifically related to critical care in this context. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g., in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity.

• (HES) Outpatients – necessary because the dataset will support research into ill health and disease, e.g., depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care.

• Mental Health Service Dataset (including its predecessors the Mental Health Minimum Data Set and Mental Health and Learning Disabilities Data Set) – necessary because these datasets will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multi morbidity, specifically to characterise states of mental ill health and disease, the care and treatment received.

• Civil Registration Mortality – necessary because this dataset will support research into ill health and disease by identifying death and cause of death, specifically to characterise states of ill health, disease, and their long-term outcomes, in this case death and cause of death.

• Cancer Registration – necessary because this will supplement existing sources and aid the analyses of the links between genetics and cancer.

• Demographics – necessary to mirror those collected independently by Genes & Health during follow-up contacts with volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

• National Diabetes Audit – necessary because this dataset will support extensive and high-quality research into diabetes and its related co-morbidities and multi morbidity. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment, and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

• Improving Access to Psychological Therapies (IAPT), Mental Health and Learning Disabilities and Community Services – necessary because these datasets are hoped to improve the completeness of data that is often under reported from other sources and allow a more accurate picture of community service use, particularly in areas like mental health.

• Maternity Services Data Set (MSDS) – necessary because this dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC. The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies.

• Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

• Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information.

Data is only accessed within the Trusted Research Environment. Download of individual level data is not permitted by researchers. Inference Control is applied to statistics involving small numbers of individuals as recommended by the UK Information Commissioners Office. No immediately identifiable Demographic Data (names, addresses, phone numbers, dates of birth, contact details) are available within the TRE. Month and year of birth; first part of postcode; geographic region (LSOA) may be available.

Data returned by NHSEngland is subsequently linked with other health data directly provided by NHS Primary Care providers, other data directly provided by NHS Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data.

The data will be minimised as follows:

• Limited to the cohort of ~70,000 volunteers with valid NHS numbers (as at time of renewal ~7/2024) who gave individual written informed consent to participate.

• Full data from earliest available to latest available on all volunteers.

Genes & Health volunteers have given written informed consent for their NHS health data to be accessed stored and processed in line with the study objectives. Volunteers also gave consent for linkage to their medical records, including from their GP and hospital records as well as from nationally held datasets such as those included in this agreement, for research purposes.

Genes & Health recruitment started in 2015 in East London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited approx. 70,000 volunteers (at March 2024) and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on East London, Luton, Bradford, West Midlands, and Greater Manchester). The cohort is expected to reach over 100,000 volunteers. Recruitment is ongoing.

Queen Mary University of London is the research Sponsor and the Data Controller as the organisation responsible for ensuring that the data will only be processed for the purpose described above.

The lawful basis for processing personal data under the UK GDPR is:

• Article 6(1)(e) - processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller.

The lawful basis for processing special category data under the UK GDPR is:

• Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

Genes & Health also has written individual consent, which provides an additional English Common Law lawful basis.

This processing is in the public interest as British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, and this resource looks to maximise its potential to deliver benefit to population health to these two groups.

The funding comes from multiple sources. Substantial funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for England Catalyst, Barts Charity, Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames), Alnylam Pharmaceuticals, Genomics PLC, EU Horizon 2020, National Institutes of Health (USA); and a Life Sciences Industry Consortium of Astra Zeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. Funding is in place until 2030. Funding to continue the work described will be sought on an ongoing basis.

The Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of Queen Mary University of London. Wellcome Sanger Institute’s role is limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data.

Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data.

Genes & Health has developed an industry consortium with partners across a range of life sciences industry companies to support research that maximises the translation from basic science to therapeutics and clinical care. Consortium partners are Astra Zeneca, Pfizer, GSK, Merck, Novo Nordisk, Maze Theraputics, BMS, and Takeda. Consortium partners are working with the Genes & Health bioresource, under rigorous regulations set out in a Consortium Agreement. At present consortium partners are using genetic and non-NHS England health datasets. If partners use individual level raw NHS England data in the future, QMUL will require these partners to sign up to the sublicense agreement. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented.

QMUL users also access NHS England data under this DSA.

A Public and Patient Involvement and Engagement group helped refine the purpose of the research. The group strongly supported the collection of the data for the purposes described above. The Genes & Health Community Advisory Board meets 3-4 times a year and provides advice by email as required for other issues.

Volunteers are regularly invited to participate in workshops, focus groups, and informal information sessions. A newsletter is distributed 2-3x per year by Mailchimp email. These activities are designed to educate volunteers about health and genetics, and to gather public opinions on Genes & Health research.

Expected output

The expected outputs of the processing will be:

• Submissions of scientific biomedical publications to peer reviewed journals – ongoing can be found under the following link https://www.genesandhealth.org/about-study/scientific-publications

• Presentations at appropriate conferences - national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications.

The outputs will not contain NHS England data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the dataset(s) from which the information was derived.

The outputs will be communicated to relevant recipients through the following dissemination channels:

- Journals

- Educational workshops and interactive web- and app involving the wider scientific community supported by the QMUL Life Sciences Initiative

- Conferences

- Social media: instagram, X, linkedin, bluesky.

- Public events the study team are guided by NIHR INVOLVE (https://www.invo.org.uk/) in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g., Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families.

- Patient Information leaflets available at www.genesandhealth.org

- Press/media engagement

- Participant newsletters

Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are design to educate participants about health and genetics, and to gather public opinions on Genes & Health research.

Expected targets for outputs are, (a) short-term, e.g., publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g., building and expanding the bioresource within 5 years and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years).

Outputs have already been generated (see Yielded Benefits, below) and further outputs are to be expected to be generated continuously until the end of the study.

Benefits reported

Genes and Health first received NHS England data in Dec 2021. Data has been integrated with other health care datasets (e.g. direct from local primary care and secondary care NHS Providers), and with genetic data and volunteer recall data. The genetic data (GSA chip genotyping, exome sequencing) has only been released in the last few years and laboratory work is ongoing.

Genes & Health has attracted 116 research applications from scientists worldwide to use the resource, in addition to 8 international life sciences companies. The Genes & Health Trusted Research Environment has 239 active users.

Genes & Health has published 44 scientific publications (https://www.genesandhealth.org/about-study/scientific-publications) , many of the very highest international excellence including 10 in Nature, 2 in Nature Medicine, 1 in Nature Genetics, 1 in Science, 1 in New England Journal of Medicine, and 2 in Cell. Many more publications will follow over the next 5 years as the exome sequencing data is fully explored.

Measurable benefits yielded include:

- 50 volunteers diagnosed with LDLR gene Familial Hypercholesterolemia, an inherited condition with very high risk of early onset heart attack. No volunteers knew their diagnosis, and almost all were either not or were undertreated, despite effective treatments being available. All have now been treated and referred to NHS Cardiology clinics for long term management and prevention. See press release https://www.qmul.ac.uk/blizard/research/featured-research/look-at-those-genes--gene-research-offers-better-health-outcomes-for-british-bangladeshi-and-pakistani-people/?ocid=BingNews%2CBingNews

- 1 of the volunteers with Familial Hypercholesterolemia took part in a novel phase 1 trial aimed at a single dose (liver gene editing) treatment to treat once and for all lifelong. The participant met the inclusion criteria due to the high quality research data from Genes & Health. Press release: https://www.qmul.ac.uk/media/news/2023/smd/a-participant-from-queen-mary-university-of-london-genes--health-study-is-the-10th-person-enrolled-in-a-gene-editing-clinical-trial-for-heart-disease-.html

- A recent discovery that ~7% of south asians in Genes & Health carry a south asian-specific genetic variant that affects the accuracy of the HbA1C test for diabetes, and leads to late diagnosis and undertreatment. Press release: https://www.qmul.ac.uk/media/news/2024/fmd/standard-diabetes-test-may-be-inaccurate-for-10000s-of-south-asian-people-in-uk.html and https://www.diabetes.org.uk/about-us/news-and-views/type-2-diabetes-test-may-be-inaccurate-thousands-south-asian-people

- The discovery of a new pathway for growth, body composition and the onset of puberty (Nature 2021, https://doi.org/10.1038/s41586-021-04088-9) through recall of a Genes & Health volunteer lacking normal melanocortin receptor MC3R function.

- The discovery of human variants influencing COVID-19 hospitalisation and severity that are more frequent in south Asian versus white European populations and explain a part of the poorer outcomes in these ethnic groups (Nature 2021, https://doi.org/10.1038/s41586-021-03767-x).

- The study of a Genes & Health volunteer lacking the HAO1 gene, leading to a new drug treatment lumasiran (targeting HAO1) for the rare disease paediatric hyperoxaluria. Genes & Health data was included in the successful FDA New Drug Application and subsequent regulatory approval (eLife 2020, https://doi.org/10.7554/eLife.54363).

- Discovery that south asians (in Genes & Health) have a high prevalence of genetic variants influencing clopidogrel metabolism (and hence effectiveness). The drug clopidogrel is used to treat heart attacks, and volunteers with the variant had a higher rate of recurrent heart attack when treated with clopidogrel ( https://doi.org/10.1016/j.jacadv.2023.100573). This is likely to change prescribing guidelines in these ethnic groups.

- Attracting £30m of international life sciences industry investment to the UK for Genes & Health Industry Consortium 1.

DARS-NIC-338864-B3Z3J-v4.2 19 July 2024 to 18 July 2025
Title
Genes and Health
Commercial
Yes
Sublicensing
Yes
Datasets
18
Files released
109

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); Demographics; Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Improving Access to Psychological Therapies (IAPT) v1.5; Improving Access to Psychological Therapies (IAPT) v2; Maternity Services Data Set (MSDS) v1.5; Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); National Diabetes Audit; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-338864-B3Z3J-v3.3

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-338864-B3Z3J-v3.3
FieldWasBecame
Start date2023-12-182024-07-19
End date2024-07-072025-07-18

Objective for processing

Queen Mary University of London (QMUL) requires requests access to NHS England data for the purpose of the following research project: Genes and & Health Study. (www.genesandhealth.org). The following is a summary of the aims of the research project provided by Genes and & Health Study on behalf of Queen Mary University of London. Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis. Genes & Health is currently used by over 120 groups of researchers worldwide, with an outstanding set of outputs and publications (https://www.genesandhealth.org/about-study/scientific-publications). Researchers access Genes & Health data via the Genes & Health Trusted Research Environment. QMUL seeks to combine data from NHS England with data from the Genes and Health Bioresource. This linked resource will be made available via sub-licensing arrangements. The strategic objective of Genes and Health is to develop and maintain a Bioresource of genetic and health record data available to the research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin through high quality research. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at QMUL have built a database to investigate genetic data and health data from Genes & Health participants, of which includes data from NHS England under this Data Sharing Agreement. This database will be used by researchers at QMUL, as well as researchers in other universities, charities, and commercial companies.” Genes & Health seeks to combine data from NHS England with existing data from genetics, other biological data from volunteer recall, health data from multiple other NHS health providers, and volunteer questionnaires data from the Genes & Health. Access to NHS England data would be via sub-licensing arrangements. The strategic objective of Genes & Health is to develop and maintain a bioresource of genetic and health record data available to the worldwide research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin people and their communities worldwide through high quality research. Non-white European ethnicities are profoundly underrepresented in large human genetic studies, biobanks, and clinical trials. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at Queen Mary University of London have built a Trusted Research Environment to investigate genetic data and health data from Genes & Health volunteers, which includes data from NHS England under this Data Sharing Agreement. This database will be used by researchers at Queen Mary University of London, as well as other researchers in other universities, academic institutions, and commercial life sciences industry companies. [2 paragraphs unchanged] • Hospital Episode Statistics (HES) Admitted Patient Care – necessary because the data [40 words unchanged] and treated during hospital admissions, as well as episodes of pregnancy care. • (HES) Accident & Emergency / Emergency Care Data Set (ECDS) – necessary [35 words unchanged] beyond simply diagnosis, giving valuable information about treatment and the care received. (HES) • HES) Critical Care – necessary because the dataset will support research into severe [67 words unchanged] episode of care when a myocardial infarction was recorded will imply severity. • (HES) Outpatients – necessary because the dataset will support research into ill [34 words unchanged] including capturing the diagnosis and monitoring of long-term conditions in outpatient care. • Mental Health Service Dataset (including its predecessors the Mental Health Minimum Data [60 words unchanged] states of mental ill health and disease, the care and treatment received. • Civil Registration Mortality – necessary because this dataset will support research into [19 words unchanged] and their long-term outcomes, in this case death and cause of death. • Cancer Registration – necessary because this will supplement existing sources and aid the analyses of the links between genetics and cancer. • Demographics – necessary to mirror those collected independently by Genes & Health during follow-up contacts with participants, volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends. • National Diabetes Audit – necessary because this dataset will support extensive and [43 words unchanged] as important patient characteristics such as the calculation of socioeconomic status indices. • Improving Access to Psychological Therapies (IAPT), Mental Health and Learning Disabilities and [27 words unchanged] accurate picture of community service use, particularly in areas like mental health. • Maternity Services Data Set (MSDS) – necessary because this dataset will support [11 words unchanged] treated during pregnancy. This dataset will complement that obtained from HES APC. The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies. The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies. • Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends. Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with participants, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends. • Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information. Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information. Data is only accessed within the Trusted Research Environment. Download of individual level data is not permitted by researchers. Inference Control is applied to statistics involving small numbers of individuals as recommended by the UK Information Commissioners Office. No immediately identifiable Demographic Data (names, addresses, phone numbers, dates of birth, contact details) are available within the TRE. Month and year of birth; first part of postcode may be available. Genetic data is available which makes the dataset Personal Data under GDPR. The level of the data will be identifiable. Data returned is subsequently linked with other health data directly provided by Primary Care providers, other data directly provided by Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data. Data returned is subsequently linked with other data directly provided by Primary Care providers, other data directly provided by Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data. [1 paragraph unchanged] • Limited to a the cohort of 65,000 patients 70,000 volunteers with valid NHS numbers (as at time of 2023) renewal ~7/2024) who consented gave individual written informed consent to participate. [2 paragraphs unchanged] Genes & Health recruitment started in 2015 in East London, where limited [5 words unchanged] through linkage to local health systems. Genes & Health has recruited approx. 65,000 61,000 volunteers (at March 2024) and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on East London, Luton, Bradford Bradford, West Midlands, and Greater Manchester). The cohort is expected to reach 100,000 volunteers. Queen Mary University of London is the research sponsor Sponsor and the controller Data Controller as the organisation responsible for ensuring that the data will only be processed for the purpose described above. [1 paragraph unchanged] • Article 6(1)(e) - processing is necessary for the performance of a task [5 words unchanged] interest or in the exercise of official authority vested in the controller. [1 paragraph unchanged] • Article 9(2)(j) - processing is necessary for archiving purposes in the public [45 words unchanged] to safeguard the fundamental rights and the interests of the data subject. Genes & Health also has written individual consent, which additionally provides a common law basis. [1 paragraph unchanged] The funding comes from multiple sources. Current Substqntial funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for [20 words unchanged] Institute for Health Research Clinical Research Network (North Thames), Alnylam Pharmaceuticals, Genomics PLC; PLC, EU Horizon 2020, National Institutes of Health (USA); and a Life Sciences Industry Consortium of Astra Zeneca PLC, Bristol-Myers Squibb [31 words unchanged] to continue the work described will be sought on an ongoing basis. Swansea University, the The Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of Queen Mary University of London. Swansea University’s and Wellcome Sanger Institute’s roles are role is limited to ensuring that the data will only be processed for the [24 words unchanged] not have access to the Genes & Health or NHS England data. Several organisations have provided data or services but are not processors of the NHS England data. These are Kings College London and London, UK Biocentre, Broad Institute (USA) who have processed DNA samples from participants. volunteers. Discovery Data Service have facilitated primary healthcare record data access from members [62 words unchanged] Genes & Health project and do not have access to the data. Genes & Health has developed an industry consortium with partners across a range of pharmaceutical life sciences industry companies to support research that maximises the translation from basic science to therapeutics and clinical care. Consortium partners are Astra Zeneca, Pfizer, GSK, Merck, Novo Nordisk, Maze Theraputics, BMS, and Takeda. Consortium partners [84 words unchanged] and UK GDPR compliance are upheld, and that commercial exploitation is prevented. A Public and Patient Involvement and Engagement group helped refine the purpose [11 words unchanged] data for the purposes described above. The Genes & Health Community Advisory Group Board meets 3-4 times a year and provides advice by email as required for other issues. Participants Volunteers are regularly invited to participate in workshops, focus groups, and informal information sessions. A newsletter is distributed 2-3x per year by Mailchimp email. These sessions activities are designed to educate participants volunteers about health and genetics, and to gather public opinions on Genes & Health research.

Processing activities

Queen Mary University of London will transfer data to NHS England. The data will consist of research volunteer identifying details (specifically NHS Number, Date of Birth, Postcode, Name, Gender, and a unique person ID) for the cohort to be linked with NHS England data. NHS England data will provide the relevant records from the HES, ECDS, [27 words unchanged] Measures and National Diabetes Audit datasets to Queen Mary University of London. The NHS England provided data will contain no direct identifying identifiable Demographic Data (name, address, date of birth, contact details and NHS number). NHS England provided data items. However the data will be identifiable but individuals will not be reidentified through linkage with other data in the possession may contain month and year of the recipient. birth, and first part of postcode. Participant level health data from a variety of datasets, along with the unique cohort study ID supplied by Genes & Health, will be transferred to Genes & Health Google Cloud TRE and Swansea University. Queen Mary University of London already has identifiable Demographic Data on research volunteers through their direct research participation. This will not be linked with NHS England Data. The data will be stored inside The UK Secure eResearch Platform (SeRP) with servers located at Swansea University and Genes & Health Google Cloud Trusted Research Environment, using only servers and storage located in the UK. Volunteer level health data from a variety of datasets from a variety of providers, including NHS England, along with the unique cohort study ID supplied by Genes & Health, will be transferred to Genes & Health Trusted Research Environment. Researchers from external institutions, including universities and commercial companies, will, with the agreement and oversight of QMUL, analyse Genes & Health data on one of two secure data environments; either inside the UK SeRP platform (provided by Swansea University), or the Trusted Research Environment administered by the Welcome Sanger Institute. UK SeRP is where NHS England data has been stored to date, with servers located at Swansea University Campuses in Wales. Both The Sanger Institute and UK SeRP/Swansea University hold Data Security and Protection Toolkits, and both environments hold ISO27001 certifications. The UK-SeRP platform is user- friendly and suitable for some types of analysis but does not have the capacity for high performance compute needed to analyse large genetic datasets. The second platform, the Trusted Research Environment (TRE) is a secure platform that has been developed by QMUL and Welcome Sanger Institute specifically for Genes & Health, based on the successful platform developed by the similar Finngen consortium (https://www.finngen.fi/en). Full ISO27001 Certification was achieved in March 2023. The TRE is hosted on UK-based Google Cloud servers located in London, England, and the IT administration is provided by Genes and Health funded staff at the Wellcome Sanger Institute. It is the intention of Genes and Health that most data users move to the TRE over the next year. The data will be stored at Queen Mary University of London including Queen Mary University of London's (as Data Controller) Genes & Health Trusted Research Environment, using only servers and storage located in England (currently, Google Cloud Platform's eu-west-2 London datacentre). Genes & Health is an open access resource, making anonymised genetic and medical record data available in a Trusted Research Environment (aka data safe haven) to external researchers and industrial partners. Researchers from external institutions, including universities and commercial companies, will, with the agreement and oversight of QMUL, analyse Genes & Health data in a secure data environment: the Genes & Health Trusted Research Environment. This is administered by the Welcome Sanger Institute, as part of the Queen Mary University of London collaboration with Wellcome Sanger Institute on the Genes & Health programme. The Wellcome Sanger Institute holds Data Security and Protection Toolkits, and full ISO27001 certification for the Genes & Health Trusted Research Environment. The Genes & Health Trusted Research Environment (TRE) is a secure platform that has been developed by QMUL and Welcome Sanger Institute specifically for Genes & Health, based on the successful platform developed by the similar Finngen consortium (https://www.finngen.fi/en). Full ISO27001 Certification was achieved in March 2023. IT administration is provided by Genes and Health funded staff at the Wellcome Sanger Institute. A Data Access Agreement (DAA) will be signed by external institutions before access to Genes & Health data is granted. A sublicense will also be signed for access to NHS England data. Genes & Health is an open access resource, available to any bona fide scientific researcher worldwide who makes an application that is approved by the Genes Executive. Sensitive applications (e.g. mental health, sexual health, consanguinity) are subject to further scrunity by the Genes & Health Community Advisory Board. Access to the TRE is subject to payment (to cover costs only), Data Access Agreement execution, and Information Governance training. The data will be accessed by authorised personnel via remote access. The data will remain on the systems data controlled by QMUL within the UK. A Data Access Agreement (DAA) will be signed by external institutions before access to Genes & Health data is granted. A sublicense will also be signed for access to NHS England data if access to NHS England Data or NHS England Manipulated Data is required. Users and institutions may choose not to have access to NHS England Data or NHS England Manipulated Data, but choose to have access to other Genes & Health data. Personnel Research users are legally prohibited prohibited, and not technically capable (via Identity and Access Management permissions), of downloading or copying data from the Trusted Research Environment to local or other devices. The data will only be accessed within the UK or the EEA.. The data will always remain in England, on the systems Data Controlled by QMUL, and will not be downloadable. Data processing will only be carried out by substantive employees of Swansea University, Analysts and researchers from Wellcome Sagner Institute and Sanger Institute, Queen Mary University of London or London, and the Institutions of approved researchers via sublicensing. sublicensing will process/analyse the NHS England data for the purposes described above. All personnel accessing the data have been appropriately trained in data protection and confidentiality. All personnel will have up to date NHS information governance training. training, annually refreshed. The data will be linked at person record level with NHS Primary Care data, NHS Secondary Care data direct from Secondary Care providers, participant volunteer questionnaire data, participant volunteer survey data, participant volunteer genetic data, and other bio sample data obtained at volunteer recall. [1 paragraph unchanged] The identifying details will be stored in a separate Genes & Health database (which does not have NHS England Data) to the linked dataset used for analysis. All analyses will use the pseudonymised dataset. Trusted Research Environment dataset using the cohort unique person ID. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised Trusted Research Environment dataset. Analysts and researchers from Swansea University, Wellcome Sanger Institute, Queen Mary University of London and the Institutions of approved researchers via sublicensing will process/analyse the data for the purposes described above.

Expected output

[15 paragraphs unchanged] Some outputs Outputs have already been published from previous data received from NHS England generated (see Yielded Benefits, below) and can be found at https://www.genesandhealth.org/about-study/scientific-publications, further outputs are to be expected to be generated continuously until the end of the study.

Expected measurable benefits

[8 paragraphs unchanged] Dissemination of new knowledge arising from the Genes & Health bioresource is [16 words unchanged] strategies, as well as finding new causes of disease or its complications. For example, Genes & Health has already generated new knowledge with the potential to benefit health care through its recent work combining information from a rare genetic variant (in the HAO1 gene) with health data (from local health data sources). Dissemination of study findings related to this gene function through academic collaboration has supported critical drug development for a rare, life-threatening metabolic disorder (primary hyperoxaluria). Identification of the effects of genetics on health (including rare gene variants, or gene differences associated with parental relatedness) could improve clinical care through better genetic counselling to families at risk and identification of risk early in the life course. [7 paragraphs unchanged] For example, charities have been involved in a recent finding from Genes & Health of a link between autozygosity and type 2 diabetes (Asthma UK, Diabetes UK) to advise on lay summary and publicity.

Benefits reported

Genes and Health received NHS England data in Dec 2021. These integrated datasets are difficult and time-consuming to produce but mean that Genes & Health has very high-quality health data that will support research projects. This linkage is within the scope of the linkages explained within this DSA. Genes and Health first received NHS England data in Dec 2021. Data has been integrated with other health care datasets (e.g. direct from local primary care and secondary care NHS Providers), and with genetic data and volunteer recall data. The genetic data (GSA chip genotyping, exome sequencing) has only been released in the last few years and laboratory work is ongoing. An example of how these improved health datasets are already being used is QMUL's contribution to the international INTERVENE project (https://www.interveneproject.eu/) which is using Artificial Intelligence and machine learning to understand the ways in which genetics contributes to disease risk. Genes & Health’s contribution to INTERVENE is really important because they provide data from people of Pakistani and Bangladeshi backgrounds, which would otherwise be underrepresented in the project. INTERVENE uses Genes & Health data, including health data that contains NHS England, to develop scores that can help doctors work out how likely it is that a patient will develop a disease. The new generation of genetic risk scores will help to diagnose and treat disease earlier, and so improve health. Genetic risk scores are in trials with the NHS, and new generations of scores built on high-quality and diverse datasets like INTERVENE will perform better and improve health further. Genes & Health has attracted 116 research applications from scientists worldwide to use the resource, in addition to 8 international life sciences companies. The Genes & Health Trusted Research Environment has 239 active users. Genes & Health continues to provide practical return of health and genetic data to participants and the community. For example, we Genes & Health have identified participants who are at a high risk of diseases like inherited high cholesterol - based on a combination of genetic data and health data (including those from NHS England) - and invited those participants to additional screening and referral for treatment. Some details are here: https://www.genesandhealth.org/research/research-studies-approved/s00013-genotype-first-recall-study-extreme-genetic-risk-atherosclerotic We also run regular workshops that aim to improving peoples' understanding of genetics and their own health in order to make informed health choices and improve uptake of existing healthcare pathways (e.g. https://www.genesandhealth.org/research-studies-approved/s00052-improving-genetics-communication). Genes & Health has published 44 scientific publications (https://www.genesandhealth.org/about-study/scientific-publications) , many of the very highest international excellence including 10 in Nature, 2 in Nature Medicine, 1 in Nature Genetics, 1 in Science, 1 in New England Journal of Medicine, and 2 in Cell. Many more publications will follow over the next 5 years as the exome sequencing data is fully explored. Analyses of data from the Genes & Health cohort have resulted in multiple high-impact publications. An up-to-date list is maintained on the study website (https://www.genesandhealth.org/about-study/scientific-publications). Measurable benefits yielded include: - 50 volunteers diagnosed with LDLR gene Familial Hypercholesterolemia, an inherited condition with very high risk of early onset heart attack. No volunteers knew their diagnosis, and almost all were either not or were undertreated, despite effective treatments being available. All have now been treated and referred to NHS Cardiology clinics for long term management and prevention. See press release https://www.qmul.ac.uk/blizard/research/featured-research/look-at-those-genes--gene-research-offers-better-health-outcomes-for-british-bangladeshi-and-pakistani-people/?ocid=BingNews%2CBingNews - 1 of the volunteers with Familial Hypercholesterolemia took part in a novel phase 1 trial aimed at a single dose (liver gene editing) treatment to treat once and for all lifelong. The participant met the inclusion criteria due to the high quality research data from Genes & Health. Press release: https://www.qmul.ac.uk/media/news/2023/smd/a-participant-from-queen-mary-university-of-london-genes--health-study-is-the-10th-person-enrolled-in-a-gene-editing-clinical-trial-for-heart-disease-.html - A recent discovery that ~7% of south asians in Genes & Health carry a south asian-specific genetic variant that affects the accuracy of the HbA1C test for diabetes, and leads to late diagnosis and undertreatment. Press release: https://www.qmul.ac.uk/media/news/2024/fmd/standard-diabetes-test-may-be-inaccurate-for-10000s-of-south-asian-people-in-uk.html and https://www.diabetes.org.uk/about-us/news-and-views/type-2-diabetes-test-may-be-inaccurate-thousands-south-asian-people - The discovery of a new pathway for growth, body composition and the onset of puberty (Nature 2021, https://doi.org/10.1038/s41586-021-04088-9) through recall of a Genes & Health volunteer lacking normal melanocortin receptor MC3R function. - The discovery of human variants influencing COVID-19 hospitalisation and severity that are more frequent in south Asian versus white European populations and explain a part of the poorer outcomes in these ethnic groups (Nature 2021, https://doi.org/10.1038/s41586-021-03767-x). - The study of a Genes & Health volunteer lacking the HAO1 gene, leading to a new drug treatment lumasiran (targeting HAO1) for the rare disease paediatric hyperoxaluria. Genes & Health data was included in the successful FDA New Drug Application and subsequent regulatory approval (eLife 2020, https://doi.org/10.7554/eLife.54363). - Discovery that south asians (in Genes & Health) have a high prevalence of genetic variants influencing clopidogrel metabolism (and hence effectiveness). The drug clopidogrel is used to treat heart attacks, and volunteers with the variant had a higher rate of recurrent heart attack when treated with clopidogrel ( https://doi.org/10.1016/j.jacadv.2023.100573). This is likely to change prescribing guidelines in these ethnic groups. - Attracting £30m of international life sciences industry investment to the UK for Genes & Health Industry Consortium 1. We are now pitching for Consortium 2 which we hope will bring a further £50m and enhance the UK's reputation as a life sciences powerhouse.

Objective for processing

Queen Mary University of London (QMUL) requests access to NHS England data for the purpose of the following research project: Genes & Health (www.genesandhealth.org).

The following is a summary of the aims of the research project provided by Genes & Health on behalf of Queen Mary University of London.

Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis. Genes & Health is currently used by over 120 groups of researchers worldwide, with an outstanding set of outputs and publications (https://www.genesandhealth.org/about-study/scientific-publications). Researchers access Genes & Health data via the Genes & Health Trusted Research Environment.

Genes & Health seeks to combine data from NHS England with existing data from genetics, other biological data from volunteer recall, health data from multiple other NHS health providers, and volunteer questionnaires data from the Genes & Health. Access to NHS England data would be via sub-licensing arrangements.

The strategic objective of Genes & Health is to develop and maintain a bioresource of genetic and health record data available to the worldwide research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin people and their communities worldwide through high quality research. Non-white European ethnicities are profoundly underrepresented in large human genetic studies, biobanks, and clinical trials. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at Queen Mary University of London have built a Trusted Research Environment to investigate genetic data and health data from Genes & Health volunteers, which includes data from NHS England under this Data Sharing Agreement. This database will be used by researchers at Queen Mary University of London, as well as other researchers in other universities, academic institutions, and commercial life sciences industry companies.

Genes & Health will charge commercial partners and academic partners for use of the bioresource, but this will cover infrastructure and sustainability costs only and will not be profit-making.

The following NHS England data will be accessed:

• Hospital Episode Statistics (HES) Admitted Patient Care – necessary because the data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. This dataset will support research into ill health and disease identified, diagnosed, and treated during hospital admissions, as well as episodes of pregnancy care.

• (HES) Accident & Emergency / Emergency Care Data Set (ECDS) – necessary because this dataset will support research into acute presentations of ill health, e.g., heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

• HES) Critical Care – necessary because the dataset will support research into severe ill health/disease, specifically related to critical care in this context. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g., in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity.

• (HES) Outpatients – necessary because the dataset will support research into ill health and disease, e.g., depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care.

• Mental Health Service Dataset (including its predecessors the Mental Health Minimum Data Set and Mental Health and Learning Disabilities Data Set) – necessary because these datasets will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multi morbidity, specifically to characterise states of mental ill health and disease, the care and treatment received.

• Civil Registration Mortality – necessary because this dataset will support research into ill health and disease by identifying death and cause of death, specifically to characterise states of ill health, disease, and their long-term outcomes, in this case death and cause of death.

• Cancer Registration – necessary because this will supplement existing sources and aid the analyses of the links between genetics and cancer.

• Demographics – necessary to mirror those collected independently by Genes & Health during follow-up contacts with volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

• National Diabetes Audit – necessary because this dataset will support extensive and high-quality research into diabetes and its related co-morbidities and multi morbidity. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment, and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

• Improving Access to Psychological Therapies (IAPT), Mental Health and Learning Disabilities and Community Services – necessary because these datasets are hoped to improve the completeness of data that is often under reported from other sources and allow a more accurate picture of community service use, particularly in areas like mental health.

• Maternity Services Data Set (MSDS) – necessary because this dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC. The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies.

• Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with volunteers, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

• Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information.

Data is only accessed within the Trusted Research Environment. Download of individual level data is not permitted by researchers. Inference Control is applied to statistics involving small numbers of individuals as recommended by the UK Information Commissioners Office. No immediately identifiable Demographic Data (names, addresses, phone numbers, dates of birth, contact details) are available within the TRE. Month and year of birth; first part of postcode may be available. Genetic data is available which makes the dataset Personal Data under GDPR.

Data returned is subsequently linked with other health data directly provided by Primary Care providers, other data directly provided by Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data.

The data will be minimised as follows:

• Limited to the cohort of 70,000 volunteers with valid NHS numbers (as at time of renewal ~7/2024) who gave individual written informed consent to participate.

• Limited to data between 1997/98 to latest available.

Genes & Health volunteers have given written informed consent for their NHS health data to be accessed stored and processed in line with the study objectives. Volunteers also gave consent for linkage to their medical records, including from their GP and hospital records as well as from nationally held datasets such as those included in this agreement, for research purposes.

Genes & Health recruitment started in 2015 in East London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited approx. 61,000 volunteers (at March 2024) and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on East London, Luton, Bradford, West Midlands, and Greater Manchester). The cohort is expected to reach 100,000 volunteers.

Queen Mary University of London is the research Sponsor and the Data Controller as the organisation responsible for ensuring that the data will only be processed for the purpose described above.

The lawful basis for processing personal data under the UK GDPR is:

• Article 6(1)(e) - processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller.

The lawful basis for processing special category data under the UK GDPR is:

• Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

Genes & Health also has written individual consent, which additionally provides a common law basis.

This processing is in the public interest as British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, and this resource looks to maximise its potential to deliver benefit to population health to these two groups.

The funding comes from multiple sources. Substqntial funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for England Catalyst, Barts Charity, Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames), Alnylam Pharmaceuticals, Genomics PLC, EU Horizon 2020, National Institutes of Health (USA); and a Life Sciences Industry Consortium of Astra Zeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. Funding is in place until 2028. Funding to continue the work described will be sought on an ongoing basis.

The Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of Queen Mary University of London. Wellcome Sanger Institute’s role is limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data.

Several organisations have provided data or services but are not processors of the NHS England data. These are Kings College London, UK Biocentre, Broad Institute (USA) who have processed DNA samples from volunteers. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership. Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data.

Genes & Health has developed an industry consortium with partners across a range of life sciences industry companies to support research that maximises the translation from basic science to therapeutics and clinical care. Consortium partners are Astra Zeneca, Pfizer, GSK, Merck, Novo Nordisk, Maze Theraputics, BMS, and Takeda. Consortium partners are working with the Genes & Health bioresource, under rigorous regulations set out in a Consortium Agreement. At present consortium partners are using genetic and non-NHS England health datasets. If partners use individual level NHS England data in the future, QMUL will require these partners to sign up to the sublicense agreements. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented.

A Public and Patient Involvement and Engagement group helped refine the purpose of the research. The group strongly supported the collection of the data for the purposes described above. The Genes & Health Community Advisory Board meets 3-4 times a year and provides advice by email as required for other issues.

Volunteers are regularly invited to participate in workshops, focus groups, and informal information sessions. A newsletter is distributed 2-3x per year by Mailchimp email. These activities are designed to educate volunteers about health and genetics, and to gather public opinions on Genes & Health research.

Expected output

The expected outputs of the processing will be:

• Submissions of scientific biomedical publications to peer reviewed journals – ongoing can be found under the following link https://www.genesandhealth.org/about-study/scientific-publications

• Presentations at appropriate conferences - national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications.

The outputs will not contain NHS England data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the dataset(s) from which the information was derived.

The outputs will be communicated to relevant recipients through the following dissemination channels:

- Journals

- Educational workshops and interactive web- and app involving the wider scientific community supported by the QMUL Life Sciences Initiative

- Webinars open to members of the Queens Mary University of London.

- Social media through twitter found at (@eastlondongenes and @bradfordgenes, and @manchestergenes).

- Public events the study team are guided by NIHR INVOLVE (https://www.invo.org.uk/) in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g., Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families.

- Patient Information leaflets available at www.genesandhealth.org

- Press/media engagement

- Participant newsletters

Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are design to educate participants about health and genetics, and to gather public opinions on Genes & Health research.

Expected targets for outputs are, (a) short-term, e.g., publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g., building and expanding the bioresource within 5 years and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years).

Outputs have already been generated (see Yielded Benefits, below) and further outputs are to be expected to be generated continuously until the end of the study.

Benefits reported

Genes and Health first received NHS England data in Dec 2021. Data has been integrated with other health care datasets (e.g. direct from local primary care and secondary care NHS Providers), and with genetic data and volunteer recall data. The genetic data (GSA chip genotyping, exome sequencing) has only been released in the last few years and laboratory work is ongoing.

Genes & Health has attracted 116 research applications from scientists worldwide to use the resource, in addition to 8 international life sciences companies. The Genes & Health Trusted Research Environment has 239 active users.

Genes & Health has published 44 scientific publications (https://www.genesandhealth.org/about-study/scientific-publications) , many of the very highest international excellence including 10 in Nature, 2 in Nature Medicine, 1 in Nature Genetics, 1 in Science, 1 in New England Journal of Medicine, and 2 in Cell. Many more publications will follow over the next 5 years as the exome sequencing data is fully explored.

Measurable benefits yielded include:

- 50 volunteers diagnosed with LDLR gene Familial Hypercholesterolemia, an inherited condition with very high risk of early onset heart attack. No volunteers knew their diagnosis, and almost all were either not or were undertreated, despite effective treatments being available. All have now been treated and referred to NHS Cardiology clinics for long term management and prevention. See press release https://www.qmul.ac.uk/blizard/research/featured-research/look-at-those-genes--gene-research-offers-better-health-outcomes-for-british-bangladeshi-and-pakistani-people/?ocid=BingNews%2CBingNews

- 1 of the volunteers with Familial Hypercholesterolemia took part in a novel phase 1 trial aimed at a single dose (liver gene editing) treatment to treat once and for all lifelong. The participant met the inclusion criteria due to the high quality research data from Genes & Health. Press release: https://www.qmul.ac.uk/media/news/2023/smd/a-participant-from-queen-mary-university-of-london-genes--health-study-is-the-10th-person-enrolled-in-a-gene-editing-clinical-trial-for-heart-disease-.html

- A recent discovery that ~7% of south asians in Genes & Health carry a south asian-specific genetic variant that affects the accuracy of the HbA1C test for diabetes, and leads to late diagnosis and undertreatment. Press release: https://www.qmul.ac.uk/media/news/2024/fmd/standard-diabetes-test-may-be-inaccurate-for-10000s-of-south-asian-people-in-uk.html and https://www.diabetes.org.uk/about-us/news-and-views/type-2-diabetes-test-may-be-inaccurate-thousands-south-asian-people

- The discovery of a new pathway for growth, body composition and the onset of puberty (Nature 2021, https://doi.org/10.1038/s41586-021-04088-9) through recall of a Genes & Health volunteer lacking normal melanocortin receptor MC3R function.

- The discovery of human variants influencing COVID-19 hospitalisation and severity that are more frequent in south Asian versus white European populations and explain a part of the poorer outcomes in these ethnic groups (Nature 2021, https://doi.org/10.1038/s41586-021-03767-x).

- The study of a Genes & Health volunteer lacking the HAO1 gene, leading to a new drug treatment lumasiran (targeting HAO1) for the rare disease paediatric hyperoxaluria. Genes & Health data was included in the successful FDA New Drug Application and subsequent regulatory approval (eLife 2020, https://doi.org/10.7554/eLife.54363).

- Discovery that south asians (in Genes & Health) have a high prevalence of genetic variants influencing clopidogrel metabolism (and hence effectiveness). The drug clopidogrel is used to treat heart attacks, and volunteers with the variant had a higher rate of recurrent heart attack when treated with clopidogrel ( https://doi.org/10.1016/j.jacadv.2023.100573). This is likely to change prescribing guidelines in these ethnic groups.

- Attracting £30m of international life sciences industry investment to the UK for Genes & Health Industry Consortium 1. We are now pitching for Consortium 2 which we hope will bring a further £50m and enhance the UK's reputation as a life sciences powerhouse.

DARS-NIC-338864-B3Z3J-v3.3 18 December 2023 to 7 July 2024
Title
Genes and Health
Commercial
Yes
Sublicensing
Yes
Datasets
18
Files released
183

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); Demographics; Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Improving Access to Psychological Therapies (IAPT) v1.5; Improving Access to Psychological Therapies (IAPT) v2; Maternity Services Data Set (MSDS) v1.5; Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); National Diabetes Audit; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-338864-B3Z3J-v2.9

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-338864-B3Z3J-v2.9
FieldWasBecame
Start date2023-06-302023-12-18

Unchanged: Objective for processing, Processing activities, Expected output, Expected measurable benefits, Benefits reported.

Objective for processing

Queen Mary University of London (QMUL) requires access to NHS England data for the purpose of the following research project: Genes and Health Study.

The following is a summary of the aims of the research project provided by Genes and Health Study on behalf of Queen Mary University of London.

Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis.

QMUL seeks to combine data from NHS England with data from the Genes and Health Bioresource. This linked resource will be made available via sub-licensing arrangements. The strategic objective of Genes and Health is to develop and maintain a Bioresource of genetic and health record data available to the research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin through high quality research. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at QMUL have built a database to investigate genetic data and health data from Genes & Health participants, of which includes data from NHS England under this Data Sharing Agreement. This database will be used by researchers at QMUL, as well as researchers in other universities, charities, and commercial companies.”

Genes & Health will charge commercial partners and academic partners for use of the bioresource, but this will cover infrastructure and sustainability costs only and will not be profit-making.

The following NHS England data will be accessed:

Hospital Episode Statistics (HES) Admitted Patient Care – necessary because the data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. This dataset will support research into ill health and disease identified, diagnosed, and treated during hospital admissions, as well as episodes of pregnancy care.

(HES) Accident & Emergency / Emergency Care Data Set (ECDS) – necessary because this dataset will support research into acute presentations of ill health, e.g., heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

(HES) Critical Care – necessary because the dataset will support research into severe ill health/disease, specifically related to critical care in this context. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g., in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity.

(HES) Outpatients – necessary because the dataset will support research into ill health and disease, e.g., depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care.

Mental Health Service Dataset (including its predecessors the Mental Health Minimum Data Set and Mental Health and Learning Disabilities Data Set) – necessary because these datasets will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multi morbidity, specifically to characterise states of mental ill health and disease, the care and treatment received.

Civil Registration Mortality – necessary because this dataset will support research into ill health and disease by identifying death and cause of death, specifically to characterise states of ill health, disease, and their long-term outcomes, in this case death and cause of death.

Cancer Registration – necessary because this will supplement existing sources and aid the analyses of the links between genetics and cancer.

Demographics – necessary to mirror those collected independently by Genes & Health during follow-up contacts with participants, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

National Diabetes Audit – necessary because this dataset will support extensive and high-quality research into diabetes and its related co-morbidities and multi morbidity. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment, and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

Improving Access to Psychological Therapies (IAPT), Mental Health and Learning Disabilities and Community Services – necessary because these datasets are hoped to improve the completeness of data that is often under reported from other sources and allow a more accurate picture of community service use, particularly in areas like mental health.

Maternity Services Data Set (MSDS) – necessary because this dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC.

The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies.

Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with participants, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information.

The level of the data will be identifiable.

Data returned is subsequently linked with other data directly provided by Primary Care providers, other data directly provided by Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data.

The data will be minimised as follows:

• Limited to a cohort of 65,000 patients (as of 2023) who consented to participate.

• Limited to data between 1997/98 to latest available.

Genes & Health Volunteers have given written informed consent for their NHS health data to be accessed stored and processed in line with the study objectives. Volunteers also gave consent for linkage to their medical records, including from their GP and hospital records as well as from nationally held datasets such as those included in this agreement, for research purposes.

Genes & Health recruitment started in 2015 in East London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited approx. 65,000 volunteers and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on East London, Luton, Bradford and Manchester). The cohort is expected to reach 100,000 volunteers.

Queen Mary University of London is the research sponsor and the controller as the organisation responsible for ensuring that the data will only be processed for the purpose described above.

The lawful basis for processing personal data under the UK GDPR is:

Article 6(1)(e) - processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller.

The lawful basis for processing special category data under the UK GDPR is:

Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

This processing is in the public interest as British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, and this resource looks to maximise its potential to deliver benefit to population health to these two groups.

The funding comes from multiple sources. Current funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for England Catalyst, Barts Charity, Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames), Alnylam Pharmaceuticals, Genomics PLC; and a Life Sciences Industry Consortium of Astra Zeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. Funding is in place until 2028. Funding to continue the work described will be sought on an ongoing basis.

Swansea University, the Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of Queen Mary University of London. Swansea University’s and Wellcome Sanger Institute’s roles are limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data.

Several organisations have provided data or services but are not processors of the NHS England data. These are Kings College London and UK Biocentre, who have processed DNA samples from participants. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership. Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data.

Genes & Health has developed an industry consortium with partners across a range of pharmaceutical companies to support research that maximises the translation from basic science to therapeutics and clinical care. Consortium partners are Pfizer, GSK, Merck, Novo Nordisk, Maze Theraputics, BMS, and Takeda. Consortium partners are working with the Genes & Health bioresource, under rigorous regulations set out in a Consortium Agreement. At present consortium partners are using genetic and non-NHS England health datasets. If partners use individual level NHS England data in the future, QMUL will require these partners to sign up to the sublicense agreements. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented.

A Public and Patient Involvement and Engagement group helped refine the purpose of the research. The group strongly supported the collection of the data for the purposes described above. The Genes & Health Community Advisory Group meets 3-4 times a year and provides advice by email as required for other issues.

Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are designed to educate participants about health and genetics, and to gather public opinions on Genes & Health research.

Expected output

The expected outputs of the processing will be:

• Submissions of scientific biomedical publications to peer reviewed journals – ongoing can be found under the following link https://www.genesandhealth.org/about-study/scientific-publications

• Presentations at appropriate conferences - national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications.

The outputs will not contain NHS England data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the dataset(s) from which the information was derived.

The outputs will be communicated to relevant recipients through the following dissemination channels:

- Journals

- Educational workshops and interactive web- and app involving the wider scientific community supported by the QMUL Life Sciences Initiative

- Webinars open to members of the Queens Mary University of London.

- Social media through twitter found at (@eastlondongenes and @bradfordgenes, and @manchestergenes).

- Public events the study team are guided by NIHR INVOLVE (https://www.invo.org.uk/) in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g., Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families.

- Patient Information leaflets available at www.genesandhealth.org

- Press/media engagement

- Participant newsletters

Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are design to educate participants about health and genetics, and to gather public opinions on Genes & Health research.

Expected targets for outputs are, (a) short-term, e.g., publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g., building and expanding the bioresource within 5 years and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years).

Some outputs have already been published from previous data received from NHS England and can be found at https://www.genesandhealth.org/about-study/scientific-publications, further outputs are to be expected to be generated continuously until the end of the study.

Benefits reported

Genes and Health received NHS England data in Dec 2021. These integrated datasets are difficult and time-consuming to produce but mean that Genes & Health has very high-quality health data that will support research projects. This linkage is within the scope of the linkages explained within this DSA.

An example of how these improved health datasets are already being used is QMUL's contribution to the international INTERVENE project (https://www.interveneproject.eu/) which is using Artificial Intelligence and machine learning to understand the ways in which genetics contributes to disease risk. Genes & Health’s contribution to INTERVENE is really important because they provide data from people of Pakistani and Bangladeshi backgrounds, which would otherwise be underrepresented in the project. INTERVENE uses Genes & Health data, including health data that contains NHS England, to develop scores that can help doctors work out how likely it is that a patient will develop a disease. The new generation of genetic risk scores will help to diagnose and treat disease earlier, and so improve health. Genetic risk scores are in trials with the NHS, and new generations of scores built on high-quality and diverse datasets like INTERVENE will perform better and improve health further.

Genes & Health continues to provide practical return of health and genetic data to participants and the community. For example, we Genes & Health have identified participants who are at a high risk of diseases like inherited high cholesterol - based on a combination of genetic data and health data (including those from NHS England) - and invited those participants to additional screening and referral for treatment. Some details are here: https://www.genesandhealth.org/research/research-studies-approved/s00013-genotype-first-recall-study-extreme-genetic-risk-atherosclerotic We also run regular workshops that aim to improving peoples' understanding of genetics and their own health in order to make informed health choices and improve uptake of existing healthcare pathways (e.g. https://www.genesandhealth.org/research-studies-approved/s00052-improving-genetics-communication).

Analyses of data from the Genes & Health cohort have resulted in multiple high-impact publications. An up-to-date list is maintained on the study website (https://www.genesandhealth.org/about-study/scientific-publications).

DARS-NIC-338864-B3Z3J-v2.9 30 June 2023 to 7 July 2024
Title
Genes and Health
Commercial
Yes
Sublicensing
Yes
Datasets
18
Files released
89

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Cancer Registration Data; Civil Registrations of Death; Community Services Data Set (CSDS); Demographics; Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Improving Access to Psychological Therapies (IAPT) v1.5; Improving Access to Psychological Therapies (IAPT) v2; Maternity Services Data Set (MSDS) v1.5; Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); National Diabetes Audit; Patient Reported Outcome Measures (Linkable to HES)

What changed from DARS-NIC-338864-B3Z3J-v1.2

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-338864-B3Z3J-v1.2
FieldWasBecame
Start date2022-07-082023-06-30
End date2023-07-072024-07-07
Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set: type of dataAnonymised - ICO Code CompliantIdentifiable
Civil Registrations of Death: type of dataAnonymised - ICO Code CompliantIdentifiable
Emergency Care Data Set (ECDS): type of dataAnonymised - ICO Code CompliantIdentifiable
Hospital Episode Statistics Accident and Emergency (HES A and E): type of dataAnonymised - ICO Code CompliantIdentifiable
Hospital Episode Statistics Admitted Patient Care (HES APC): type of dataAnonymised - ICO Code CompliantIdentifiable
Hospital Episode Statistics Critical Care (HES Critical Care): type of dataAnonymised - ICO Code CompliantIdentifiable
Hospital Episode Statistics Outpatients (HES OP): type of dataAnonymised - ICO Code CompliantIdentifiable
MSDS (Maternity Services Data Set) v1.5: type of dataAnonymised - ICO Code CompliantIdentifiable
Mental Health Minimum Data Set (MHMDS): type of dataAnonymised - ICO Code CompliantIdentifiable
Mental Health Services Data Set (MHSDS): type of dataAnonymised - ICO Code CompliantIdentifiable
Mental Health and Learning Disabilities Data Set (MHLDDS): type of dataAnonymised - ICO Code CompliantIdentifiable
National Diabetes Audit: type of dataAnonymised - ICO Code CompliantIdentifiable

Datasets: + Cancer Registration Data; + Community Services Data Set (CSDS); + Demographics; + Improving Access to Psychological Therapies (IAPT) v2; + Improving Access to Psychological Therapies Data Set_v1.5; + Patient Reported Outcome Measures (Linkable to HES)

Objective for processing

This data request comes from the Genes & Health Study, which is based at Queen Mary University of London (QMUL). (QMUL) requires access to NHS England data for the purpose of the following research project: Genes and Health Study. The following is a summary of the aims of the research project provided by Genes and Health Study on behalf of Queen Mary University of London. [1 paragraph unchanged] QMUL wish to combine data from NHS Digital with data from the Genes and Health BioResource. This linked resource will be made available via sub-licensing arrangements. QMUL seeks to combine data from NHS England with data from the Genes and Health Bioresource. This linked resource will be made available via sub-licensing arrangements. The strategic objective of Genes and Health is to develop and maintain a Bioresource of genetic and health record data available to the research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin through high quality research. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at QMUL have built a database to investigate genetic data and health data from Genes & Health participants, of which includes data from NHS England under this Data Sharing Agreement. This database will be used by researchers at QMUL, as well as researchers in other universities, charities, and commercial companies.” The strategic objective of Genes & Health is to develop and maintain a BioResource of genetic and health record data available to the research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin through high quality research. Genes & Health will charge commercial partners and academic partners for use of the bioresource, but this will cover infrastructure and sustainability costs only and will not be profit-making. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their lifecourse. The Genes & Health study team at QMUL have built a database to investigate genetic data and health data from Genes & Health participants, and hope to include data from NHS Digital. This database will be used by researchers at QMUL, as well as researchers in other universities, charities, and commercial companies, once they have been approved by an application review committee, and signed a sublicensing agreement. The following NHS England data will be accessed: Genes & Health Volunteers have given written informed consent for their health data to be accessed stored and processed in line with the study objectives. The study is organised as follows: Hospital Episode Statistics (HES) Admitted Patient Care – necessary because the data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. This dataset will support research into ill health and disease identified, diagnosed, and treated during hospital admissions, as well as episodes of pregnancy care. (i) Stage 1 research: genetic information from saliva samples provided at recruitment by Genes & Health volunteers is combined with health record data to characterise genetic variation and correlate this to physical characteristics, and states of health and disease. (HES) Accident & Emergency / Emergency Care Data Set (ECDS) – necessary because this dataset will support research into acute presentations of ill health, e.g., heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. (ii) Stage 2 research; information from genetic studies and/or health record information are used to guide focused translational research studies, e.g. investigating how a rare gene variant might impact on health. (HES) Critical Care – necessary because the dataset will support research into severe ill health/disease, specifically related to critical care in this context. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g., in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity. The article 6 justification for processing the requested data is stated in 1. (e) “processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller” (HES) Outpatients – necessary because the dataset will support research into ill health and disease, e.g., depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care. This data request is in the public interest under Article 9(2)(j). Specifically, it is in the public interest for scientific and statistical purposes in accordance with Article 89(1). Furthermore, the request is proportionate to the aim pursued, respects the essence of the right to data protection, and provides for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject. Mental Health Service Dataset (including its predecessors the Mental Health Minimum Data Set and Mental Health and Learning Disabilities Data Set) – necessary because these datasets will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multi morbidity, specifically to characterise states of mental ill health and disease, the care and treatment received. Dissemination of results arising from this data request, including those published from sublicenced data, will be fully anonymised, and will inform study volunteers and the wider public of the research findings at summary level. These findings will comprise new knowledge of health and disease and will be disseminated using lay language. Where findings might be sensitive, e.g. reporting an increased risk of a severe disease in the study population, or its genetic cause, the Community Advisory Group will be consulted to ensure effective and appropriate means of communication. Civil Registration Mortality – necessary because this dataset will support research into ill health and disease by identifying death and cause of death, specifically to characterise states of ill health, disease, and their long-term outcomes, in this case death and cause of death. Volunteers give consent for linkage to their medical records, including from their GP and hospital records as well as from nationally-held datasets such as those included in this agreement, for research purposes. Genes & Health is a multi-purpose bioresource, and combines health data with genetic data (obtained from saliva samples collected at recruitment) to undertake detailed characterisation of states of health and disease and generate new knowledge relating to these. Genes & Health recruitment started in 2015 in east London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited over 47,000 volunteers and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on east London, Luton, Bradford and Manchester). The cohort is expected to reach 100,000 volunteers. Linkage to national datasets is required to (a) expand the geographical coverage of data linkage where local datasets are not available or where healthcare is accessed beyond its limits, and (b) to enrich the datasets available using high quality multisource national data that are not available through local health systems, e.g. civil registration of deaths. Cancer Registration – necessary because this will supplement existing sources and aid the analyses of the links between genetics and cancer. Genes & Health receives funding from leading funders including the Wellcome Trust, Medical Research Council, National Institute for Health Research Clinical Research Networks and others. People of British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, including UK Biobank. Genes & Health is an open access resource, making anonymised genetic and medical record data available in a data safe haven to external researchers and industrial partners to deliver its aims and maximise its potential to deliver benefit to population health. The infrastructure of Genes & Health has, at its core, rigorous protocols to ensure the appropriate use of data, data safety and volunteer acceptability. Demographics – necessary to mirror those collected independently by Genes & Health during follow-up contacts with participants, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends. The study complements, but is distinct from, existing bioresources (e.g. the well-known Cambridge BioResource and UK Biobank) by its focus on two distinct UK ethnic minority groups with a different spectrum of genetic variation, high levels of health deprivation, and substantial parental relatedness. As such, Genes & Health will deliver new knowledge relevant to an understudied population at high risk of ill health. Genes & Health also complements existing large-scale open access health data resources, e.g. the Clinical Practice Research Datalink, by its linkage of health data to genetic data. National Diabetes Audit – necessary because this dataset will support extensive and high-quality research into diabetes and its related co-morbidities and multi morbidity. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment, and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. The research taking place within Genes & Health reflects community prioritisation, health need, research expertise and emerging knowledge. Furthermore, Genes & Health is an adaptable and responsive resource that can accommodate changing priority and need, an example being its recent rapid contribution to the International COVID-19 Host Genetic Initiative. Improving Access to Psychological Therapies (IAPT), Mental Health and Learning Disabilities and Community Services – necessary because these datasets are hoped to improve the completeness of data that is often under reported from other sources and allow a more accurate picture of community service use, particularly in areas like mental health. Use of data from key national health datasets will support the multipurpose, bioresource function of Genes & Health by expanding the depth and breadth of data available and by increasing the availability of health data on volunteers receiving care outside of local health systems where linkage is already in place. The scope of the data request outlined in this application is broad, and reflects the design and need of Genes & Health as a bioresource supporting a large portfolio of research projects in Stage 1 and Stage 2 studies: Maternity Services Data Set (MSDS) – necessary because this dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC. (a) Stage 1 research: genetic information from saliva samples provided at recruitment by Genes & Health volunteers is combined with health record data to characterise genetic variation and correlate this to biomedical traits and phenotypes, including states of health and disease. These studies generally require population-based analysis of association (e.g. in case-control studies of the genetic associations of prevalent disease, or the prediction of disease onset using risk scores) or longitudinal analysis of disease progression (e.g. using survival analysis incorporating risk factors, incident disease, and progression to complications and death). All analyses require careful adjustment for confounding variables (e.g. socioeconomic status, comorbidity/multimorbidity, smoking) and use of time-dependent covariates (e.g. pregnancy, progression from at-risk states). The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies. (b) Stage 2 research; information from genetic studies and/or health record information are used to guide focused translational research studies, e.g. investigating how a rare gene variant might impact on health. These studies are typically smaller in scale compared to Stage 1 studies, and might involve a small number of individuals (e.g. 5-100) carrying a rare genetic variant whose function is unknown. In such examples, health record data will be used to identify novel associations between genotype and health/disease status. Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with participants, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends. Genes & Health invites the following people to participate: Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information. • People living within reach of its research sites (currently east London, Luton, Bradford and Manchester) The level of the data will be identifiable. • Self described ethnicity: Bangladeshi, British-Bangladeshi, Pakistani, British-Pakistani Data returned is subsequently linked with other data directly provided by Primary Care providers, other data directly provided by Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data. • Age 16 or over (no upper limit) The data will be minimised as follows: • Willingness to give a saliva DNA sample for analysis • Limited to a cohort of 65,000 patients (as of 2023) who consented to participate. • Willingness to grant investigators access to medical records (GP, Hospital, and national NHS) for the full duration (stage 1 and stage 2) of the study. • Limited to data between 1997/98 to latest available. • Willingness to receive invitations for recall (stage 2 research activity) for the duration of the study (actual participation in all stage 2 research activity would be after a second informed consent process). Genes & Health Volunteers have given written informed consent for their NHS health data to be accessed stored and processed in line with the study objectives. Volunteers also gave consent for linkage to their medical records, including from their GP and hospital records as well as from nationally held datasets such as those included in this agreement, for research purposes. QMUL seek linkage of NHS Digital datasets at record level to all Genes & Health volunteers. This agreement is requesting pseudonymised, record-level data. Genes & Health recruitment started in 2015 in East London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited approx. 65,000 volunteers and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on East London, Luton, Bradford and Manchester). The cohort is expected to reach 100,000 volunteers. Genes & Health collects saliva samples from all participants, which QMUL use to generate genetic data like SNP genotyping and DNA sequencing. From a select group of volunteers QMUL also collect samples like blood and urine, and make clinical measurements like height and weight. To understand how a person's genetics interact with their health, to cause bad health or to protect against disease, researchers need to be able to analyse genetic data from Genes & Health alongside health data from NHS Digital – one set of data alone won’t help understand the connection between genetics and health. Queen Mary University of London is the research sponsor and the controller as the organisation responsible for ensuring that the data will only be processed for the purpose described above. The partners who will use the Genes & Health resource include academic research groups at UK universities (the vast majority of users) and at universities in the EEA. Partners will also include technology and pharmaceutical companies based in the UK, and EEA, who want to access Genes & Health data for research and development purposes. All partners, academic and commercial, are required to submit an application for a defined project for review, have a summary of their project published online, and making appropriate results of their work available to the research community. To date 45 applications have been received, of which only three were from commercial companies. The lawful basis for processing personal data under the UK GDPR is: Carefully curated, and minimised, data from the following resources are requested in this agreement: Article 6(1)(e) - processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller. ¥ Hospital Episode Statistics (HES), including HES A&E, OP, critical care and APC. The lawful basis for processing special category data under the UK GDPR is: ¥ HES APC: Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject. This dataset will support research into ill health and disease identified, diagnosed and treated during hospital admissions, as well as episodes of pregnancy care. The data obtained from HES APC will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. A number of data fields have been requested in addition to STUDY_ID to ensure adequate linkage between the Genes & Health ID and pseudonymised HES identifiers across HES datasets, This processing is in the public interest as British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, and this resource looks to maximise its potential to deliver benefit to population health to these two groups. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and pregnancy. This data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. The funding comes from multiple sources. Current funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for England Catalyst, Barts Charity, Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames), Alnylam Pharmaceuticals, Genomics PLC; and a Life Sciences Industry Consortium of Astra Zeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. Funding is in place until 2028. Funding to continue the work described will be sought on an ongoing basis. ¥ HES OP: Swansea University, the Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of Queen Mary University of London. Swansea University’s and Wellcome Sanger Institute’s roles are limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data. This dataset will support research into ill health and disease, e.g. depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. The data obtained from HES OP will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved), discharge and outpatient follow up. Several organisations have provided data or services but are not processors of the NHS England data. These are Kings College London and UK Biocentre, who have processed DNA samples from participants. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership. Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data. This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care. Genes & Health has developed an industry consortium with partners across a range of pharmaceutical companies to support research that maximises the translation from basic science to therapeutics and clinical care. Consortium partners are Pfizer, GSK, Merck, Novo Nordisk, Maze Theraputics, BMS, and Takeda. Consortium partners are working with the Genes & Health bioresource, under rigorous regulations set out in a Consortium Agreement. At present consortium partners are using genetic and non-NHS England health datasets. If partners use individual level NHS England data in the future, QMUL will require these partners to sign up to the sublicense agreements. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented. ¥ HES AE: A Public and Patient Involvement and Engagement group helped refine the purpose of the research. The group strongly supported the collection of the data for the purposes described above. The Genes & Health Community Advisory Group meets 3-4 times a year and provides advice by email as required for other issues. This dataset will support research into acute presentations of ill health, e.g. heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. The data obtained from HES A&E will complement the other HES datasets that are being requested also, giving QMUL health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. A number of data fields have been requested in addition to STUDY_ID to ensure adequate linkage between the study ID and pseudonymised HES identifiers across HES datasets, Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are designed to educate participants about health and genetics, and to gather public opinions on Genes & Health research. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. ¥ HES CC This dataset will support research into severe ill health/disease, specifically related to critical care in this context. The data obtained from HES Critical Care will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g. in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity. ¥ Emergency Care Data Set (ECDS)/ This dataset will support research into acute presentations of ill health, e.g. heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. The data obtained from ECDS will complement HES A&E and other HES datasets that are being requested also, to give health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. ¥ The Mental Health Services Data Set (MHSDS), and its predecessors, the Mental Health Minimum Data Set (MHMDS) and Mental Health and Learning Disabilities Data Set (MHLDDS) ¥ MHSDS: This dataset will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multimorbidity. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease, including those related to mental health. This data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. ¥ MHMDS/MHLDDS This dataset will support research into mental ill health in the Genes and Health volunteers, including the study of rare and common disease, and their genetic basis. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of mental ill health and disease, the care and treatment received. ¥ Maternity Service Data Set (MSDS) This dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g. body mass index at booking, or social factors) will also add valuable detail to the characterisation of our female study volunteers more broadly, for cross-sectional and longitudinal studies. ¥ Civil Registrations (Deaths) This dataset will support research into ill health and disease by identifying death and cause of death. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and their long-term outcomes, in this case death and cause of death. ¥ National Diabetes Audit This dataset will support extensive and high quality research into diabetes and its related co-morbidities and multimorbidity. The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and pregnancy. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. ¥ Bridging files HES:civil registrations, HES:MHSDS To allow the datasets to be abridged. Use of data from key national health datasets will support the multipurpose, bioresource function of Genes & Health by expanding the depth and breadth of data available and by increasing the availability of health data on volunteers receiving care outside of local health systems where linkage is already in place. The content of data analysis has also been designed carefully to mirror the NHS Digital data available via UK Biobank and Clinical Practice Research Datalink so that replication and validation studies can be performed across all of these valuable resources. This mirroring is to allow similar analyses to be replicated between datasets, no direct linkage of BioBank, CPRD and Genes & Health data is anticipated. The data requested is required for the following types of analyses: (a) Association-based genetic analysis of prevalent disease(s), its severity and characteristics (b) Prediction of disease risk and severity using polygenic risk scores (c) Longitudinal analysis of states of health and disease(s) through the lifecourse (including pregnancy), and their association with genetic factors (d) Discovery-based analysis of the impact of rare genetic variation on health and disease (e) Assessment of, and adjustment for, confounding influences of the above using covariates. (f) Assessment of causal relationships in health and disease using genetic data in Mendelian randomisation (g) Validation studies using multi-source health data to quality control self-report data (e.g. the study questionnaire and local health data) and data from other bioresources (e.g. UK Biobank, Clinical Practice Research Datalink) (h) Replication studies of known genetic factors on health and disease in an understudied population of British Pakistanis and British Bangladeshis. The study incorporates longitudinal analysis of states of health and disease and includes the study of historic diagnoses. For example, an admission with myocardial infarction in 1999, is important as it will show the earliest time a patient was diagnosed. This will be vital for understanding early onset disease. QMUL also have ongoing funding and recruitment to 2023 at least and expect Genes & Health to serve as a long-term bioresource in the future, therefore QMUL will be requesting an annual update to this data linkage to update previously linked records and add linkage to newly-recruited volunteers. The study is recruiting nationally across multiple sites (currently London, Luton, Bradford and Manchester), with the possibility of opening additional sites in the future. It is important for Genes & Health to be able to link to national data in order to capture health data on volunteers who move house. There are no alternative data sources that would achieve the full aims and objectives of the Genes & Health study. Local health system data has already been used to its maximum potential, but this is limited to those systems able to supply data, and the data returned are limited to patients treated within those local systems. Currently this includes Barts Health, and a small number of local CCGs who have signed data transfer agreements, but not primary or secondary care providers further afield. Genes & Health is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on east London, Luton and Bradford) and linkage to national datasets is required to (a) expand the geographical coverage of data linkage where local datasets are not available or where healthcare is accessed beyond its limits, and (b) to enrich the datasets available using high quality multisource national data that are not available through local health systems, e.g. civil registration of deaths. Although data has been requested from multiple NHS Digital datasets, stringent variable selection has been undertaken for the purposes of data minimisation. Genes & Health have taken the following steps to ensure that the variables requested are strictly necessary and achieve data minimisation: 1) Detailed curation of required variables to meet the Genes & Health study objectives, achieved through discussion and consensus with the core Genes & Health research and Executive teams. The focus of this data request is on diagnoses, clinical observations and related dates, rather than organisational information that does not have a clear purpose in the research objectives. 2) Non-replication of data already held – Genes & Health already hold data, such as full postcode, particularly where it is identifiable or sensitive. A limited set of non-identifiable data (year and month of birth, gender, postcode district and ethnicity) has been requested as a quality control check across datasets. 3) Genes & Health have followed previous large-scale bioresource studies who have used NHS Digital datasets to guide the choice of variables, including UK Biobank and the INTERVAL and COMPARE studies of blood donors, as well as the Clinical Practice Research Datalink. This data request mirrors the set of variables available in these studies that are relevant to the Genes & Health study objectives, allowing for cross-study validation and replication. This will significantly expand the utility of the data requested. 4) Genes & Health have taken extensive and detailed guidance from the DARS team (case officer and data production team member) to guide the selection of variables, in particular where multiple datasets exist for the same clinical areas (e.g. MHSDS, MHLDDS and MHMDS). 5) The request is also minimised to the size of the cohort, currently a little over 47,000 participants. Genes & Health is still actively recruiting, and the number of participants is expected to reach 100,000 by approximately 2023. QMUL is the sole data controller for Genes and Health data, including these data requested from NHS Digital. Swansea University - who provide the UK Secure eResearch Platform (UK SeRP), will store the data and are a data processor. Researchers from external institutions, including universities and commercial companies, will, with the agreement, direction and oversight of QMUL, analyse Genes & Health data on a platform provided by UK SeRP (part of Swansea University). This will be facilitated by sublicensing agreements with all organisations. An access register will detail these projects and organisations, and a summary of all projects will be published on the Genes & Health website. A Data Access Agreement (DAA) will be signed by external institutions before access to Genes & Health data is granted. This agreement sets out the scope of data processing activities that may be undertaken and the legal and technical data protection requirements. To try and reinforce good practice, individual researchers are also asked to sign, having read a letter summarising their data protection responsibilities in the project. Several organisations have provided data or services but are not data processors for the NHS Digital data. These are King’s College London and UK Biocentre, who have processed DNA samples from participants, and Wellcome Sanger Institute who have analysed DNA samples. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership. Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data. Sub-licenses will be granted upon successful application, where applicants satisfy section viii of the terms of reference – which includes data security, public benefit, impact on participants. Applications covering a sensitive topic (for example sexual health, consanguinity, mental health) will be reviewed by a community panel. Conditions will be compliance with the terms of the agreement, which include strict data security requirements imposed by our research environment and to only undertake work within in the remit of their application. QMUL are able to view and audit the work of collaborators, and as the environment is export controlled QMUL will actively review every request for data-out to make sure they are complying with the terms. Data disseminated by NHS Digital data is potentially identifiable in the hands of the Data Controller as QMUL have the ability to link the Study ID's to the identifiers. Data will, however, always be pseudonymised in the hands of the sub-licensees.

Processing activities

A list of individuals will be supplied to NHS digital, and will include patient NHS numbers, gender, dates of birth and postcodes. A unique cohort study ID will be included. No other health data will be supplied to NHS Digital. The cohort size is approx 50,000. Queen Mary University of London will transfer data to NHS England. The data will consist of identifying details (specifically NHS Number, Date of Birth, Postcode, Name, Gender, and a unique person ID) for the cohort to be linked with NHS England data. Participant level health data from a variety of datasets, along with the unique cohort study ID supplied by Genes & Health, will be transferred to Genes & Health. The requested data include details of recent and future use of mental health services, categorised as high risk data. NHS England data will provide the relevant records from the HES, ECDS, Cancer Registration, Civil Registration of Death, Demographics, Maternity Services, Mental Health and Learning Disabilities, Mental Health Minimum, Community Services, Improving Access to Psychological Therapies, Patient Reported Outcome Measures and National Diabetes Audit datasets to Queen Mary University of London. The data will contain no direct identifying data items. However the data will be identifiable but individuals will not be reidentified through linkage with other data in the possession of the recipient. Participant-level data information from NHS Digital will be stored in a data safe haven (currently UK SeRP (part of Swansea University) along with the study ID and genetic data generated by the Genes & Health study. Approved collaborating researchers will analyse genetic and health data inside UK SeRP (part of Swansea University). Researchers do not have access to identifiable participant information, and will only be able to export summary results (i.e. without participant-level data). Participant level health data from a variety of datasets, along with the unique cohort study ID supplied by Genes & Health, will be transferred to Genes & Health Google Cloud TRE and Swansea University. All export requests are reviewed by a delegated member of the Genes & Health executive committee. Access to each data source is restricted to researchers working on relevant projects. The data will be stored inside The UK Secure eResearch Platform (SeRP) with servers located at Swansea University and Genes & Health Google Cloud Trusted Research Environment, using only servers and storage located in the UK. Researchers may be based in the UK or EEA. To date applications to use Genes and Health data have been received from the EEA, the USA, and Australia. Regardless of a researchers country, the participant level data is stored and processed inside the same Data Safe Haven located in England/Wales, and cannot be downloaded. Researchers from external institutions, including universities and commercial companies, will, with the agreement and oversight of QMUL, analyse Genes & Health data on one of two secure data environments; either inside the UK SeRP platform (provided by Swansea University), or the Trusted Research Environment administered by the Welcome Sanger Institute. UK SeRP is where NHS England data has been stored to date, with servers located at Swansea University Campuses in Wales. Both The Sanger Institute and UK SeRP/Swansea University hold Data Security and Protection Toolkits, and both environments hold ISO27001 certifications. The UK-SeRP platform is user- friendly and suitable for some types of analysis but does not have the capacity for high performance compute needed to analyse large genetic datasets. The second platform, the Trusted Research Environment (TRE) is a secure platform that has been developed by QMUL and Welcome Sanger Institute specifically for Genes & Health, based on the successful platform developed by the similar Finngen consortium (https://www.finngen.fi/en). Full ISO27001 Certification was achieved in March 2023. The TRE is hosted on UK-based Google Cloud servers located in London, England, and the IT administration is provided by Genes and Health funded staff at the Wellcome Sanger Institute. It is the intention of Genes and Health that most data users move to the TRE over the next year. This agreement permits use in the EEA only. Genes & Health is an open access resource, making anonymised genetic and medical record data available in a Trusted Research Environment (aka data safe haven) to external researchers and industrial partners. This project does not require any subsequent flows of data. A Data Access Agreement (DAA) will be signed by external institutions before access to Genes & Health data is granted. A sublicense will also be signed for access to NHS England data. QMUL employees (members of the core Genes & Health team) will organise and adapt NHS Digital data into formats for transfer to the UK SeRP (part of Swansea University) The data will be accessed by authorised personnel via remote access. The data will remain on the systems data controlled by QMUL within the UK. Swansea University who provide UK SeRP will be data processors, responsible for storing the data, and maintaining the platform and tools that are used to analyse the data. Personnel are legally prohibited and not technically capable of downloading or copying data from the Trusted Research Environment to local devices. QMUL employees will maintain and update datasets, and consult and use the data for the purposes of analysis. Once their institution has signed a Data Access Agreement, employees of external institutions will consult and use the data for the purposes of analysis within the scope agreed with QMUL. The data will only be accessed within the UK or the EEA.. Following analyses, summary results (not including identifiable or participant-level data) will be exported for the purposes of publication and dissemination. QMUL employees (delegated Genes & Health team members) will review each export. Only aggregated results/outputs with small number suppression will be used in publications or dissemination. Data processing will only be carried out by substantive employees of Swansea University, Wellcome Sagner Institute and Queen Mary University of London or the Institutions of approved researchers via sublicensing. Genotyping data (all participants), exome sequencing data (some participants) primary and secondary care health record data (all participants), and the requested NHS Digital data will be available for data linkage. Access to each data set is only granted as required, and researchers agree to limit linkage to that described in their project plan approved by the Genes & Health Executive Committee. Researchers may import additional datasets in to the data safe haven for further linkage, once approved by the Genes & Health Executive Committee. All personnel accessing the data have been appropriately trained in data protection and confidentiality. All personnel will have up to date NHS information governance training. Project applications will specify the datasets to be used and linked, including imported datasets. The executive committee will not approve projects where the risk of re-identification through linkage is high. The data will be linked at person record level with NHS Primary Care data, NHS Secondary Care data direct from Secondary Care providers, participant questionnaire data, participant survey data, participant genetic data, and other bio sample data obtained at recall. The access agreement signed by researchers from external institutions will prohibit linkage outside the scope of their approved project proposal. This will be monitored by review of imported datasets, review of exported summary data on each export, and review of progress and publications arising from use of Genes & Health data. There will be no requirement and no attempt to reidentify individuals by using NHS England provided data. On rare occasions projects may wish to export summary results containing a small number of participants possessing a rare genetic variant. On these occasions care will be taken to export the minimum amount of data to avoid identification of participants. For example, researchers will be expected to use an age range rather than age for each participant. All small numbers will be suppressed in line with the HES Analysis guide. The identifying details will be stored in a separate database to the linked dataset used for analysis. All analyses will use the pseudonymised dataset. There will be no requirement and no attempt to reidentify individuals when using the pseudonymised dataset. Any applications to link participant data will be assessed by the Executive Committee in its formal application review process. It is not expected that Genes & Health data will be linked to identifiable publicly available datasets (no applications to do so have been received). Analysts and researchers from Swansea University, Wellcome Sanger Institute, Queen Mary University of London and the Institutions of approved researchers via sublicensing will process/analyse the data for the purposes described above. Researchers may use a combination of genetic and health record data to flag participants for further follow-up. Examples might include a number of individuals with a rare genetic variant of interest, or whose genetic makeup indicates they may be at a high risk of high cholesterol. Researchers, who can see pseodonymised data, will provide study IDs to the Genes & Health team, who are based at QMUL and have access to identifiable details like consent forms and contact details questionnaires, and who will contact participants. This ‘recall’ of volunteers is routine, all participants of Genes & Health have agreed to be contacted, and it will allow better characterisation of volunteers’ health or genetics. Participants’ identifiable details will not be made available to researchers. There will be no attempt to link the data provided by NHS Digital directly to identifiable details. Data processing will only be carried out by substantive employees of Swansea University, QMUL or the Institutions of approved researchers via sublicensing. All employees with access will have been appropriately trained, at a minimum with the e-Learning for Health Data Security Awareness training (or equivalent for institutions that do not have access). The UK Secure eResearch Platform (SeRP), run by Swansea University, will be used to store data and provide an analysis platform. UK SeRP hold a Data Security and Protection Toolkit, and are compliant with ISO27001. Data will be stored at named institutions.

Expected output

Many health research and genetic studies have focused more on other populations (for example, white European origin people), even though South Asian communities see high rates of diabetes, heart disease, rare genetic diseases, and many other conditions. Genes & Health was set up to make sure that advances in genetics and healthcare are available to this underserved community, and to help provide insights into some of the health inequalities that exist within the UK. The ability to recall participants with unique genetic makeup or combination of genetics and health status, also provides a powerful opportunity to better understand how genetics and health are linked, which may be of benefit to everyone. The expected outputs of the processing will be: The results of data processing will be new knowledge related to health and disease in people of British Bangladeshi and British Pakistani origin. The Genes & Health findings will be shared with the wider scientific community via presentations at national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications. Genes & Health has already published several peer-reviewed publications, including a cohort profile (Finer et al, International Journal of Epidemiology, PMID 31504546), and empirical research demonstrating the potential of the bioresource to build new knowledge on health and disease in understudied ethnic groups and improve patient care (McGregor et al, eLife, PMID 26940866; Narasimhan et al, Science, PMID 32207686). Outputs arising from use of the Genes & Health bioresource is likely to generate a significant number of high impact publications over the next 5 years, and preprint servers and open access journals will be used in preference. Study research findings are also summarized on the Genes & Health website (www.genesandhealth.org) and via its twitter account (@eastlondongenes and @bradfordgenes, and Manchester when they are ready to recruit). • Submissions of scientific biomedical publications to peer reviewed journals – ongoing can be found under the following link https://www.genesandhealth.org/about-study/scientific-publications Genes & Health has an established protocol for sharing data outputs, and will continue to use this for outputs arising from NHS Digital datasets. The protocol for sharing data outputs is as follows: • Presentations at appropriate conferences - national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications. Level 1: Fully open data. Genes & Health distribute summary level analysed data on the study website and via twitter. For example summary phenotype counts, e.g. numbers of volunteers with diabetes, or summary statistics of clinical observations such as blood lipid measurements. Small number suppression will be used when the numbers of volunteers is very low and there was a risk of being able to identify individuals or families from these data. The outputs will not contain NHS England data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the dataset(s) from which the information was derived. Level 2: Fully anonymised genetic data is available under a Data Access Agreement with the European Genome- phenome Archive (EGA, https://www.ebi.ac.uk/ega/home). This data will not be linked to any health data, including that obtained from NHS Digital. The outputs will be communicated to relevant recipients through the following dissemination channels: Level 3: Data access and analysis within Data Safe Haven. Data security is central to ELGH (and written into the study Ethics and Governance), a data breach would irretrievably damage the study and the community trust that has been built. Applicants wishing to analyse phenotype data, e.g. NHS Digital health data, do this within an ISO27001 and NHS Information Governance compliant Data Safe Haven environment (currently UK-SERP, which provides a Virtual Desktop Infrastructure on the end user device, with export controls). Data outputs and publications will be summarised and presented in aggregate, and it will not be possible to identify individuals. - Journals All collaborative research involving NHS Digital datasets will require a Data Sharing Agreement to ensure a rigorous approach to data safety and onward use. - Educational workshops and interactive web- and app involving the wider scientific community supported by the QMUL Life Sciences Initiative The applicants will take the following approaches: - Webinars open to members of the Queens Mary University of London. a) Academic dissemination: this has been described above and will include communication of results through peer-reviewed publication and conference presentations, as well as through other clinical and academic networks, relevant professional interest groups and social media. - Social media through twitter found at (@eastlondongenes and @bradfordgenes, and @manchestergenes). b) Stakeholder dissemination: Genes & Health disseminates regular updates regarding study activities (e.g. recruitment, outputs) to a range of stakeholders, including funders, clinicians, policymakers, community representatives and volunteers. This dissemination takes place through its website (which is updated regularly) and by social media. - Public events the study team are guided by NIHR INVOLVE (https://www.invo.org.uk/) in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g., Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families. c) Open science: Genes & Health is committed to open science. It openly shares its methods, data dictionaries, codelists and manuscripts on open access platforms and via its website. - Patient Information leaflets available at www.genesandhealth.org d) Public engagement: The study team are guided by NIHR INVOLVE in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g. Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families. Genes & Health undertake innovative engagement activities with the award-winning Centre of the Cell, to deliver educational workshops and interactive web- and app- content based on the study objectives, and are supported by the QMUL Life Sciences Initiative, QMUL Centre for Public Engagement and a recent Wellcome Trust Public Engagement Award (PI D van Heel, ref 102627/B/13/A). The Helix Champions are also critical to ongoing and effective communication with volunteers to keep them engaged in the study and informed about its outcomes, and to engage with community leaders, religious organisations and schools. The research team have presented on local and national radio (Betar Bangla, BBC Asian Network and World Service) and BBC London TV. Genes & Health has an active Community Advisory Panel who support all ELGH activities through the entire research pathway, from prioritisation of topics for study, to dissemination of research findings. - Press/media engagement The Genes & Health open access policy is described above and ensures that, where possible, there is no restrictive ownership of data or outputs in Genes & Health. It is possible that work on combined genetic and clinical phenotypes could derive significant commercial interest, e.g. to pharmaceutical companies for the development of new drugs, or genomics companies developing new disease risk algorithms. Genes & Health has a rigorous system for reviewing applications to work with its bioresource, involving its Executive team and Community Advisory Group. All collaborations and partnerships involving commercial organisations are pre-competitive and therefore not commercially exploitable. - Participant newsletters Expected targets for outputs are, (a) short-term, e.g. publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g. building and expanding the bioresource within 5 years, and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years). Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are design to educate participants about health and genetics, and to gather public opinions on Genes & Health research. Some examples of current Genes & Health supported research is summarised below to illustrate the scope of the data request applied for: Expected targets for outputs are, (a) short-term, e.g., publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g., building and expanding the bioresource within 5 years and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years). • Type 2 diabetes – this condition disproportionately affects British south Asians and studies are currently taking place/planned with Genes & Health to investigate the influence of common (polygenic risk scores) and rare genetic variants (changes) in type 2 diabetes, misclassification of diabetes types in British south Asians, and gestational diabetes. These studies require detailed longitudinal clinical data from HES, National Diabetes Audit, MHSDS, MSDS, cancer registration, and civil registration of deaths in order to describe the onset of diabetes (including progression from at-risk states such as gestational diabetes), acute diabetes emergencies, glucose control and uptake of diabetes care, its comorbidities (including mental health disorders and cancer), complications and outcomes. Association-based analysis will be used to quantify the relationship between genetic variation and clinical outcomes, using survival analysis to understand trends over time. Adjustment for confounding variables, such as socioeconomic status, will be included in analyses. Some outputs have already been published from previous data received from NHS England and can be found at https://www.genesandhealth.org/about-study/scientific-publications, further outputs are to be expected to be generated continuously until the end of the study. • Familial hypercholesterolaemia (high cholesterol) – this rare genetic condition is a cause of excess death from cardiovascular disease. The condition is under-diagnosed in British south Asians, and work within Genes & Health is identifying volunteers with the condition through their genetic sequence data, and correlating this to clinical data to improve diagnosis and treatment. Longitudinal data from HES and ECDS is required to capture diagnoses and relevant hospital admissions with cardiac emergencies, and long-term outcomes including death from civil registration data. The analysis of health data will be used to design translational Stage 2 recall studies based on genotype and identify volunteers who need specific clinical intervention (e.g. initiation of specific drugs). • Multimorbidity – a large programme of MRC-funded work is currently underway investigating the clustering of multimorbidity in British Bangladeshis and British Pakistanis, and their trajectories across the lifecourse. This research is taking a novel, data-driven approach to identify clusters of multimorbidity across multiple single conditions. The use of multi-source, linked medical record data will increase data quality (e.g. by validating diagnoses across datasets) and the ability to generate novel and meaningful multimorbidity clusters (e.g. encompassing both physical and mental health disorders by using HES and MHSDS/MHMDS/MHLDDS data). Historic data will allow analysis of a patients’ risk from multimorbidity across their lifecourse, and data from HES and MSDS will be used to investigate the impact of specific lifecourse events such as pregnancy, on the development of multimorbidity. • COVID-19 – Genes & Health is contributing to international efforts to identify risk of disease, and its severity, in the host genome. The availability of HES data, including diagnoses from, and episodes in, emergency care (and ECDS), admitted patient care and critical care, and civil registration of deaths will support this work across all study volunteers. These data will be used in genome-wide association studies, with likely subgroup analysis according to disease severity/hospitalisation, and adjustment for confounding variables such as age and socioeconomic status (Index of Multiple Deprivation) • Mental health and dementia – work is underway to better understand the genetic influences on mental health conditions, e.g. depression and anxiety, and dementia. This work will include longitudinal analysis of risk factors for these diseases their diagnosis and severity and associated mortality and therefore requires linked multisource data, including MHSDS, HES, and civil registration of deaths. Analyses will comprise genome-wide association studies, calculation of polygenic risk scores (including assessment of their performance in predicting disease onset). • Discovery analyses of rare genetic variants (changes) – one of the unique features of the Genes & Health study is the ability to investigate the impact of rare genetic variation on health and disease due to its large scale and focus on a population with high rates of parental relatedness. The impact of rare genetic variants on an individual’s health requires careful study, particularly where the genetic variation is novel and its impact is unknown. A discovery-based approach is required to study such genetic variants, as they may have broad health consequences and novel disease associations. Conversely, rare genetic variants may offer protection from disease and the absence of diagnoses and hospital episodes would be highly informative. A proof-of-principal study of a rare genetic variant in the HAO1 gene has shown the importance of such studies using health data: an individual carrying a variant in this important metabolic regulator gene, had no ill effects on their health (determined by medical record data) and this knowledge provided critical information to support the development of a drug that targets this gene in a rare metabolic illness. All diagnoses and details of episodes of care are required to inform this novel genetic discovery-based research and direct subsequent translational research and drug development. • The impact of inherited genetic variants on health – another unique feature of the Genes & Health study is the ability to investigate the impact of rare genetic variants arising from parental relatedness (called autozygous variants), on health. It is known that autozygosity can increase the risk of developmental disorders (the relevance of the Mental Health Minimum Dataset and its inclusion of diagnoses related to learning disability) are critical here. Additionally, there is a recent understanding that autozygosity may impact on a range of long-term conditions and reproductive health/fertility. It is particularly important to study these associations further in British south Asians due to the higher rates of parental relatedness. Discovery-based analyses are planned across a large range of traits and phenotypes to characterise these associations further, and data from HES, NDA, maternity and mental health datasets will be highly informative. • Pregnancy-based studies – QMUL has active research investigating rare and severe pregnancy-based conditions with a known genetic basis (e.g. intrahepatic cholestasis of pregnancy) which will require detailed health record information to identify cases and outcome. Studies are underway within Genes & Health investigating common conditions in pregnancy (such as gestational diabetes or pre-eclampsia) their genetic basis (e.g. polygenic risk) and how this affects future risk of disease (such a progression to type 2 diabetes or cardiovascular disease). Additionally, QMUL have planned research investigating causal associations between maternal genetics offspring traits such as birth weight. All pregnancy based studies will require a core set of data from antenatal care, through the peripartum period, to immediate postnatal care to determine severity and outcome to the affected mother and her child. A limited set of data from offspring (e.g. birth weight, Apgar score, neonatal intensive care admission) has been requested to determine immediate pregnancy outcomes. The above summary is not exhaustive and is to give an overview of the types of research currently funded and being undertaken in Genes & Health in order to justify the broad scope of data included in this application. It also reflects the need to be responsive to the ongoing development and use of Genes & Health as a bioresource that currently receives 2-4 new applications per month for new collaborative research studies. The data request has been designed to be comprehensive and support current and future research with the Genes & Health bioresource, but with the minimum of data fields required to meet its objectives.

Expected measurable benefits

It is anticipated that the potential impact of this research will be broad, and cover a range of domains, including academic, health policy and economic impact, as well as direct impact on the individuals contributing to the research through public engagement. Dissemination will largely bring benefit through the acquisition of new knowledge relating to health and disease in a previously under-represented and understudied population. New interdisciplinary knowledge will be derived from the combined analysis of health and genomic data on a population scale. Detailed health data from NHS Digital will allow Genes & Health users to generate this knowledge for an otherwise under-studied British South Asian population who experience high rates of, and premature, disease. This new knowledge will be highly complementary to other major bioresources, e.g. UK Biobank and the Clinical Practice Research Datalink, and the availability of linked data curated across both will support high quality replication and validation studies thereby maximising outputs. The use of routine health data in this work will support direct and feasible translation of findings back to clinical care in the NHS Additionally, the research will deliver new methodological insights related to the integration of health data science and genomics that could have direct impact to the academic community, including through the training of junior academic researchers. The findings of this research study are expected to contribute to evidence-based decision-making for policymakers, local decision-makers such as doctors, and patients to inform best practice to improve the care, treatment and experience of health care users relevant to the subject matter of the study. Dissemination of new knowledge arising from the Genes & Health bioresource is likely to benefit health and social care through the development of new treatments and risk prediction strategies, as well as finding new causes of disease or its complications. For example, Genes & Health has already generated new knowledge with the potential to benefit health care through its recent work combining information from a rare genetic variant (in the HAO1 gene) with health data (from local health data sources). Dissemination of study findings related to this gene function through academic collaboration has supported critical drug development for a rare, life-threatening metabolic disorder (primary hyperoxaluria). Identification of the effects of genetics on health (including rare gene variants, or gene differences associated with parental relatedness) could improve clinical care through better genetic counselling to families at risk and identification of risk early in the lifecourse. The use of the data could: Longer-term, Genes & Health anticipate that the research will support policy-level improvements in health care with potential economic impact, e.g. through targeting prevention strategies to those at the highest risk, delivering more effective treatments, and reducing health inequalities. - Help the system to better understand the health and care needs of populations. Sharing new knowledge arising from Genes & Health is likely to benefit the public through raised awareness and understanding of health and disease, and improving access and availability to effective treatment. These benefits will relate to British Bangladeshi and British Pakistanis, and may be generalisable to other south Asian groups in the UK or globally. Genes & Health community partners and advisory group will have an active role in dissemination of results to ensure that they are effectively and sensitively disseminated to the communities the study represents. - Lead to the identification or improvement of treatments or interventions, or health and care system design to improve health and care outcomes or experience. All research projects using the Genes & Health bioresource go through an approval process to ensure that their scope is within the remit of the Genes & Health research programme and its objectives. Genes & Health will monitor the outputs of research involving the bioresource via annual reporting to the Executive group. - Advance understanding of the need for, or effectiveness of, preventative health and care measures for particular populations or conditions such as obesity and diabetes. This monitoring will include ensuring that research outputs meet their original objectives and do not extend beyond the scope of their original approval. - Provide a mechanism for checking the quality of care. This could include identifying areas of good practice to learn from, or areas of poorer practice which need to be addressed. - Support knowledge creation or exploratory research (and the innovations and developments that might result from that exploratory work). It is anticipated that the potential impact of this research will be broad, and cover a range of domains, including academic, health policy and economic impact, as well as direct impact on the individuals contributing to the research through public engagement. It is hoped dissemination will largely bring benefit through the acquisition of new knowledge relating to health and disease in a previously under-represented and understudied population. New interdisciplinary knowledge will be derived from the combined analysis of health and genomic data on a population scale. Detailed health data from NHS England should allow Genes & Health users to generate this knowledge for an otherwise under-studied British South Asian population who experience high rates of, and premature, disease. This new knowledge could be highly complementary to other major bioresources, e.g., UK Biobank and the Clinical Practice Research Datalink, and the availability of linked data curated across both could support high quality replication and validation studies thereby hopefully maximising outputs. The use of routine health data in this work could support direct and feasible translation of findings back to clinical care in the NHS. Additionally, the research could deliver new methodological insights related to the integration of health data science and genomics that could have direct impact to the academic community, including through the training of junior academic researchers. Dissemination of new knowledge arising from the Genes & Health bioresource is likely to benefit health and social care through the development of new treatments and risk prediction strategies, as well as finding new causes of disease or its complications. For example, Genes & Health has already generated new knowledge with the potential to benefit health care through its recent work combining information from a rare genetic variant (in the HAO1 gene) with health data (from local health data sources). Dissemination of study findings related to this gene function through academic collaboration has supported critical drug development for a rare, life-threatening metabolic disorder (primary hyperoxaluria). Identification of the effects of genetics on health (including rare gene variants, or gene differences associated with parental relatedness) could improve clinical care through better genetic counselling to families at risk and identification of risk early in the life course. Longer-term, Genes & Health anticipate that the research will support policy-level improvements in health care with potential economic impact, e.g., through targeting prevention strategies to those at the highest risk, delivering more effective treatments, and reducing health inequalities. Sharing new knowledge arising from Genes & Health is likely to benefit the public through raised awareness and understanding of health and disease and improving access and availability to effective treatment. These benefits will relate to British Bangladeshi and British Pakistanis and may be generalisable to other south Asian groups in the UK or globally. Genes & Health community partners and advisory group will have an active role in dissemination of results to ensure that they are effectively and sensitively disseminated to the communities the study represents. [3 paragraphs unchanged] PhD students may be part of the research teams working with Genes & Health, both from within QMUL and from other academic organisations. All PhD students will have appropriate supervision, employment contracts with host research institutions, and have undergone relevant information governance training. It is hoped that through publication of findings in appropriate media, the findings of this research will add to the body of evidence that is considered by the bodies, organisations and individual care practitioners charged with making policy decisions for or within the NHS or treatment decisions in relation to specific patients. Genes & Health have recently employed a Communications and Engagement Manger, who will ensure that, in addition to research and NHS groups, local government and third sector organisations are involved in the development and dissemination of the research programme. For example, charities have been involved in a recent finding from Genes & Health of a link between autozygosity and type 2 diabetes (Asthma UK, Diabetes UK) to advise on lay summary and publicity.

Benefits reported

Genes and Health received NHS Digital England data in Dec 2021. In the few months that they had access to the data they have been able to incorporate these data with health data from other sources, such as local NHS Trusts and primary care providers as a source of quality control and additional geographic coverage. These integrated datasets are difficult and time-consuming to produce, produce but mean that Genes & Health has very high-quality health data that will support research projects. This linkage is within the scope of the linkages explained within this DSA DSA. An example of how these improved health datasets are already being used is QMUL's contribution to the international INTERVENE project (https://www.interveneproject.eu/) which is using AI Artificial Intelligence and machine learning to understand the ways in which genetics contributes to [30 words unchanged] INTERVENE uses Genes & Health data, including health data that contains NHS Digital, England, to develop scores that can help doctors work out how likely it [44 words unchanged] and diverse datasets like INTERVENE will perform better and improve health further. Analyses of data from the Genes & Health cohort have resulted in multiple high-impact publications. An up-to-date list is maintained on the study website (https://www.genesandhealth.org/about-study/scientific-publications) and a short summary of recent publications are included here: Genes & Health continues to provide practical return of health and genetic data to participants and the community. For example, we Genes & Health have identified participants who are at a high risk of diseases like inherited high cholesterol - based on a combination of genetic data and health data (including those from NHS England) - and invited those participants to additional screening and referral for treatment. Some details are here: https://www.genesandhealth.org/research/research-studies-approved/s00013-genotype-first-recall-study-extreme-genetic-risk-atherosclerotic We also run regular workshops that aim to improving peoples' understanding of genetics and their own health in order to make informed health choices and improve uptake of existing healthcare pathways (e.g. https://www.genesandhealth.org/research-studies-approved/s00052-improving-genetics-communication). Whole genome sequencing reveals host factors underlying critical Covid-19 (Genes & Health contributed as one of many international cohorts to the multiple phenotype analyses of the COVID-19 Host Genetics Initiative release 6): Analyses of data from the Genes & Health cohort have resulted in multiple high-impact publications. An up-to-date list is maintained on the study website (https://www.genesandhealth.org/about-study/scientific-publications). Nature 2022. DOI https://doi.org/10.1038/s41586-022-04576-6 Transferability of genetic loci and polygenic scores for cardiometabolic traits in British Pakistanis and Bangladeshis: medRxiv 2021 https://www.medrxiv.org/content/10.1101/2021.06.22.21259323v1 Harnessing the power of polygenic risk scores to predict type 2 diabetes and its subtypes in a high-risk population of British Pakistanis and Bangladeshis in a routine healthcare setting: medRxiv 2021 https://www.medrxiv.org/content/10.1101/2021.07.12.21259837v1 Global Biobank Meta-analysis Initiative: powering genetic discovery across human diseases: medRxiv 2021 https://doi.org/10.1101/2021.11.19.21266436 The power of genetic diversity in genome-wide association studies of lipids: Nature 2021. DOI https://doi.org/10.1038/s41586-021-04064-3 MC3R links nutritional state to childhood growth and the timing of puberty: Nature 2021. DOI https://doi.org/10.1038/s41586-021-04088-9 Mapping the human genetic architecture of COVID-19: Nature 2021 https://doi.org/10.1038/s41586-021-03767-x

Objective for processing

Queen Mary University of London (QMUL) requires access to NHS England data for the purpose of the following research project: Genes and Health Study.

The following is a summary of the aims of the research project provided by Genes and Health Study on behalf of Queen Mary University of London.

Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis.

QMUL seeks to combine data from NHS England with data from the Genes and Health Bioresource. This linked resource will be made available via sub-licensing arrangements. The strategic objective of Genes and Health is to develop and maintain a Bioresource of genetic and health record data available to the research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin through high quality research. Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their life course. The Genes & Health study team at QMUL have built a database to investigate genetic data and health data from Genes & Health participants, of which includes data from NHS England under this Data Sharing Agreement. This database will be used by researchers at QMUL, as well as researchers in other universities, charities, and commercial companies.”

Genes & Health will charge commercial partners and academic partners for use of the bioresource, but this will cover infrastructure and sustainability costs only and will not be profit-making.

The following NHS England data will be accessed:

Hospital Episode Statistics (HES) Admitted Patient Care – necessary because the data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices. This dataset will support research into ill health and disease identified, diagnosed, and treated during hospital admissions, as well as episodes of pregnancy care.

(HES) Accident & Emergency / Emergency Care Data Set (ECDS) – necessary because this dataset will support research into acute presentations of ill health, e.g., heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

(HES) Critical Care – necessary because the dataset will support research into severe ill health/disease, specifically related to critical care in this context. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g., in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity.

(HES) Outpatients – necessary because the dataset will support research into ill health and disease, e.g., depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care.

Mental Health Service Dataset (including its predecessors the Mental Health Minimum Data Set and Mental Health and Learning Disabilities Data Set) – necessary because these datasets will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multi morbidity, specifically to characterise states of mental ill health and disease, the care and treatment received.

Civil Registration Mortality – necessary because this dataset will support research into ill health and disease by identifying death and cause of death, specifically to characterise states of ill health, disease, and their long-term outcomes, in this case death and cause of death.

Cancer Registration – necessary because this will supplement existing sources and aid the analyses of the links between genetics and cancer.

Demographics – necessary to mirror those collected independently by Genes & Health during follow-up contacts with participants, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

National Diabetes Audit – necessary because this dataset will support extensive and high-quality research into diabetes and its related co-morbidities and multi morbidity. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment, and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

Improving Access to Psychological Therapies (IAPT), Mental Health and Learning Disabilities and Community Services – necessary because these datasets are hoped to improve the completeness of data that is often under reported from other sources and allow a more accurate picture of community service use, particularly in areas like mental health.

Maternity Services Data Set (MSDS) – necessary because this dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC.

The MSDS data s relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g., body mass index at booking, or social factors) will also add valuable detail to the characterisation of the female study volunteers more broadly, for cross-sectional and longitudinal studies.

Patient Reported Outcome Measures (Linkable to HES) - These are required to mirror those collected independently by Genes & Health during follow-up contacts with participants, and those collected by UK Biobank, along with the newly requested demographics dataset these data will allow better analyses of population genomics and health trends.

Community Services Data Set - necessary because it provides important supplementary information to other datasets, including social and demographic information and health care utilisation information.

The level of the data will be identifiable.

Data returned is subsequently linked with other data directly provided by Primary Care providers, other data directly provided by Secondary Care Providers, NHS Trusts, Questionnaire Data, Survey Data and Genetic data.

The data will be minimised as follows:

• Limited to a cohort of 65,000 patients (as of 2023) who consented to participate.

• Limited to data between 1997/98 to latest available.

Genes & Health Volunteers have given written informed consent for their NHS health data to be accessed stored and processed in line with the study objectives. Volunteers also gave consent for linkage to their medical records, including from their GP and hospital records as well as from nationally held datasets such as those included in this agreement, for research purposes.

Genes & Health recruitment started in 2015 in East London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited approx. 65,000 volunteers and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on East London, Luton, Bradford and Manchester). The cohort is expected to reach 100,000 volunteers.

Queen Mary University of London is the research sponsor and the controller as the organisation responsible for ensuring that the data will only be processed for the purpose described above.

The lawful basis for processing personal data under the UK GDPR is:

Article 6(1)(e) - processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller.

The lawful basis for processing special category data under the UK GDPR is:

Article 9(2)(j) - processing is necessary for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes in accordance with Article 89(1) based on Union or Member State law which shall be proportionate to the aim pursued, respect the essence of the right to data protection and provide for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

This processing is in the public interest as British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, and this resource looks to maximise its potential to deliver benefit to population health to these two groups.

The funding comes from multiple sources. Current funders include: Wellcome Trust, Medical Research Council, Higher Education Funding Council for England Catalyst, Barts Charity, Health Data Research UK (for London substantive site), and research delivery support from the NHS National Institute for Health Research Clinical Research Network (North Thames), Alnylam Pharmaceuticals, Genomics PLC; and a Life Sciences Industry Consortium of Astra Zeneca PLC, Bristol-Myers Squibb Company, GlaxoSmithKline Research and Development Limited, Maze Therapeutics Inc, Merck Sharp & Dohme LLC, Novo Nordisk A/S, Pfizer Inc, Takeda Development Centre Americas Inc. Funding is in place until 2028. Funding to continue the work described will be sought on an ongoing basis.

Swansea University, the Wellcome Sanger Institute and Google UK Limited are processors acting under the instructions of Queen Mary University of London. Swansea University’s and Wellcome Sanger Institute’s roles are limited to ensuring that the data will only be processed for the purpose described above. Google UK Limited’s role is to provide a storage and data processing platform via Cloud. Employees of Google UK Limited do not have access to the Genes & Health or NHS England data.

Several organisations have provided data or services but are not processors of the NHS England data. These are Kings College London and UK Biocentre, who have processed DNA samples from participants. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership. Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data.

Genes & Health has developed an industry consortium with partners across a range of pharmaceutical companies to support research that maximises the translation from basic science to therapeutics and clinical care. Consortium partners are Pfizer, GSK, Merck, Novo Nordisk, Maze Theraputics, BMS, and Takeda. Consortium partners are working with the Genes & Health bioresource, under rigorous regulations set out in a Consortium Agreement. At present consortium partners are using genetic and non-NHS England health datasets. If partners use individual level NHS England data in the future, QMUL will require these partners to sign up to the sublicense agreements. These agreements and procedures will ensure that research using the data in the Genes & Health bioresource, are used to meet the objectives of the study, that appropriate levels of information governance and UK GDPR compliance are upheld, and that commercial exploitation is prevented.

A Public and Patient Involvement and Engagement group helped refine the purpose of the research. The group strongly supported the collection of the data for the purposes described above. The Genes & Health Community Advisory Group meets 3-4 times a year and provides advice by email as required for other issues.

Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are designed to educate participants about health and genetics, and to gather public opinions on Genes & Health research.

Expected output

The expected outputs of the processing will be:

• Submissions of scientific biomedical publications to peer reviewed journals – ongoing can be found under the following link https://www.genesandhealth.org/about-study/scientific-publications

• Presentations at appropriate conferences - national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications.

The outputs will not contain NHS England data and will only contain aggregated information with small numbers suppressed as appropriate in line with the relevant disclosure rules for the dataset(s) from which the information was derived.

The outputs will be communicated to relevant recipients through the following dissemination channels:

- Journals

- Educational workshops and interactive web- and app involving the wider scientific community supported by the QMUL Life Sciences Initiative

- Webinars open to members of the Queens Mary University of London.

- Social media through twitter found at (@eastlondongenes and @bradfordgenes, and @manchestergenes).

- Public events the study team are guided by NIHR INVOLVE (https://www.invo.org.uk/) in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g., Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families.

- Patient Information leaflets available at www.genesandhealth.org

- Press/media engagement

- Participant newsletters

Participants are regularly invited to participate in workshops, focus groups, and informal information sessions. These sessions are design to educate participants about health and genetics, and to gather public opinions on Genes & Health research.

Expected targets for outputs are, (a) short-term, e.g., publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g., building and expanding the bioresource within 5 years and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years).

Some outputs have already been published from previous data received from NHS England and can be found at https://www.genesandhealth.org/about-study/scientific-publications, further outputs are to be expected to be generated continuously until the end of the study.

Benefits reported

Genes and Health received NHS England data in Dec 2021. These integrated datasets are difficult and time-consuming to produce but mean that Genes & Health has very high-quality health data that will support research projects. This linkage is within the scope of the linkages explained within this DSA.

An example of how these improved health datasets are already being used is QMUL's contribution to the international INTERVENE project (https://www.interveneproject.eu/) which is using Artificial Intelligence and machine learning to understand the ways in which genetics contributes to disease risk. Genes & Health’s contribution to INTERVENE is really important because they provide data from people of Pakistani and Bangladeshi backgrounds, which would otherwise be underrepresented in the project. INTERVENE uses Genes & Health data, including health data that contains NHS England, to develop scores that can help doctors work out how likely it is that a patient will develop a disease. The new generation of genetic risk scores will help to diagnose and treat disease earlier, and so improve health. Genetic risk scores are in trials with the NHS, and new generations of scores built on high-quality and diverse datasets like INTERVENE will perform better and improve health further.

Genes & Health continues to provide practical return of health and genetic data to participants and the community. For example, we Genes & Health have identified participants who are at a high risk of diseases like inherited high cholesterol - based on a combination of genetic data and health data (including those from NHS England) - and invited those participants to additional screening and referral for treatment. Some details are here: https://www.genesandhealth.org/research/research-studies-approved/s00013-genotype-first-recall-study-extreme-genetic-risk-atherosclerotic We also run regular workshops that aim to improving peoples' understanding of genetics and their own health in order to make informed health choices and improve uptake of existing healthcare pathways (e.g. https://www.genesandhealth.org/research-studies-approved/s00052-improving-genetics-communication).

Analyses of data from the Genes & Health cohort have resulted in multiple high-impact publications. An up-to-date list is maintained on the study website (https://www.genesandhealth.org/about-study/scientific-publications).

DARS-NIC-338864-B3Z3J-v1.2 8 July 2022 to 7 July 2023
Title
Genes and Health
Commercial
Yes
Sublicensing
Yes
Datasets
12
Files released
0

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Civil Registrations of Death; Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Maternity Services Data Set (MSDS) v1.5; Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); National Diabetes Audit

What changed from DARS-NIC-338864-B3Z3J-v0.12

Text removed is struck through; text added is underlined. Unchanged paragraphs are summarised rather than repeated.

Fields changed from DARS-NIC-338864-B3Z3J-v0.12
FieldWasBecame
Start date2021-07-092022-07-08
End date2022-07-082023-07-07

Expected output

Many health research and genetic studies have focused more on other populations [8 words unchanged] South Asian communities see high rates of diabetes, heart disease, rare genetic diseases,and diseases, and many other conditions. Genes & Health was set up to make sure [53 words unchanged] genetics and health are linked, which may be of benefit to everyone. [23 paragraphs unchanged]

Benefits reported

Yielded Benefits is not a requirement for new applications. Genes and Health received NHS Digital data in Dec 2021. In the few months that they had access to the data they have been able to incorporate these data with health data from other sources, such as local NHS Trusts and primary care providers as a source of quality control and additional geographic coverage. These integrated datasets are difficult and time-consuming to produce, but mean that Genes & Health has very high-quality health data that will support research projects. This linkage is within the scope of the linkages explained within this DSA An example of how these improved health datasets are already being used is QMUL's contribution to the international INTERVENE project (https://www.interveneproject.eu/) which is using AI and machine learning to understand the ways in which genetics contributes to disease risk. Genes & Health’s contribution to INTERVENE is really important because they provide data from people of Pakistani and Bangladeshi backgrounds, which would otherwise be underrepresented in the project. INTERVENE uses Genes & Health data, including health data that contains NHS Digital, to develop scores that can help doctors work out how likely it is that a patient will develop a disease. The new generation of genetic risk scores will help to diagnose and treat disease earlier, and so improve health. Genetic risk scores are in trials with the NHS, and new generations of scores built on high-quality and diverse datasets like INTERVENE will perform better and improve health further. Analyses of data from the Genes & Health cohort have resulted in multiple high-impact publications. An up-to-date list is maintained on the study website (https://www.genesandhealth.org/about-study/scientific-publications) and a short summary of recent publications are included here: Whole genome sequencing reveals host factors underlying critical Covid-19 (Genes & Health contributed as one of many international cohorts to the multiple phenotype analyses of the COVID-19 Host Genetics Initiative release 6): Nature 2022. DOI https://doi.org/10.1038/s41586-022-04576-6 Transferability of genetic loci and polygenic scores for cardiometabolic traits in British Pakistanis and Bangladeshis: medRxiv 2021 https://www.medrxiv.org/content/10.1101/2021.06.22.21259323v1 Harnessing the power of polygenic risk scores to predict type 2 diabetes and its subtypes in a high-risk population of British Pakistanis and Bangladeshis in a routine healthcare setting: medRxiv 2021 https://www.medrxiv.org/content/10.1101/2021.07.12.21259837v1 Global Biobank Meta-analysis Initiative: powering genetic discovery across human diseases: medRxiv 2021 https://doi.org/10.1101/2021.11.19.21266436 The power of genetic diversity in genome-wide association studies of lipids: Nature 2021. DOI https://doi.org/10.1038/s41586-021-04064-3 MC3R links nutritional state to childhood growth and the timing of puberty: Nature 2021. DOI https://doi.org/10.1038/s41586-021-04088-9 Mapping the human genetic architecture of COVID-19: Nature 2021 https://doi.org/10.1038/s41586-021-03767-x

Unchanged: Objective for processing, Processing activities, Expected measurable benefits.

Objective for processing

This data request comes from the Genes & Health Study, which is based at Queen Mary University of London (QMUL).

Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis.

QMUL wish to combine data from NHS Digital with data from the Genes and Health BioResource. This linked resource will be made available via sub-licensing arrangements.

The strategic objective of Genes & Health is to develop and maintain a BioResource of genetic and health record data available to the research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin through high quality research.

Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their lifecourse. The Genes & Health study team at QMUL have built a database to investigate genetic data and health data from Genes & Health participants, and hope to include data from NHS Digital. This database will be used by researchers at QMUL, as well as researchers in other universities, charities, and commercial companies, once they have been approved by an application review committee, and signed a sublicensing agreement.

Genes & Health Volunteers have given written informed consent for their health data to be accessed stored and processed in line with the study objectives. The study is organised as follows:

(i) Stage 1 research: genetic information from saliva samples provided at recruitment by Genes & Health volunteers is combined with health record data to characterise genetic variation and correlate this to physical characteristics, and states of health and disease.

(ii) Stage 2 research; information from genetic studies and/or health record information are used to guide focused translational research studies, e.g. investigating how a rare gene variant might impact on health.

The article 6 justification for processing the requested data is stated in 1. (e) “processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller”

This data request is in the public interest under Article 9(2)(j). Specifically, it is in the public interest for scientific and statistical purposes in accordance with Article 89(1). Furthermore, the request is proportionate to the aim pursued, respects the essence of the right to data protection, and provides for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

Dissemination of results arising from this data request, including those published from sublicenced data, will be fully anonymised, and will inform study volunteers and the wider public of the research findings at summary level. These findings will comprise new knowledge of health and disease and will be disseminated using lay language. Where findings might be sensitive, e.g. reporting an increased risk of a severe disease in the study population, or its genetic cause, the Community Advisory Group will be consulted to ensure effective and appropriate means of communication.

Volunteers give consent for linkage to their medical records, including from their GP and hospital records as well as from nationally-held datasets such as those included in this agreement, for research purposes. Genes & Health is a multi-purpose bioresource, and combines health data with genetic data (obtained from saliva samples collected at recruitment) to undertake detailed characterisation of states of health and disease and generate new knowledge relating to these. Genes & Health recruitment started in 2015 in east London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited over 47,000 volunteers and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on east London, Luton, Bradford and Manchester). The cohort is expected to reach 100,000 volunteers. Linkage to national datasets is required to (a) expand the geographical coverage of data linkage where local datasets are not available or where healthcare is accessed beyond its limits, and (b) to enrich the datasets available using high quality multisource national data that are not available through local health systems, e.g. civil registration of deaths.

Genes & Health receives funding from leading funders including the Wellcome Trust, Medical Research Council, National Institute for Health Research Clinical Research Networks and others. People of British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, including UK Biobank. Genes & Health is an open access resource, making anonymised genetic and medical record data available in a data safe haven to external researchers and industrial partners to deliver its aims and maximise its potential to deliver benefit to population health. The infrastructure of Genes & Health has, at its core, rigorous protocols to ensure the appropriate use of data, data safety and volunteer acceptability.

The study complements, but is distinct from, existing bioresources (e.g. the well-known Cambridge BioResource and UK Biobank) by its focus on two distinct UK ethnic minority groups with a different spectrum of genetic variation, high levels of health deprivation, and substantial parental relatedness. As such, Genes & Health will deliver new knowledge relevant to an understudied population at high risk of ill health. Genes & Health also complements existing large-scale open access health data resources, e.g. the Clinical Practice Research Datalink, by its linkage of health data to genetic data.

The research taking place within Genes & Health reflects community prioritisation, health need, research expertise and emerging knowledge. Furthermore, Genes & Health is an adaptable and responsive resource that can accommodate changing priority and need, an example being its recent rapid contribution to the International COVID-19 Host Genetic Initiative.

Use of data from key national health datasets will support the multipurpose, bioresource function of Genes & Health by expanding the depth and breadth of data available and by increasing the availability of health data on volunteers receiving care outside of local health systems where linkage is already in place. The scope of the data request outlined in this application is broad, and reflects the design and need of Genes & Health as a bioresource supporting a large portfolio of research projects in Stage 1 and Stage 2 studies:

(a) Stage 1 research: genetic information from saliva samples provided at recruitment by Genes & Health volunteers is combined with health record data to characterise genetic variation and correlate this to biomedical traits and phenotypes, including states of health and disease. These studies generally require population-based analysis of association (e.g. in case-control studies of the genetic associations of prevalent disease, or the prediction of disease onset using risk scores) or longitudinal analysis of disease progression (e.g. using survival analysis incorporating risk factors, incident disease, and progression to complications and death). All analyses require careful adjustment for confounding variables (e.g. socioeconomic status, comorbidity/multimorbidity, smoking) and use of time-dependent covariates (e.g. pregnancy, progression from at-risk states).

(b) Stage 2 research; information from genetic studies and/or health record information are used to guide focused translational research studies, e.g. investigating how a rare gene variant might impact on health. These studies are typically smaller in scale compared to Stage 1 studies, and might involve a small number of individuals (e.g. 5-100) carrying a rare genetic variant whose function is unknown. In such examples, health record data will be used to identify novel associations between genotype and health/disease status.

Genes & Health invites the following people to participate:

• People living within reach of its research sites (currently east London, Luton, Bradford and Manchester)

• Self described ethnicity: Bangladeshi, British-Bangladeshi, Pakistani, British-Pakistani

• Age 16 or over (no upper limit)

• Willingness to give a saliva DNA sample for analysis

• Willingness to grant investigators access to medical records (GP, Hospital, and national NHS) for the full duration (stage 1 and stage 2) of the study.

• Willingness to receive invitations for recall (stage 2 research activity) for the duration of the study (actual participation in all stage 2 research activity would be after a second informed consent process).

QMUL seek linkage of NHS Digital datasets at record level to all Genes & Health volunteers. This agreement is requesting pseudonymised, record-level data.

Genes & Health collects saliva samples from all participants, which QMUL use to generate genetic data like SNP genotyping and DNA sequencing. From a select group of volunteers QMUL also collect samples like blood and urine, and make clinical measurements like height and weight. To understand how a person's genetics interact with their health, to cause bad health or to protect against disease, researchers need to be able to analyse genetic data from Genes & Health alongside health data from NHS Digital – one set of data alone won’t help understand the connection between genetics and health.

The partners who will use the Genes & Health resource include academic research groups at UK universities (the vast majority of users) and at universities in the EEA. Partners will also include technology and pharmaceutical companies based in the UK, and EEA, who want to access Genes & Health data for research and development purposes. All partners, academic and commercial, are required to submit an application for a defined project for review, have a summary of their project published online, and making appropriate results of their work available to the research community. To date 45 applications have been received, of which only three were from commercial companies.

Carefully curated, and minimised, data from the following resources are requested in this agreement:

¥ Hospital Episode Statistics (HES), including HES A&E, OP, critical care and APC.

¥ HES APC:

This dataset will support research into ill health and disease identified, diagnosed and treated during hospital admissions, as well as episodes of pregnancy care. The data obtained from HES APC will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. A number of data fields have been requested in addition to STUDY_ID to ensure adequate linkage between the Genes & Health ID and pseudonymised HES identifiers across HES datasets,

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and pregnancy. This data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

¥ HES OP:

This dataset will support research into ill health and disease, e.g. depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. The data obtained from HES OP will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved), discharge and outpatient follow up.

This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care.

¥ HES AE:

This dataset will support research into acute presentations of ill health, e.g. heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. The data obtained from HES A&E will complement the other HES datasets that are being requested also, giving QMUL health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. A number of data fields have been requested in addition to STUDY_ID to ensure adequate linkage between the study ID and pseudonymised HES identifiers across HES datasets,

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

¥ HES CC

This dataset will support research into severe ill health/disease, specifically related to critical care in this context. The data obtained from HES Critical Care will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g. in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity.

¥ Emergency Care Data Set (ECDS)/

This dataset will support research into acute presentations of ill health, e.g. heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. The data obtained from ECDS will complement HES A&E and other HES datasets that are being requested also, to give health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

¥ The Mental Health Services Data Set (MHSDS), and its predecessors, the Mental Health Minimum Data Set (MHMDS) and Mental Health and Learning Disabilities Data Set (MHLDDS)

¥ MHSDS: This dataset will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multimorbidity.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease, including those related to mental health. This data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

¥ MHMDS/MHLDDS

This dataset will support research into mental ill health in the Genes and Health volunteers, including the study of rare and common disease, and their genetic basis.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of mental ill health and disease, the care and treatment received.

¥ Maternity Service Data Set (MSDS)

This dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g. body mass index at booking, or social factors) will also add valuable detail to the characterisation of our female study volunteers more broadly, for cross-sectional and longitudinal studies.

¥ Civil Registrations (Deaths)

This dataset will support research into ill health and disease by identifying death and cause of death.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and their long-term outcomes, in this case death and cause of death.

¥ National Diabetes Audit

This dataset will support extensive and high quality research into diabetes and its related co-morbidities and multimorbidity.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and pregnancy. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

¥ Bridging files HES:civil registrations, HES:MHSDS

To allow the datasets to be abridged.

Use of data from key national health datasets will support the multipurpose, bioresource function of Genes & Health by expanding the depth and breadth of data available and by increasing the availability of health data on volunteers receiving care outside of local health systems where linkage is already in place. The content of data analysis has also been designed carefully to mirror the NHS Digital data available via UK Biobank and Clinical Practice Research Datalink so that replication and validation studies can be performed across all of these valuable resources. This mirroring is to allow similar analyses to be replicated between datasets, no direct linkage of BioBank, CPRD and Genes & Health data is anticipated.

The data requested is required for the following types of analyses:

(a) Association-based genetic analysis of prevalent disease(s), its severity and characteristics

(b) Prediction of disease risk and severity using polygenic risk scores

(c) Longitudinal analysis of states of health and disease(s) through the lifecourse (including pregnancy), and their association with genetic factors

(d) Discovery-based analysis of the impact of rare genetic variation on health and disease

(e) Assessment of, and adjustment for, confounding influences of the above using covariates.

(f) Assessment of causal relationships in health and disease using genetic data in Mendelian randomisation

(g) Validation studies using multi-source health data to quality control self-report data (e.g. the study questionnaire and local health data) and data from other bioresources (e.g. UK Biobank, Clinical Practice Research Datalink)

(h) Replication studies of known genetic factors on health and disease in an understudied population of British Pakistanis and British Bangladeshis.

The study incorporates longitudinal analysis of states of health and disease and includes the study of historic diagnoses. For example, an admission with myocardial infarction in 1999, is important as it will show the earliest time a patient was diagnosed. This will be vital for understanding early onset disease.

QMUL also have ongoing funding and recruitment to 2023 at least and expect Genes & Health to serve as a long-term bioresource in the future, therefore QMUL will be requesting an annual update to this data linkage to update previously linked records and add linkage to newly-recruited volunteers. The study is recruiting nationally across multiple sites (currently London, Luton, Bradford and Manchester), with the possibility of opening additional sites in the future. It is important for Genes & Health to be able to link to national data in order to capture health data on volunteers who move house.

There are no alternative data sources that would achieve the full aims and objectives of the Genes & Health study. Local health system data has already been used to its maximum potential, but this is limited to those systems able to supply data, and the data returned are limited to patients treated within those local systems. Currently this includes Barts Health, and a small number of local CCGs who have signed data transfer agreements, but not primary or secondary care providers further afield. Genes & Health is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on east London, Luton and Bradford) and linkage to national datasets is required to (a) expand the geographical coverage of data linkage where local datasets are not available or where healthcare is accessed beyond its limits, and (b) to enrich the datasets available using high quality multisource national data that are not available through local health systems, e.g. civil registration of deaths.

Although data has been requested from multiple NHS Digital datasets, stringent variable selection has been undertaken for the purposes of data minimisation. Genes & Health have taken the following steps to ensure that the variables requested are strictly necessary and achieve data minimisation:

1) Detailed curation of required variables to meet the Genes & Health study objectives, achieved through discussion and consensus with the core Genes & Health research and Executive teams. The focus of this data request is on diagnoses, clinical observations and related dates, rather than organisational information that does not have a clear purpose in the research objectives.

2) Non-replication of data already held – Genes & Health already hold data, such as full postcode, particularly where it is identifiable or sensitive. A limited set of non-identifiable data (year and month of birth, gender, postcode district and ethnicity) has been requested as a quality control check across datasets.

3) Genes & Health have followed previous large-scale bioresource studies who have used NHS Digital datasets to guide the choice of variables, including UK Biobank and the INTERVAL and COMPARE studies of blood donors, as well as the Clinical Practice Research Datalink. This data request mirrors the set of variables available in these studies that are relevant to the Genes & Health study objectives, allowing for cross-study validation and replication. This will significantly expand the utility of the data requested.

4) Genes & Health have taken extensive and detailed guidance from the DARS team (case officer and data production team member) to guide the selection of variables, in particular where multiple datasets exist for the same clinical areas (e.g. MHSDS, MHLDDS and MHMDS).

5) The request is also minimised to the size of the cohort, currently a little over 47,000 participants. Genes & Health is still actively recruiting, and the number of participants is expected to reach 100,000 by approximately 2023.

QMUL is the sole data controller for Genes and Health data, including these data requested from NHS Digital. Swansea University - who provide the UK Secure eResearch Platform (UK SeRP), will store the data and are a data processor.

Researchers from external institutions, including universities and commercial companies, will, with the agreement, direction and oversight of QMUL, analyse Genes & Health data on a platform provided by UK SeRP (part of Swansea University). This will be facilitated by sublicensing agreements with all organisations. An access register will detail these projects and organisations, and a summary of all projects will be published on the Genes & Health website.

A Data Access Agreement (DAA) will be signed by external institutions before access to Genes & Health data is granted. This agreement sets out the scope of data processing activities that may be undertaken and the legal and technical data protection requirements. To try and reinforce good practice, individual researchers are also asked to sign, having read a letter summarising their data protection responsibilities in the project.

Several organisations have provided data or services but are not data processors for the NHS Digital data. These are King’s College London and UK Biocentre, who have processed DNA samples from participants, and Wellcome Sanger Institute who have analysed DNA samples. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership.

Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data.

Sub-licenses will be granted upon successful application, where applicants satisfy section viii of the terms of reference – which includes data security, public benefit, impact on participants. Applications covering a sensitive topic (for example sexual health, consanguinity, mental health) will be reviewed by a community panel.

Conditions will be compliance with the terms of the agreement, which include strict data security requirements imposed by our research environment and to only undertake work within in the remit of their application. QMUL are able to view and audit the work of collaborators, and as the environment is export controlled QMUL will actively review every request for data-out to make sure they are complying with the terms.

Data disseminated by NHS Digital data is potentially identifiable in the hands of the Data Controller as QMUL have the ability to link the Study ID's to the identifiers. Data will, however, always be pseudonymised in the hands of the sub-licensees.

Expected output

Many health research and genetic studies have focused more on other populations (for example, white European origin people), even though South Asian communities see high rates of diabetes, heart disease, rare genetic diseases, and many other conditions. Genes & Health was set up to make sure that advances in genetics and healthcare are available to this underserved community, and to help provide insights into some of the health inequalities that exist within the UK. The ability to recall participants with unique genetic makeup or combination of genetics and health status, also provides a powerful opportunity to better understand how genetics and health are linked, which may be of benefit to everyone.

The results of data processing will be new knowledge related to health and disease in people of British Bangladeshi and British Pakistani origin. The Genes & Health findings will be shared with the wider scientific community via presentations at national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications. Genes & Health has already published several peer-reviewed publications, including a cohort profile (Finer et al, International Journal of Epidemiology, PMID 31504546), and empirical research demonstrating the potential of the bioresource to build new knowledge on health and disease in understudied ethnic groups and improve patient care (McGregor et al, eLife, PMID 26940866; Narasimhan et al, Science, PMID 32207686). Outputs arising from use of the Genes & Health bioresource is likely to generate a significant number of high impact publications over the next 5 years, and preprint servers and open access journals will be used in preference. Study research findings are also summarized on the Genes & Health website (www.genesandhealth.org) and via its twitter account (@eastlondongenes and @bradfordgenes, and Manchester when they are ready to recruit).

Genes & Health has an established protocol for sharing data outputs, and will continue to use this for outputs arising from NHS Digital datasets. The protocol for sharing data outputs is as follows:

Level 1: Fully open data. Genes & Health distribute summary level analysed data on the study website and via twitter. For example summary phenotype counts, e.g. numbers of volunteers with diabetes, or summary statistics of clinical observations such as blood lipid measurements. Small number suppression will be used when the numbers of volunteers is very low and there was a risk of being able to identify individuals or families from these data.

Level 2: Fully anonymised genetic data is available under a Data Access Agreement with the European Genome- phenome Archive (EGA, https://www.ebi.ac.uk/ega/home). This data will not be linked to any health data, including that obtained from NHS Digital.

Level 3: Data access and analysis within Data Safe Haven. Data security is central to ELGH (and written into the study Ethics and Governance), a data breach would irretrievably damage the study and the community trust that has been built. Applicants wishing to analyse phenotype data, e.g. NHS Digital health data, do this within an ISO27001 and NHS Information Governance compliant Data Safe Haven environment (currently UK-SERP, which provides a Virtual Desktop Infrastructure on the end user device, with export controls). Data outputs and publications will be summarised and presented in aggregate, and it will not be possible to identify individuals.

All collaborative research involving NHS Digital datasets will require a Data Sharing Agreement to ensure a rigorous approach to data safety and onward use.

The applicants will take the following approaches:

a) Academic dissemination: this has been described above and will include communication of results through peer-reviewed publication and conference presentations, as well as through other clinical and academic networks, relevant professional interest groups and social media.

b) Stakeholder dissemination: Genes & Health disseminates regular updates regarding study activities (e.g. recruitment, outputs) to a range of stakeholders, including funders, clinicians, policymakers, community representatives and volunteers. This dissemination takes place through its website (which is updated regularly) and by social media.

c) Open science: Genes & Health is committed to open science. It openly shares its methods, data dictionaries, codelists and manuscripts on open access platforms and via its website.

d) Public engagement: The study team are guided by NIHR INVOLVE in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g. Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families. Genes & Health undertake innovative engagement activities with the award-winning Centre of the Cell, to deliver educational workshops and interactive web- and app- content based on the study objectives, and are supported by the QMUL Life Sciences Initiative, QMUL Centre for Public Engagement and a recent Wellcome Trust Public Engagement Award (PI D van Heel, ref 102627/B/13/A). The Helix Champions are also critical to ongoing and effective communication with volunteers to keep them engaged in the study and informed about its outcomes, and to engage with community leaders, religious organisations and schools. The research team have presented on local and national radio (Betar Bangla, BBC Asian Network and World Service) and BBC London TV. Genes & Health has an active Community Advisory Panel who support all ELGH activities through the entire research pathway, from prioritisation of topics for study, to dissemination of research findings.

The Genes & Health open access policy is described above and ensures that, where possible, there is no restrictive ownership of data or outputs in Genes & Health. It is possible that work on combined genetic and clinical phenotypes could derive significant commercial interest, e.g. to pharmaceutical companies for the development of new drugs, or genomics companies developing new disease risk algorithms. Genes & Health has a rigorous system for reviewing applications to work with its bioresource, involving its Executive team and Community Advisory Group. All collaborations and partnerships involving commercial organisations are pre-competitive and therefore not commercially exploitable.

Expected targets for outputs are, (a) short-term, e.g. publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g. building and expanding the bioresource within 5 years, and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years).

Some examples of current Genes & Health supported research is summarised below to illustrate the scope of the data request applied for:

• Type 2 diabetes – this condition disproportionately affects British south Asians and studies are currently taking place/planned with Genes & Health to investigate the influence of common (polygenic risk scores) and rare genetic variants (changes) in type 2 diabetes, misclassification of diabetes types in British south Asians, and gestational diabetes. These studies require detailed longitudinal clinical data from HES, National Diabetes Audit, MHSDS, MSDS, cancer registration, and civil registration of deaths in order to describe the onset of diabetes (including progression from at-risk states such as gestational diabetes), acute diabetes emergencies, glucose control and uptake of diabetes care, its comorbidities (including mental health disorders and cancer), complications and outcomes. Association-based analysis will be used to quantify the relationship between genetic variation and clinical outcomes, using survival analysis to understand trends over time. Adjustment for confounding variables, such as socioeconomic status, will be included in analyses.

• Familial hypercholesterolaemia (high cholesterol) – this rare genetic condition is a cause of excess death from cardiovascular disease. The condition is under-diagnosed in British south Asians, and work within Genes & Health is identifying volunteers with the condition through their genetic sequence data, and correlating this to clinical data to improve diagnosis and treatment. Longitudinal data from HES and ECDS is required to capture diagnoses and relevant hospital admissions with cardiac emergencies, and long-term outcomes including death from civil registration data. The analysis of health data will be used to design translational Stage 2 recall studies based on genotype and identify volunteers who need specific clinical intervention (e.g. initiation of specific drugs).

• Multimorbidity – a large programme of MRC-funded work is currently underway investigating the clustering of multimorbidity in British Bangladeshis and British Pakistanis, and their trajectories across the lifecourse. This research is taking a novel, data-driven approach to identify clusters of multimorbidity across multiple single conditions. The use of multi-source, linked medical record data will increase data quality (e.g. by validating diagnoses across datasets) and the ability to generate novel and meaningful multimorbidity clusters (e.g. encompassing both physical and mental health disorders by using HES and MHSDS/MHMDS/MHLDDS data). Historic data will allow analysis of a patients’ risk from multimorbidity across their lifecourse, and data from HES and MSDS will be used to investigate the impact of specific lifecourse events such as pregnancy, on the development of multimorbidity.

• COVID-19 – Genes & Health is contributing to international efforts to identify risk of disease, and its severity, in the host genome. The availability of HES data, including diagnoses from, and episodes in, emergency care (and ECDS), admitted patient care and critical care, and civil registration of deaths will support this work across all study volunteers. These data will be used in genome-wide association studies, with likely subgroup analysis according to disease severity/hospitalisation, and adjustment for confounding variables such as age and socioeconomic status (Index of Multiple Deprivation)

• Mental health and dementia – work is underway to better understand the genetic influences on mental health conditions, e.g. depression and anxiety, and dementia. This work will include longitudinal analysis of risk factors for these diseases their diagnosis and severity and associated mortality and therefore requires linked multisource data, including MHSDS, HES, and civil registration of deaths. Analyses will comprise genome-wide association studies, calculation of polygenic risk scores (including assessment of their performance in predicting disease onset).

• Discovery analyses of rare genetic variants (changes) – one of the unique features of the Genes & Health study is the ability to investigate the impact of rare genetic variation on health and disease due to its large scale and focus on a population with high rates of parental relatedness. The impact of rare genetic variants on an individual’s health requires careful study, particularly where the genetic variation is novel and its impact is unknown. A discovery-based approach is required to study such genetic variants, as they may have broad health consequences and novel disease associations. Conversely, rare genetic variants may offer protection from disease and the absence of diagnoses and hospital episodes would be highly informative. A proof-of-principal study of a rare genetic variant in the HAO1 gene has shown the importance of such studies using health data: an individual carrying a variant in this important metabolic regulator gene, had no ill effects on their health (determined by medical record data) and this knowledge provided critical information to support the development of a drug that targets this gene in a rare metabolic illness. All diagnoses and details of episodes of care are required to inform this novel genetic discovery-based research and direct subsequent translational research and drug development.

• The impact of inherited genetic variants on health – another unique feature of the Genes & Health study is the ability to investigate the impact of rare genetic variants arising from parental relatedness (called autozygous variants), on health. It is known that autozygosity can increase the risk of developmental disorders (the relevance of the Mental Health Minimum Dataset and its inclusion of diagnoses related to learning disability) are critical here. Additionally, there is a recent understanding that autozygosity may impact on a range of long-term conditions and reproductive health/fertility. It is particularly important to study these associations further in British south Asians due to the higher rates of parental relatedness. Discovery-based analyses are planned across a large range of traits and phenotypes to characterise these associations further, and data from HES, NDA, maternity and mental health datasets will be highly informative.

• Pregnancy-based studies – QMUL has active research investigating rare and severe pregnancy-based conditions with a known genetic basis (e.g. intrahepatic cholestasis of pregnancy) which will require detailed health record information to identify cases and outcome. Studies are underway within Genes & Health investigating common conditions in pregnancy (such as gestational diabetes or pre-eclampsia) their genetic basis (e.g. polygenic risk) and how this affects future risk of disease (such a progression to type 2 diabetes or cardiovascular disease). Additionally, QMUL have planned research investigating causal associations between maternal genetics offspring traits such as birth weight. All pregnancy based studies will require a core set of data from antenatal care, through the peripartum period, to immediate postnatal care to determine severity and outcome to the affected mother and her child. A limited set of data from offspring (e.g. birth weight, Apgar score, neonatal intensive care admission) has been requested to determine immediate pregnancy outcomes.

The above summary is not exhaustive and is to give an overview of the types of research currently funded and being undertaken in Genes & Health in order to justify the broad scope of data included in this application. It also reflects the need to be responsive to the ongoing development and use of Genes & Health as a bioresource that currently receives 2-4 new applications per month for new collaborative research studies. The data request has been designed to be comprehensive and support current and future research with the Genes & Health bioresource, but with the minimum of data fields required to meet its objectives.

Benefits reported

Genes and Health received NHS Digital data in Dec 2021. In the few months that they had access to the data they have been able to incorporate these data with health data from other sources, such as local NHS Trusts and primary care providers as a source of quality control and additional geographic coverage. These integrated datasets are difficult and time-consuming to produce, but mean that Genes & Health has very high-quality health data that will support research projects. This linkage is within the scope of the linkages explained within this DSA

An example of how these improved health datasets are already being used is QMUL's contribution to the international INTERVENE project (https://www.interveneproject.eu/) which is using AI and machine learning to understand the ways in which genetics contributes to disease risk. Genes & Health’s contribution to INTERVENE is really important because they provide data from people of Pakistani and Bangladeshi backgrounds, which would otherwise be underrepresented in the project. INTERVENE uses Genes & Health data, including health data that contains NHS Digital, to develop scores that can help doctors work out how likely it is that a patient will develop a disease. The new generation of genetic risk scores will help to diagnose and treat disease earlier, and so improve health. Genetic risk scores are in trials with the NHS, and new generations of scores built on high-quality and diverse datasets like INTERVENE will perform better and improve health further.

Analyses of data from the Genes & Health cohort have resulted in multiple high-impact publications. An up-to-date list is maintained on the study website (https://www.genesandhealth.org/about-study/scientific-publications) and a short summary of recent publications are included here:

Whole genome sequencing reveals host factors underlying critical Covid-19 (Genes & Health contributed as one of many international cohorts to the multiple phenotype analyses of the COVID-19 Host Genetics Initiative release 6):

Nature 2022. DOI https://doi.org/10.1038/s41586-022-04576-6

Transferability of genetic loci and polygenic scores for cardiometabolic traits in British Pakistanis and Bangladeshis: medRxiv 2021 https://www.medrxiv.org/content/10.1101/2021.06.22.21259323v1

Harnessing the power of polygenic risk scores to predict type 2 diabetes and its subtypes in a high-risk population of British Pakistanis and Bangladeshis in a routine healthcare setting: medRxiv 2021 https://www.medrxiv.org/content/10.1101/2021.07.12.21259837v1

Global Biobank Meta-analysis Initiative: powering genetic discovery across human diseases: medRxiv 2021 https://doi.org/10.1101/2021.11.19.21266436

The power of genetic diversity in genome-wide association studies of lipids: Nature 2021. DOI https://doi.org/10.1038/s41586-021-04064-3

MC3R links nutritional state to childhood growth and the timing of puberty: Nature 2021. DOI https://doi.org/10.1038/s41586-021-04088-9

Mapping the human genetic architecture of COVID-19: Nature 2021 https://doi.org/10.1038/s41586-021-03767-x

DARS-NIC-338864-B3Z3J-v0.12 9 July 2021 to 8 July 2022
Title
Genes and Health
Commercial
Yes
Sublicensing
Yes
Datasets
12
Files released
152

Datasets: Bridge file: Hospital Episode Statistics to Mental Health Minimum Data Set; Civil Registrations of Death; Emergency Care Data Set (ECDS); Hospital Episode Statistics Accident and Emergency (HES A and E); Hospital Episode Statistics Admitted Patient Care (HES APC); Hospital Episode Statistics Critical Care (HES Critical Care); Hospital Episode Statistics Outpatients (HES OP); Maternity Services Data Set (MSDS) v1.5; Mental Health and Learning Disabilities Data Set (MHLDDS); Mental Health Minimum Data Set (MHMDS); Mental Health Services Data Set (MHSDS); National Diabetes Audit

Objective for processing

This data request comes from the Genes & Health Study, which is based at Queen Mary University of London (QMUL).

Genes & Health is a major UK-based research programme of health and disease in British Bangladeshis and British Pakistanis.

QMUL wish to combine data from NHS Digital with data from the Genes and Health BioResource. This linked resource will be made available via sub-licensing arrangements.

The strategic objective of Genes & Health is to develop and maintain a BioResource of genetic and health record data available to the research community (academic and industrial partners) to improve the health of people of British Bangladeshi and British Pakistani origin through high quality research.

Health record data is a key requisite of this objective, as it is used to characterise in detail health and disease in Genes & Health study volunteers across their lifecourse. The Genes & Health study team at QMUL have built a database to investigate genetic data and health data from Genes & Health participants, and hope to include data from NHS Digital. This database will be used by researchers at QMUL, as well as researchers in other universities, charities, and commercial companies, once they have been approved by an application review committee, and signed a sublicensing agreement.

Genes & Health Volunteers have given written informed consent for their health data to be accessed stored and processed in line with the study objectives. The study is organised as follows:

(i) Stage 1 research: genetic information from saliva samples provided at recruitment by Genes & Health volunteers is combined with health record data to characterise genetic variation and correlate this to physical characteristics, and states of health and disease.

(ii) Stage 2 research; information from genetic studies and/or health record information are used to guide focused translational research studies, e.g. investigating how a rare gene variant might impact on health.

The article 6 justification for processing the requested data is stated in 1. (e) “processing is necessary for the performance of a task carried out in the public interest or in the exercise of official authority vested in the controller”

This data request is in the public interest under Article 9(2)(j). Specifically, it is in the public interest for scientific and statistical purposes in accordance with Article 89(1). Furthermore, the request is proportionate to the aim pursued, respects the essence of the right to data protection, and provides for suitable and specific measures to safeguard the fundamental rights and the interests of the data subject.

Dissemination of results arising from this data request, including those published from sublicenced data, will be fully anonymised, and will inform study volunteers and the wider public of the research findings at summary level. These findings will comprise new knowledge of health and disease and will be disseminated using lay language. Where findings might be sensitive, e.g. reporting an increased risk of a severe disease in the study population, or its genetic cause, the Community Advisory Group will be consulted to ensure effective and appropriate means of communication.

Volunteers give consent for linkage to their medical records, including from their GP and hospital records as well as from nationally-held datasets such as those included in this agreement, for research purposes. Genes & Health is a multi-purpose bioresource, and combines health data with genetic data (obtained from saliva samples collected at recruitment) to undertake detailed characterisation of states of health and disease and generate new knowledge relating to these. Genes & Health recruitment started in 2015 in east London, where limited health data has been obtainable through linkage to local health systems. Genes & Health has recruited over 47,000 volunteers and is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on east London, Luton, Bradford and Manchester). The cohort is expected to reach 100,000 volunteers. Linkage to national datasets is required to (a) expand the geographical coverage of data linkage where local datasets are not available or where healthcare is accessed beyond its limits, and (b) to enrich the datasets available using high quality multisource national data that are not available through local health systems, e.g. civil registration of deaths.

Genes & Health receives funding from leading funders including the Wellcome Trust, Medical Research Council, National Institute for Health Research Clinical Research Networks and others. People of British Bangladeshi and British Pakistani origin are significantly underrepresented in other research studies, including UK Biobank. Genes & Health is an open access resource, making anonymised genetic and medical record data available in a data safe haven to external researchers and industrial partners to deliver its aims and maximise its potential to deliver benefit to population health. The infrastructure of Genes & Health has, at its core, rigorous protocols to ensure the appropriate use of data, data safety and volunteer acceptability.

The study complements, but is distinct from, existing bioresources (e.g. the well-known Cambridge BioResource and UK Biobank) by its focus on two distinct UK ethnic minority groups with a different spectrum of genetic variation, high levels of health deprivation, and substantial parental relatedness. As such, Genes & Health will deliver new knowledge relevant to an understudied population at high risk of ill health. Genes & Health also complements existing large-scale open access health data resources, e.g. the Clinical Practice Research Datalink, by its linkage of health data to genetic data.

The research taking place within Genes & Health reflects community prioritisation, health need, research expertise and emerging knowledge. Furthermore, Genes & Health is an adaptable and responsive resource that can accommodate changing priority and need, an example being its recent rapid contribution to the International COVID-19 Host Genetic Initiative.

Use of data from key national health datasets will support the multipurpose, bioresource function of Genes & Health by expanding the depth and breadth of data available and by increasing the availability of health data on volunteers receiving care outside of local health systems where linkage is already in place. The scope of the data request outlined in this application is broad, and reflects the design and need of Genes & Health as a bioresource supporting a large portfolio of research projects in Stage 1 and Stage 2 studies:

(a) Stage 1 research: genetic information from saliva samples provided at recruitment by Genes & Health volunteers is combined with health record data to characterise genetic variation and correlate this to biomedical traits and phenotypes, including states of health and disease. These studies generally require population-based analysis of association (e.g. in case-control studies of the genetic associations of prevalent disease, or the prediction of disease onset using risk scores) or longitudinal analysis of disease progression (e.g. using survival analysis incorporating risk factors, incident disease, and progression to complications and death). All analyses require careful adjustment for confounding variables (e.g. socioeconomic status, comorbidity/multimorbidity, smoking) and use of time-dependent covariates (e.g. pregnancy, progression from at-risk states).

(b) Stage 2 research; information from genetic studies and/or health record information are used to guide focused translational research studies, e.g. investigating how a rare gene variant might impact on health. These studies are typically smaller in scale compared to Stage 1 studies, and might involve a small number of individuals (e.g. 5-100) carrying a rare genetic variant whose function is unknown. In such examples, health record data will be used to identify novel associations between genotype and health/disease status.

Genes & Health invites the following people to participate:

• People living within reach of its research sites (currently east London, Luton, Bradford and Manchester)

• Self described ethnicity: Bangladeshi, British-Bangladeshi, Pakistani, British-Pakistani

• Age 16 or over (no upper limit)

• Willingness to give a saliva DNA sample for analysis

• Willingness to grant investigators access to medical records (GP, Hospital, and national NHS) for the full duration (stage 1 and stage 2) of the study.

• Willingness to receive invitations for recall (stage 2 research activity) for the duration of the study (actual participation in all stage 2 research activity would be after a second informed consent process).

QMUL seek linkage of NHS Digital datasets at record level to all Genes & Health volunteers. This agreement is requesting pseudonymised, record-level data.

Genes & Health collects saliva samples from all participants, which QMUL use to generate genetic data like SNP genotyping and DNA sequencing. From a select group of volunteers QMUL also collect samples like blood and urine, and make clinical measurements like height and weight. To understand how a person's genetics interact with their health, to cause bad health or to protect against disease, researchers need to be able to analyse genetic data from Genes & Health alongside health data from NHS Digital – one set of data alone won’t help understand the connection between genetics and health.

The partners who will use the Genes & Health resource include academic research groups at UK universities (the vast majority of users) and at universities in the EEA. Partners will also include technology and pharmaceutical companies based in the UK, and EEA, who want to access Genes & Health data for research and development purposes. All partners, academic and commercial, are required to submit an application for a defined project for review, have a summary of their project published online, and making appropriate results of their work available to the research community. To date 45 applications have been received, of which only three were from commercial companies.

Carefully curated, and minimised, data from the following resources are requested in this agreement:

¥ Hospital Episode Statistics (HES), including HES A&E, OP, critical care and APC.

¥ HES APC:

This dataset will support research into ill health and disease identified, diagnosed and treated during hospital admissions, as well as episodes of pregnancy care. The data obtained from HES APC will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. A number of data fields have been requested in addition to STUDY_ID to ensure adequate linkage between the Genes & Health ID and pseudonymised HES identifiers across HES datasets,

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and pregnancy. This data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

¥ HES OP:

This dataset will support research into ill health and disease, e.g. depression treated in a specialist psychiatry outpatient clinic, dementia diagnosed in a memory clinic, and the investigations/care/treatment received related to it. The data obtained from HES OP will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved), discharge and outpatient follow up.

This data will inform a more extensive clinical picture, including capturing the diagnosis and monitoring of long-term conditions in outpatient care.

¥ HES AE:

This dataset will support research into acute presentations of ill health, e.g. heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. The data obtained from HES A&E will complement the other HES datasets that are being requested also, giving QMUL health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge. A number of data fields have been requested in addition to STUDY_ID to ensure adequate linkage between the study ID and pseudonymised HES identifiers across HES datasets,

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

¥ HES CC

This dataset will support research into severe ill health/disease, specifically related to critical care in this context. The data obtained from HES Critical Care will complement the other HES datasets that are being requested also, providing health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received. For example, Genes & Health researchers will be able to determine severity of illnesses under close study, e.g. in a study of cardiovascular disease, data showing a patient required a long critical care admission during an episode of care when a myocardial infarction was recorded will imply severity.

¥ Emergency Care Data Set (ECDS)/

This dataset will support research into acute presentations of ill health, e.g. heart attack, an acute diabetes emergency, and the investigations/care/treatment received related to it. The data obtained from ECDS will complement HES A&E and other HES datasets that are being requested also, to give health data from emergency presentation, through admission, procedures, critical care (where involved) and discharge.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease. This data will inform a more extensive clinical picture beyond simply diagnosis, giving valuable information about treatment and the care received.

¥ The Mental Health Services Data Set (MHSDS), and its predecessors, the Mental Health Minimum Data Set (MHMDS) and Mental Health and Learning Disabilities Data Set (MHLDDS)

¥ MHSDS: This dataset will support research into mental ill health in volunteers, including the study of rare and common disease, and their genetic basis. The study will be able to research diagnoses, as well as care received, across multiple studies encompassing mental health and multimorbidity.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health and disease, including those related to mental health. This data will inform a more extensive clinical picture that includes diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

¥ MHMDS/MHLDDS

This dataset will support research into mental ill health in the Genes and Health volunteers, including the study of rare and common disease, and their genetic basis.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of mental ill health and disease, the care and treatment received.

¥ Maternity Service Data Set (MSDS)

This dataset will support research into physical traits, ill health and disease identified, diagnosed and treated during pregnancy. This dataset will complement that obtained from HES APC.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise physical traits and states of ill health, disease. Genes & Health has a number of studies specific to pregnancy, including studies of rare and common pregnancy disorders. Pregnancy-based characterisation of physical traits (e.g. body mass index at booking, or social factors) will also add valuable detail to the characterisation of our female study volunteers more broadly, for cross-sectional and longitudinal studies.

¥ Civil Registrations (Deaths)

This dataset will support research into ill health and disease by identifying death and cause of death.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and their long-term outcomes, in this case death and cause of death.

¥ National Diabetes Audit

This dataset will support extensive and high quality research into diabetes and its related co-morbidities and multimorbidity.

The data requested is relevant to the strategic objective of Genes & Health, specifically to characterise states of ill health, disease and pregnancy. This data will inform a more extensive clinical picture related to people with diabetes (who make up 18% of the study volunteers), including clinical measurements, diagnosis, treatment and care received, as well as important patient characteristics such as the calculation of socioeconomic status indices.

¥ Bridging files HES:civil registrations, HES:MHSDS

To allow the datasets to be abridged.

Use of data from key national health datasets will support the multipurpose, bioresource function of Genes & Health by expanding the depth and breadth of data available and by increasing the availability of health data on volunteers receiving care outside of local health systems where linkage is already in place. The content of data analysis has also been designed carefully to mirror the NHS Digital data available via UK Biobank and Clinical Practice Research Datalink so that replication and validation studies can be performed across all of these valuable resources. This mirroring is to allow similar analyses to be replicated between datasets, no direct linkage of BioBank, CPRD and Genes & Health data is anticipated.

The data requested is required for the following types of analyses:

(a) Association-based genetic analysis of prevalent disease(s), its severity and characteristics

(b) Prediction of disease risk and severity using polygenic risk scores

(c) Longitudinal analysis of states of health and disease(s) through the lifecourse (including pregnancy), and their association with genetic factors

(d) Discovery-based analysis of the impact of rare genetic variation on health and disease

(e) Assessment of, and adjustment for, confounding influences of the above using covariates.

(f) Assessment of causal relationships in health and disease using genetic data in Mendelian randomisation

(g) Validation studies using multi-source health data to quality control self-report data (e.g. the study questionnaire and local health data) and data from other bioresources (e.g. UK Biobank, Clinical Practice Research Datalink)

(h) Replication studies of known genetic factors on health and disease in an understudied population of British Pakistanis and British Bangladeshis.

The study incorporates longitudinal analysis of states of health and disease and includes the study of historic diagnoses. For example, an admission with myocardial infarction in 1999, is important as it will show the earliest time a patient was diagnosed. This will be vital for understanding early onset disease.

QMUL also have ongoing funding and recruitment to 2023 at least and expect Genes & Health to serve as a long-term bioresource in the future, therefore QMUL will be requesting an annual update to this data linkage to update previously linked records and add linkage to newly-recruited volunteers. The study is recruiting nationally across multiple sites (currently London, Luton, Bradford and Manchester), with the possibility of opening additional sites in the future. It is important for Genes & Health to be able to link to national data in order to capture health data on volunteers who move house.

There are no alternative data sources that would achieve the full aims and objectives of the Genes & Health study. Local health system data has already been used to its maximum potential, but this is limited to those systems able to supply data, and the data returned are limited to patients treated within those local systems. Currently this includes Barts Health, and a small number of local CCGs who have signed data transfer agreements, but not primary or secondary care providers further afield. Genes & Health is expanding nationally with recruitment taking place across multiple geographical regions in the UK (currently focusing on east London, Luton and Bradford) and linkage to national datasets is required to (a) expand the geographical coverage of data linkage where local datasets are not available or where healthcare is accessed beyond its limits, and (b) to enrich the datasets available using high quality multisource national data that are not available through local health systems, e.g. civil registration of deaths.

Although data has been requested from multiple NHS Digital datasets, stringent variable selection has been undertaken for the purposes of data minimisation. Genes & Health have taken the following steps to ensure that the variables requested are strictly necessary and achieve data minimisation:

1) Detailed curation of required variables to meet the Genes & Health study objectives, achieved through discussion and consensus with the core Genes & Health research and Executive teams. The focus of this data request is on diagnoses, clinical observations and related dates, rather than organisational information that does not have a clear purpose in the research objectives.

2) Non-replication of data already held – Genes & Health already hold data, such as full postcode, particularly where it is identifiable or sensitive. A limited set of non-identifiable data (year and month of birth, gender, postcode district and ethnicity) has been requested as a quality control check across datasets.

3) Genes & Health have followed previous large-scale bioresource studies who have used NHS Digital datasets to guide the choice of variables, including UK Biobank and the INTERVAL and COMPARE studies of blood donors, as well as the Clinical Practice Research Datalink. This data request mirrors the set of variables available in these studies that are relevant to the Genes & Health study objectives, allowing for cross-study validation and replication. This will significantly expand the utility of the data requested.

4) Genes & Health have taken extensive and detailed guidance from the DARS team (case officer and data production team member) to guide the selection of variables, in particular where multiple datasets exist for the same clinical areas (e.g. MHSDS, MHLDDS and MHMDS).

5) The request is also minimised to the size of the cohort, currently a little over 47,000 participants. Genes & Health is still actively recruiting, and the number of participants is expected to reach 100,000 by approximately 2023.

QMUL is the sole data controller for Genes and Health data, including these data requested from NHS Digital. Swansea University - who provide the UK Secure eResearch Platform (UK SeRP), will store the data and are a data processor.

Researchers from external institutions, including universities and commercial companies, will, with the agreement, direction and oversight of QMUL, analyse Genes & Health data on a platform provided by UK SeRP (part of Swansea University). This will be facilitated by sublicensing agreements with all organisations. An access register will detail these projects and organisations, and a summary of all projects will be published on the Genes & Health website.

A Data Access Agreement (DAA) will be signed by external institutions before access to Genes & Health data is granted. This agreement sets out the scope of data processing activities that may be undertaken and the legal and technical data protection requirements. To try and reinforce good practice, individual researchers are also asked to sign, having read a letter summarising their data protection responsibilities in the project.

Several organisations have provided data or services but are not data processors for the NHS Digital data. These are King’s College London and UK Biocentre, who have processed DNA samples from participants, and Wellcome Sanger Institute who have analysed DNA samples. Discovery Data Service have facilitated primary healthcare record data access from members of the East London Health and Care Partnership.

Genes & Health funders are not involved in the day to day running of the project, and do not have access to the data. Commissioners or representatives were involved in reviewing the proposal for data sharing and are involved in overseeing the wider data sharing partnerships, but they are not directly involved in the Genes & Health project and do not have access to the data.

Sub-licenses will be granted upon successful application, where applicants satisfy section viii of the terms of reference – which includes data security, public benefit, impact on participants. Applications covering a sensitive topic (for example sexual health, consanguinity, mental health) will be reviewed by a community panel.

Conditions will be compliance with the terms of the agreement, which include strict data security requirements imposed by our research environment and to only undertake work within in the remit of their application. QMUL are able to view and audit the work of collaborators, and as the environment is export controlled QMUL will actively review every request for data-out to make sure they are complying with the terms.

Data disseminated by NHS Digital data is potentially identifiable in the hands of the Data Controller as QMUL have the ability to link the Study ID's to the identifiers. Data will, however, always be pseudonymised in the hands of the sub-licensees.

Expected output

Many health research and genetic studies have focused more on other populations (for example, white European origin people), even though South Asian communities see high rates of diabetes, heart disease, rare genetic diseases,and many other conditions. Genes & Health was set up to make sure that advances in genetics and healthcare are available to this underserved community, and to help provide insights into some of the health inequalities that exist within the UK. The ability to recall participants with unique genetic makeup or combination of genetics and health status, also provides a powerful opportunity to better understand how genetics and health are linked, which may be of benefit to everyone.

The results of data processing will be new knowledge related to health and disease in people of British Bangladeshi and British Pakistani origin. The Genes & Health findings will be shared with the wider scientific community via presentations at national and international conferences and professional meetings across a range of audiences including genomics, clinical and health data communities and scientific publications. Genes & Health has already published several peer-reviewed publications, including a cohort profile (Finer et al, International Journal of Epidemiology, PMID 31504546), and empirical research demonstrating the potential of the bioresource to build new knowledge on health and disease in understudied ethnic groups and improve patient care (McGregor et al, eLife, PMID 26940866; Narasimhan et al, Science, PMID 32207686). Outputs arising from use of the Genes & Health bioresource is likely to generate a significant number of high impact publications over the next 5 years, and preprint servers and open access journals will be used in preference. Study research findings are also summarized on the Genes & Health website (www.genesandhealth.org) and via its twitter account (@eastlondongenes and @bradfordgenes, and Manchester when they are ready to recruit).

Genes & Health has an established protocol for sharing data outputs, and will continue to use this for outputs arising from NHS Digital datasets. The protocol for sharing data outputs is as follows:

Level 1: Fully open data. Genes & Health distribute summary level analysed data on the study website and via twitter. For example summary phenotype counts, e.g. numbers of volunteers with diabetes, or summary statistics of clinical observations such as blood lipid measurements. Small number suppression will be used when the numbers of volunteers is very low and there was a risk of being able to identify individuals or families from these data.

Level 2: Fully anonymised genetic data is available under a Data Access Agreement with the European Genome- phenome Archive (EGA, https://www.ebi.ac.uk/ega/home). This data will not be linked to any health data, including that obtained from NHS Digital.

Level 3: Data access and analysis within Data Safe Haven. Data security is central to ELGH (and written into the study Ethics and Governance), a data breach would irretrievably damage the study and the community trust that has been built. Applicants wishing to analyse phenotype data, e.g. NHS Digital health data, do this within an ISO27001 and NHS Information Governance compliant Data Safe Haven environment (currently UK-SERP, which provides a Virtual Desktop Infrastructure on the end user device, with export controls). Data outputs and publications will be summarised and presented in aggregate, and it will not be possible to identify individuals.

All collaborative research involving NHS Digital datasets will require a Data Sharing Agreement to ensure a rigorous approach to data safety and onward use.

The applicants will take the following approaches:

a) Academic dissemination: this has been described above and will include communication of results through peer-reviewed publication and conference presentations, as well as through other clinical and academic networks, relevant professional interest groups and social media.

b) Stakeholder dissemination: Genes & Health disseminates regular updates regarding study activities (e.g. recruitment, outputs) to a range of stakeholders, including funders, clinicians, policymakers, community representatives and volunteers. This dissemination takes place through its website (which is updated regularly) and by social media.

c) Open science: Genes & Health is committed to open science. It openly shares its methods, data dictionaries, codelists and manuscripts on open access platforms and via its website.

d) Public engagement: The study team are guided by NIHR INVOLVE in all dissemination activities with the public. Genes & Health is a community-embedded study that keeps engagement at its core, supported by its Community Advisory Group, 'Helix Champions' (community researchers) and third sector organisations, e.g. Social Action for Health. The study and QMUL support regular community engagement and health education activities sharing new knowledge amongst volunteers and their families. Genes & Health undertake innovative engagement activities with the award-winning Centre of the Cell, to deliver educational workshops and interactive web- and app- content based on the study objectives, and are supported by the QMUL Life Sciences Initiative, QMUL Centre for Public Engagement and a recent Wellcome Trust Public Engagement Award (PI D van Heel, ref 102627/B/13/A). The Helix Champions are also critical to ongoing and effective communication with volunteers to keep them engaged in the study and informed about its outcomes, and to engage with community leaders, religious organisations and schools. The research team have presented on local and national radio (Betar Bangla, BBC Asian Network and World Service) and BBC London TV. Genes & Health has an active Community Advisory Panel who support all ELGH activities through the entire research pathway, from prioritisation of topics for study, to dissemination of research findings.

The Genes & Health open access policy is described above and ensures that, where possible, there is no restrictive ownership of data or outputs in Genes & Health. It is possible that work on combined genetic and clinical phenotypes could derive significant commercial interest, e.g. to pharmaceutical companies for the development of new drugs, or genomics companies developing new disease risk algorithms. Genes & Health has a rigorous system for reviewing applications to work with its bioresource, involving its Executive team and Community Advisory Group. All collaborations and partnerships involving commercial organisations are pre-competitive and therefore not commercially exploitable.

Expected targets for outputs are, (a) short-term, e.g. publication of results within 1-2 years of receiving data for analysis, and (b) medium- to long-term, e.g. building and expanding the bioresource within 5 years, and maintaining an open access research resource to be used in global consortium-based work and replication studies (5-10 years).

Some examples of current Genes & Health supported research is summarised below to illustrate the scope of the data request applied for:

• Type 2 diabetes – this condition disproportionately affects British south Asians and studies are currently taking place/planned with Genes & Health to investigate the influence of common (polygenic risk scores) and rare genetic variants (changes) in type 2 diabetes, misclassification of diabetes types in British south Asians, and gestational diabetes. These studies require detailed longitudinal clinical data from HES, National Diabetes Audit, MHSDS, MSDS, cancer registration, and civil registration of deaths in order to describe the onset of diabetes (including progression from at-risk states such as gestational diabetes), acute diabetes emergencies, glucose control and uptake of diabetes care, its comorbidities (including mental health disorders and cancer), complications and outcomes. Association-based analysis will be used to quantify the relationship between genetic variation and clinical outcomes, using survival analysis to understand trends over time. Adjustment for confounding variables, such as socioeconomic status, will be included in analyses.

• Familial hypercholesterolaemia (high cholesterol) – this rare genetic condition is a cause of excess death from cardiovascular disease. The condition is under-diagnosed in British south Asians, and work within Genes & Health is identifying volunteers with the condition through their genetic sequence data, and correlating this to clinical data to improve diagnosis and treatment. Longitudinal data from HES and ECDS is required to capture diagnoses and relevant hospital admissions with cardiac emergencies, and long-term outcomes including death from civil registration data. The analysis of health data will be used to design translational Stage 2 recall studies based on genotype and identify volunteers who need specific clinical intervention (e.g. initiation of specific drugs).

• Multimorbidity – a large programme of MRC-funded work is currently underway investigating the clustering of multimorbidity in British Bangladeshis and British Pakistanis, and their trajectories across the lifecourse. This research is taking a novel, data-driven approach to identify clusters of multimorbidity across multiple single conditions. The use of multi-source, linked medical record data will increase data quality (e.g. by validating diagnoses across datasets) and the ability to generate novel and meaningful multimorbidity clusters (e.g. encompassing both physical and mental health disorders by using HES and MHSDS/MHMDS/MHLDDS data). Historic data will allow analysis of a patients’ risk from multimorbidity across their lifecourse, and data from HES and MSDS will be used to investigate the impact of specific lifecourse events such as pregnancy, on the development of multimorbidity.

• COVID-19 – Genes & Health is contributing to international efforts to identify risk of disease, and its severity, in the host genome. The availability of HES data, including diagnoses from, and episodes in, emergency care (and ECDS), admitted patient care and critical care, and civil registration of deaths will support this work across all study volunteers. These data will be used in genome-wide association studies, with likely subgroup analysis according to disease severity/hospitalisation, and adjustment for confounding variables such as age and socioeconomic status (Index of Multiple Deprivation)

• Mental health and dementia – work is underway to better understand the genetic influences on mental health conditions, e.g. depression and anxiety, and dementia. This work will include longitudinal analysis of risk factors for these diseases their diagnosis and severity and associated mortality and therefore requires linked multisource data, including MHSDS, HES, and civil registration of deaths. Analyses will comprise genome-wide association studies, calculation of polygenic risk scores (including assessment of their performance in predicting disease onset).

• Discovery analyses of rare genetic variants (changes) – one of the unique features of the Genes & Health study is the ability to investigate the impact of rare genetic variation on health and disease due to its large scale and focus on a population with high rates of parental relatedness. The impact of rare genetic variants on an individual’s health requires careful study, particularly where the genetic variation is novel and its impact is unknown. A discovery-based approach is required to study such genetic variants, as they may have broad health consequences and novel disease associations. Conversely, rare genetic variants may offer protection from disease and the absence of diagnoses and hospital episodes would be highly informative. A proof-of-principal study of a rare genetic variant in the HAO1 gene has shown the importance of such studies using health data: an individual carrying a variant in this important metabolic regulator gene, had no ill effects on their health (determined by medical record data) and this knowledge provided critical information to support the development of a drug that targets this gene in a rare metabolic illness. All diagnoses and details of episodes of care are required to inform this novel genetic discovery-based research and direct subsequent translational research and drug development.

• The impact of inherited genetic variants on health – another unique feature of the Genes & Health study is the ability to investigate the impact of rare genetic variants arising from parental relatedness (called autozygous variants), on health. It is known that autozygosity can increase the risk of developmental disorders (the relevance of the Mental Health Minimum Dataset and its inclusion of diagnoses related to learning disability) are critical here. Additionally, there is a recent understanding that autozygosity may impact on a range of long-term conditions and reproductive health/fertility. It is particularly important to study these associations further in British south Asians due to the higher rates of parental relatedness. Discovery-based analyses are planned across a large range of traits and phenotypes to characterise these associations further, and data from HES, NDA, maternity and mental health datasets will be highly informative.

• Pregnancy-based studies – QMUL has active research investigating rare and severe pregnancy-based conditions with a known genetic basis (e.g. intrahepatic cholestasis of pregnancy) which will require detailed health record information to identify cases and outcome. Studies are underway within Genes & Health investigating common conditions in pregnancy (such as gestational diabetes or pre-eclampsia) their genetic basis (e.g. polygenic risk) and how this affects future risk of disease (such a progression to type 2 diabetes or cardiovascular disease). Additionally, QMUL have planned research investigating causal associations between maternal genetics offspring traits such as birth weight. All pregnancy based studies will require a core set of data from antenatal care, through the peripartum period, to immediate postnatal care to determine severity and outcome to the affected mother and her child. A limited set of data from offspring (e.g. birth weight, Apgar score, neonatal intensive care admission) has been requested to determine immediate pregnancy outcomes.

The above summary is not exhaustive and is to give an overview of the types of research currently funded and being undertaken in Genes & Health in order to justify the broad scope of data included in this application. It also reflects the need to be responsive to the ongoing development and use of Genes & Health as a bioresource that currently receives 2-4 new applications per month for new collaborative research studies. The data request has been designed to be comprehensive and support current and future research with the Genes & Health bioresource, but with the minimum of data fields required to meet its objectives.

Benefits reported

Yielded Benefits is not a requirement for new applications.

Register history

When this agreement appeared in, or was edited in, each monthly edition of the register. Built by comparing every edition this site holds.

"Amended in place" means NHS England changed the record without issuing a new version number. The register publishes no changelog for those edits; this site infers them by comparing editions. An edit is attributed to the edition it first appears in, not to the date it was made.

Cite this page

NHS England (2026) Data Uses Register, September 2026 edition, agreement DARS-NIC-338864-B3Z3J, “Genes and Health”. Read via NHS Data Access Explorer (unofficial), https://healthdatauses.uk/agreements/dars-nic-338864-b3z3j/ (accessed [date]).

This address stays the same, but the page is rebuilt with each monthly edition, so the citation names the edition it shows. Every edition's data is kept in the facts store.

Source: datausesregister_september2026.xlsx, September 2026 edition of the NHS England Data Uses Register. Search that workbook for DARS-NIC-338864-B3Z3J to see the original rows.