Written evidence from Peter Freeman, Lecturer in Healthcare Sciences (Clinical Bioinformatics and Genomics), The University of Manchester and support by Policy@Manchester (PMA0037)
Generative AI was used to assist in the initial drafting of this response. The author accepts responsibility for the contents of this submission.
Executive summary
- This submission responds specifically to Question 3: Health Data Research Infrastructure. Drawing on my technical leadership in genomic data standards, international editorial experience, and work training NHS bioinformaticians, it outlines the key systemic deficiencies in the UK’s health and genomic data infrastructure and provides concrete recommendations to ensure the NHS can fully realise the benefits of personalised medicine and AI.
- The key barrier to NHS genomics innovation is not sequencing or computational limitations, but inconsistent and inaccurate variant data standards, which prevent reliable linkage across labs, electronic health records, research systems and AI tools.
- Variant nomenclature errors are universal and systemic, undermining diagnostic accuracy, evidence discovery and AI training; without enforced HGVS‑aligned standards, national datasets will remain fragmented.
- The NHS lacks a central, standards‑aligned genomic database and validated data‑egress tools, resulting in divergent representations across Genomic Laboratory Hubs, and blocking national reanalysis, AI deployment and evidence pooling.
- Government action is needed to mandate standards, co‑commission specialist national tooling, and invest in the genomics bioinformatics workforce, ensuring interoperability, public trust, improved diagnostics and UK leadership in personalised medicine.
Background
- I have over a decade of experience in genomic data annotation, clinical bioinformatics, and the development of internationally adopted frameworks for the accurate reporting of genetic variation. My work focuses on ensuring the accuracy, reproducibility, and interoperability of genomic data – essential foundations for personalised medicine and for the safe and effective deployment of AI across the NHS.
- Since 2019, I have been a major contributor to the NHS Scientist Training Programme in Clinical Bioinformatics (Genomics), leading national modules in diagnostic sequencing, variant annotation, and software engineering for clinical genomics. Through training healthcare scientists across all seven NHS Genomic Laboratory Hubs (GLHs), I have gained substantial insight into the technical, organisational, and infrastructural challenges that affect data quality, hinder adoption of best practice, and slow the integration of AI and personalised medicine into routine NHS care.
- I am the creator and Principal Investigator of VariantValidator, used across all seven GLHs and internationally to standardise variant representation, reduce diagnostic errors, and support compliance with global genomic reporting standards. Its impact has been documented in Nature Genetics (Freeman et al., 2024), where we demonstrated that improved nomenclature accuracy directly increases diagnostic rates.
- I serve as Lead Technical Editor at Genetics in Medicine and GIMO, where I lead a specialist editorial team dedicated to ensuring accurate data representation in peer‑reviewed genomics manuscripts. This work includes co‑authoring a study (Lansdon et al., 2026) which systematically demonstrated the near‑universal presence of nomenclature errors in submitted manuscripts and underscored the structural challenges caused by inconsistent genomic reporting – challenges that are amplified in AI training datasets.
- I am a senior member of 2the Reporting of Sequence Variants Working Group, part of the Human Genome Organisation (HUGO), where I co‑authored guidance for verifying variant nomenclature in scientific manuscripts (Higgins et al., 2021). I am also a core contributor to the multi‑organisation project “Standards for journals, authors and clinical laboratories for the reporting and sharing of interpreted genomic variation”, involving:
- The American College of Medical Genetics and Genomics (ACMG)
- The American Society of Human Genetics (ASHG)
- The Association for Clinical Genomic Science (ACGS)
- The Association for Molecular Pathology (AMP)
- The Canadian College of Medical Geneticists (CCMG)
- The College of American Pathologists (CAP)
- The European Molecular Genetics Quality Network (EMQN)
- The Human Genome Organisation (HUGO)
-
-
-
-
-
-
-
-
- This collaboration focuses on defining robust, interoperable, computationally reliable genomic data standards that support global consistency and enable AI‑readiness.
Health Data Research Infrastructure
Question 3: What further research infrastructure is needed to support personalised medicine and AI, and where are the gaps?
- Personalised genomic medicine depends on robust linkage between genomic variation and clinical evidence (Richards et al., 2015); phenotypic datasets, which link physical traits to genetic data (Köhler et al., 2019); and health‑system data, all encoded in a standardised and fully computable framework. However, a current core obstacle lies not in sequencing capacity or computational power, but in the fragmented and inconsistent ways that genomic variation is represented, stored, and communicated across the NHS and globally.
- In clinical practice, multiple systems are required and utilised to represent and exchange genomic data, each with its own strengths and limitations. The Variant Call Format (VCF) (Danecek et al., 2011) remains widely used in diagnostic workflows, but frequently loses essential metadata required for linkage, interpretation, and reproducibility. Newer computable frameworks, such as the GA4GH Variation Representation Specification (VRS) (Goar et al., 2022), have been developed to provide globally interoperable, unambiguous representations of genomic variation, and the SPDI model (Holmes et al., 2020) has emerged as another attempt to deliver stable, computation‑ready descriptions of variants. Despite their computational strengths, these models remain insufficient for human interpretation, particularly because they cannot effectively represent the consequences of genetic variation at a sufficient depth as required in clinical practice. Consequently, the globally adopted and clinically mandated nomenclature for reporting sequence variants remains the Human Genome Variation Society (HGVS) standard (den Dunnen et al., 2016).
- HGVS offers a human‑readable and clinically interpretable description of variants and is the nomenclature understood by clinicians and clinical scientists. It is explicitly mandated by the ACMG and the AMP in guidelines for variant interpretation (Richards et al., 2015), which underpin UK practice through modification and adoption by the ACGS and the EMQN.
- Because the different naming systems were developed for different purposes, and serve fundamentally different audiences, inconsistencies arise at exactly the points where data must be linked. Variant identifiers cannot be reliably matched between systems without stable, standards‑compliant transformation pipelines. This represents a foundational barrier to personalised medicine and AI, because effective linkage requires variant descriptions that are consistent, computable, and traceable across sequencing pipelines, clinical reports, electronic health records (EHRs), phenotypic repositories, and national‑scale databases.
- Despite the long‑standing mandate to use HGVS nomenclature – and the fact that HGVS remains the globally adopted clinical standard and the primary entry point for retrieving diagnostic evidence from the literature and curated databases – its application remains inconsistent across the NHS and globally. HGVS descriptions are often omitted from reports or substituted with imprecise or outdated representations, and where it is used, it is frequently incorrect. As part of the HUGO Reporting of Sequence Variants Working Group, the Genetics in Medicine technical editing team, which I lead, conducted a systematic assessment of HGVS usage across the scientific literature (Lansdon et al., 2026).
- The results were striking: 100% of submitted manuscripts contained erroneous and/or missing HGVS and gene‑symbol naming/identification. The study further demonstrated that inconsistent and inaccurate implementation of HGVS hinders the findability of diagnostic evidence during routine database searches and the utilisation of specialised AI platforms designed to identify published accounts of interpreted genetic data, providing direct evidence that inaccurate nomenclature limits the capabilities of AIs used in clinical practice. Although the key focus of the review was published clinical data in journal articles, the errors are also endemic in critical databases used in the diagnostic workflow, such as ClinVar (Landrum et al., 2015). Failures at the most basic representational level undermine the entire ecosystem of personalised medicine and AI: without accurate variant descriptions, downstream systems cannot compensate; AI systems cannot learn reliably; and linkage becomes ambiguous, or impossible.
- Within the NHS Genomic Medicine Service, these problems are compounded by the heterogeneity and non‑standardisation of data‑storage systems. Across the seven Genomic Laboratory Hubs (GLHs), variant data is held in a mixture of information management platforms, custom databases, flat files, and Excel spreadsheets. While HGVS, HGNC, and reference‑sequence standards are used, differing annotation tools and inconsistent validation mean that identical variants frequently end up represented differently across GLH systems, leading to divergence and loss of concordance. Many laboratories do not store all levels of HGVS description, and many do not retain stable gene identifiers. Gene symbols may no longer align with current HGNC standards as they evolve over time, and failure to record the stable, immutable HGNC gene identifier hinders reliable gene‑level searches. Similarly, failing to capture the full genomic‑level variant description results in the loss of the original variant discovery and, in some cases, reference‑sequence drift makes the originating variant impossible to reproduce accurately, and therefore impossible to validate.
- Equally problematic is the inconsistency in reference sequence usage by GLHs. Some use Ensembl transcripts, others RefSeq, and many use a mixture of both depending on the historical evolution of their pipelines. In some workflows, including tools historically used by Genomics England, RefSeq transcripts are mis‑treated as pseudo‑genomic resources, collapsing transcript‑specific context and producing ambiguous or incorrect variant descriptions. Because these tools continue to assign the unique RefSeq identifiers to these pseudo‑genomic representations, they introduce positional errors and inaccurate mappings when variants are projected from the genome reference sequence onto the true RefSeq transcript coordinates. This results in variants that cannot be reliably matched across centres, and since RefSeq is used more prevalently on a global scale, these inaccuracies break interoperability with international databases and disrupt access to the wider evidence base.
- In the most severe cases, transcript processing distorts gene structures entirely, excluding sequences from analysis or introducing artificial content. For example, excluding bona‑fide coding exons from routine analyses, as seen in clinically relevant genes such as SHANK3 (HGNC:14294), or introducing artificial intronic and exonic content, as observed in RYBP (HGNC:10480). More subtle discrepancies occur as well, for example, on GRCh37[1], the gene NR2E3 (HGNC:7974) exhibits a one‑base‑pair alignment mismatch between transcript models that shifts downstream genomic coordinates. Although this was corrected in GRCh38, historical use of GRCh37 across the NHS means that disease-causing frameshift variants at this position may not have been scrutinised correctly and could have been missed. Likewise, small exon‑boundary mismatches in the RYR1 (HGNC:10483) gene on GRCh38 continue to generate inconsistent genomic to transcript projections across NHS laboratories.
- For the NHS to deliver personalised medicine and AI at a national scale, it must establish a central, standards‑aligned genomic database capable of receiving and integrating variant data from all GLHs. To populate such a system, historical variant data must be extracted from local databases, normalised, corrected, and updated. However, this work is currently underfunded and under‑resourced. Critically, the use of disparate tools to cleanse legacy data introduces further ambiguity rather than resolving it. Outsourced or general‑purpose tools frequently mis‑apply HGVS, mishandle RefSeq transcripts, or collapse transcript structures, and because each GLH relies on different solutions, their outputs diverge instead of converging. As a result, NHS healthcare scientists are often left to manually repair errors, frequently using specialist platforms such as VariantValidator to restore correct HGVS expressions, resolve Ensembl–RefSeq discrepancies, and normalise variant representations post hoc.
- This is neither efficient nor sustainable. A coherent national approach requires a validated, specialist egress tool capable of enforcing HGVS correctness, performing accurate mapping between Ensembl and RefSeq transcript sets, and providing robust variant normalisation. While VariantValidator does not yet generate VRS or SPDI identifiers, these capabilities are close to completion, and with formal collaboration between the NHS, Genomics England, the tool’s developers, and working with the professional societies with which VariantValidator is embedded, VRS and SPDI support can be integrated directly and to a globally accepted standard. By contrast, outsourcing to generic tools that replicate the same inaccuracies observed in Genomics England software risks embedding those errors into the national database. This is not a criticism of Genomics England’s efforts, but a recognition that every representation must be accurate at all levels if the foundation of the national dataset is to be stable, and effort is required to align Genomics England and NHS tooling with gold standard platforms and societies tackling these critical issues, rather than ignoring them. This ethos also applies to the generation of new data currently passing through the NHS Genomics Medicine Service and Genomics England pipelines, which must also be aligned to ensure accuracy going forward.
- Finally, this work requires sustained investment in the healthcare genomics bioinformatics workforce. Data harmonisation and standards compliance are currently under‑resourced, despite being essential to the success of a central NHS genomics database. Without specialist tools, unified standards, and adequately supported expert staff, the national dataset will remain fractured, and the promise of personalised medicine and AI will remain out of reach.
- A central NHS database must adhere explicitly to the emerging global technical standard for variant representation being developed collaboratively by the organisations listed in the introduction. Internally, the system may adopt optimal computational descriptors such as SPDI or VRS, but the underlying gene, transcript, and variant information must be robust, authoritative, and fully concordant with clinical HGVS expressions. Only then can computable IDs remain synchronised with the human‑readable nomenclature clinicians use for diagnosis.
- If implemented correctly, the benefits are substantial.
- First, single‑point variant findability becomes possible: any variant can be found with a single search, whether using HGVS, SPDI, VRS, or genomic coordinates, replacing the current reality of multiple discordant searches (Lansdon et al., 2026).
- Second, human–machine interoperability is achieved, enabling clinicians, researchers, and AI systems to reference the same underlying variant record reliably.
- Third, AI systems finally gain access to consistent, high‑quality training data, avoiding the propagation of errors that arise from inconsistent variant descriptions.
- Fourth, diagnostic rates increase through national‑scale evidence pooling.
- Fifth, and critically, a unified dataset enables the establishment of a national re‑analysis service, allowing automated reassessment of unsolved cases, notification of clinicians when new evidence emerges, and systematic reduction of the backlog of unresolved genomic cases, an area where the NHS currently lags.
- Finally, by adopting and enforcing these standards, the UK secures international leadership in clinical genomics.
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
- The National Institutes of Health (NIH) in the United States is already working directly with professional societies and specialist developers, including VariantValidator, to harmonise ClinVar submissions and their supporting tools. The UK must not fall behind. If the NHS continues to use tools that mis-handle HGVS, incorrectly map reference sequences, or apply inconsistent transcript logic, the result will be continued fragmentation, reduced diagnostic accuracy, and loss of global leadership.
- In summary, the essential research infrastructure required is a standards‑aligned national genomic database with validated data‑egress pipelines, strict nomenclature enforcement, and robust mapping between human‑readable and computable variant descriptors, supported by investment in the healthcare scientist workforce. Without this, the promises of personalised medicine and AI will remain unrealised.
Question 3a: Recommendations for HDR UK, Genomics England, and the Genomics AI Network
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
- HDR UK, Genomics England, and the Genomics AI Network should jointly mandate and enforce a national standard for variant representation that aligns with international best practice and professional technical standards. Central to this is the adoption of authoritative HGVS, HGNC, and stable transcript references, supported by bi-directional mappings to SPDI and VRS.
- These bodies should co‑commission a validated national software platform for data egress and harmonisation, developed in partnership with specialist platforms already recognised by leading global societies and quality bodies, including ACGS and EMQN. Investment in healthcare scientists with expertise in variant curation, bioinformatics, and standards compliance is essential.
- Finally, these bodies should collaborate actively with the international community via the professional societies to demonstrate that UK systems remain focused on adherence to standards-driven data in genomics and encourage others to contribute equally robust and interoperable data into global databases such as ClinVar and LOVD so that they enhance the UKs own diagnostic dataset.
Question 3b: Government progress in linking health data
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
- Government initiatives such as the 100,000 Genomes Project, the establishment of the NHS Genomic Medicine Service, and the development of Trusted Research Environments represent important progress. However, they have not yet delivered national-scale data linkage. Variant data across GLHs remains discordant, inconsistent, and in many cases non‑computable. There is no unified schema for variant representation, no enforced nomenclature standard, no reliable linkage between GLH variant data and EHR systems, and no national re‑analysis service. As a result, much of the evidence required for personalised medicine remains inaccessible or unlinked.
Question 3c: NHS digital and IT infrastructure barriers
- The NHS digital estate remains a major barrier to personalised medicine. EHR systems are fragmented, lack interoperability, and are generally unable to ingest structured genomic data, let alone computable variant representations such as HGVS, VRS, or SPDI. Many NHS Trusts operate legacy systems that cannot communicate with genomic laboratories or national platforms. Existing government interoperability initiatives do not address the specific technical requirements of genomic data, which include strict standardisation, transcript‑level resolution, and variant‑level precision. Without investment and enforcement at this level, personalised medicine pathways cannot scale.
Question 3d: Public trust and data governance
- Public trust in genomic data sharing depends on transparency, clear governance, visibility of data flows, and confidence that data are used safely and appropriately. A national genomic database must therefore include transparent audit trails, publicly visible data‑access logs, and clear explanations of how data are used to improve patient care. Strong oversight of industrial partnerships is essential to ensure that commercial use of NHS genomic data aligns with public expectations. Crucially, by enforcing internationally recognised standards and ensuring variant data is accurate, traceable, and clinically interpretable, the NHS strengthens public confidence in the safety and integrity of the genomic infrastructure.
References
Freeman PJ, Wagstaff JF, Fokkema IFAC, et al. Standardizing variant naming in literature with VariantValidator to increase diagnostic rates. Nature Genetics. 2024;56:2284–2286. https://doi.org/10.1038/s41588-024-01938-w
Higgins J, Dalgleish R, den Dunnen JT, et al. Verifying nomenclature of DNA variants in submitted manuscripts: Guidance for journals. Human Mutation. 2021;42:3–7. https://doi.org/10.1002/humu.24144
Lansdon LA, Porath B, Mori M, Miller DT, Dunham Drexler D, Wattenberg C, Richardson M, Steiner RD, Freeman PJ. Universal Presence of Gene/Variant Nomenclature Errors in Journal Manuscript Submissions. Clinical Chemistry. 2026;hvag010. https://doi.org/10.1093/clinchem/hvag010
Goar W, Babb L, Chamala S, Cline M, Freimuth RR, et al. Development and application of a computable genotype model in the GA4GH Variation Representation Specification. Genomics, Proteomics & Bioinformatics. 2022. https://doi.org/10.1142/9789811270611_0035
Danecek P, Auton A, Abecasis G, Albers CA, Banks E, et al. The variant call format and VCFtools. Bioinformatics. 2011;27(15):2156–2158. https://doi.org/10.1093/bioinformatics/btr330
Holmes JB, Moyer E, Phan L, Maglott D, Kattman B. SPDI: data model for variants and applications at NCBI. Bioinformatics. 2020;36(6):1902–1907. https://doi.org/10.1093/bioinformatics/btz856
Richards S, Aziz N, Bale S, Bick D, Das S, Gastier‑Foster J, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of ACMG and AMP. Genetics in Medicine. 2015;17(5):405–424. https://doi.org/10.1038/gim.2015.30
den Dunnen JT, Dalgleish R, Maglott DR, Hart RK, Greenblatt MS, et al. HGVS Recommendations for the Description of Sequence Variants: 2016 Update. Human Mutation. 2016;37:564–569. https://doi.org/10.1002/humu.22981
Köhler S, Carmody L, Vasilevsky N, Jacobsen JOB, Danis D, et al. Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources. Nucleic Acids Research. 2019;47(D1):D1018–D1027. https://doi.org/10.1093/nar/gky1105
Landrum MJ, Lee JM, Benson M, Brown G, Chao C, et al. ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Research. 2016;44(D1):D862–D868. https://doi.org/10.1093/nar/gkv1222