Written evidence from Peter Freeman, Lecturer in Healthcare Sciences (Clinical Bioinformatics and Genomics), The University of Manchester and support by Policy@Manchester (PMA0037)

Generative AI was used to assist in the initial drafting of this response. The author accepts responsibility for the contents of this submission.

Executive summary

 

  1. This submission responds specifically to Question 3: Health Data Research Infrastructure. Drawing on my technical leadership in genomic data standards, international editorial experience, and work training NHS bioinformaticians, it outlines the key systemic deficiencies in the UK’s health and genomic data infrastructure and provides concrete recommendations to ensure the NHS can fully realise the benefits of personalised medicine and AI.

Background

 

  1. I have over a decade of experience in genomic data annotation, clinical bioinformatics, and the development of internationally adopted frameworks for the accurate reporting of genetic variation. My work focuses on ensuring the accuracy, reproducibility, and interoperability of genomic data – essential foundations for personalised medicine and for the safe and effective deployment of AI across the NHS.
  2. Since 2019, I have been a major contributor to the NHS Scientist Training Programme in Clinical Bioinformatics (Genomics), leading national modules in diagnostic sequencing, variant annotation, and software engineering for clinical genomics. Through training healthcare scientists across all seven NHS Genomic Laboratory Hubs (GLHs), I have gained substantial insight into the technical, organisational, and infrastructural challenges that affect data quality, hinder adoption of best practice, and slow the integration of AI and personalised medicine into routine NHS care.
  3. I am the creator and Principal Investigator of VariantValidator, used across all seven GLHs and internationally to standardise variant representation, reduce diagnostic errors, and support compliance with global genomic reporting standards. Its impact has been documented in Nature Genetics (Freeman et al., 2024), where we demonstrated that improved nomenclature accuracy directly increases diagnostic rates.
  4. I serve as Lead Technical Editor at Genetics in Medicine and GIMO, where I lead a specialist editorial team dedicated to ensuring accurate data representation in peerreviewed genomics manuscripts. This work includes coauthoring a study (Lansdon et al., 2026) which systematically demonstrated the nearuniversal presence of nomenclature errors in submitted manuscripts and underscored the structural challenges caused by inconsistent genomic reporting – challenges that are amplified in AI training datasets.
  5. I am a senior member of 2the Reporting of Sequence Variants Working Group, part of the Human Genome Organisation (HUGO), where I coauthored guidance for verifying variant nomenclature in scientific manuscripts (Higgins et al., 2021). I am also a core contributor to the multiorganisation project “Standards for journals, authors and clinical laboratories for the reporting and sharing of interpreted genomic variation”, involving:
  1.  
  2.  
  3.  
  4.  
  5.  
  6.  
  7.  
  8.  
  1. This collaboration focuses on defining robust, interoperable, computationally reliable genomic data standards that support global consistency and enable AIreadiness.

 

 

 

 

 

 

Health Data Research Infrastructure

 

Question 3: What further research infrastructure is needed to support personalised medicine and AI, and where are the gaps?

 

  1. Personalised genomic medicine depends on robust linkage between genomic variation and clinical evidence (Richards et al., 2015); phenotypic datasets, which link physical traits to genetic data (Köhler et al., 2019); and healthsystem data, all encoded in a standardised and fully computable framework. However, a current core obstacle lies not in sequencing capacity or computational power, but in the fragmented and inconsistent ways that genomic variation is represented, stored, and communicated across the NHS and globally.

 

  1. In clinical practice, multiple systems are required and utilised to represent and exchange genomic data, each with its own strengths and limitations. The Variant Call Format (VCF) (Danecek et al., 2011) remains widely used in diagnostic workflows, but frequently loses essential metadata required for linkage, interpretation, and reproducibility. Newer computable frameworks, such as the GA4GH Variation Representation Specification (VRS) (Goar et al., 2022), have been developed to provide globally interoperable, unambiguous representations of genomic variation, and the SPDI model (Holmes et al., 2020) has emerged as another attempt to deliver stable, computationready descriptions of variants. Despite their computational strengths, these models remain insufficient for human interpretation, particularly because they cannot effectively represent the consequences of genetic variation at a sufficient depth as required in clinical practice. Consequently, the globally adopted and clinically mandated nomenclature for reporting sequence variants remains the Human Genome Variation Society (HGVS) standard (den Dunnen et al., 2016).

 

  1.         HGVS offers a humanreadable and clinically interpretable description of variants and is the nomenclature understood by clinicians and clinical scientists. It is explicitly mandated by the ACMG and the AMP in guidelines for variant interpretation (Richards et al., 2015), which underpin UK practice through modification and adoption by the ACGS and the EMQN.

 

  1.         Because the different naming systems were developed for different purposes, and serve fundamentally different audiences, inconsistencies arise at exactly the points where data must be linked. Variant identifiers cannot be reliably matched between systems without stable, standardscompliant transformation pipelines. This represents a foundational barrier to personalised medicine and AI, because effective linkage requires variant descriptions that are consistent, computable, and traceable across sequencing pipelines, clinical reports, electronic health records (EHRs), phenotypic repositories, and nationalscale databases.

 

  1.         Despite the longstanding mandate to use HGVS nomenclature – and the fact that HGVS remains the globally adopted clinical standard and the primary entry point for retrieving diagnostic evidence from the literature and curated databases – its application remains inconsistent across the NHS and globally. HGVS descriptions are often omitted from reports or substituted with imprecise or outdated representations, and where it is used, it is frequently incorrect. As part of the HUGO Reporting of Sequence Variants Working Group, the Genetics in Medicine technical editing team, which I lead, conducted a systematic assessment of HGVS usage across the scientific literature (Lansdon et al., 2026).

 

  1.         The results were striking: 100% of submitted manuscripts contained erroneous and/or missing HGVS and genesymbol naming/identification. The study further demonstrated that inconsistent and inaccurate implementation of HGVS hinders the findability of diagnostic evidence during routine database searches and the utilisation of specialised AI platforms designed to identify published accounts of interpreted genetic data, providing direct evidence that inaccurate nomenclature limits the capabilities of AIs used in clinical practice. Although the key focus of the review was published clinical data in journal articles, the errors are also endemic in critical databases used in the diagnostic workflow, such as ClinVar (Landrum et al., 2015). Failures at the most basic representational level undermine the entire ecosystem of personalised medicine and AI: without accurate variant descriptions, downstream systems cannot compensate; AI systems cannot learn reliably; and linkage becomes ambiguous, or impossible.

 

  1.         Within the NHS Genomic Medicine Service, these problems are compounded by the heterogeneity and nonstandardisation of datastorage systems. Across the seven Genomic Laboratory Hubs (GLHs), variant data is held in a mixture of information management platforms, custom databases, flat files, and Excel spreadsheets. While HGVS, HGNC, and referencesequence standards are used, differing annotation tools and inconsistent validation mean that identical variants frequently end up represented differently across GLH systems, leading to divergence and loss of concordance. Many laboratories do not store all levels of HGVS description, and many do not retain stable gene identifiers. Gene symbols may no longer align with current HGNC standards as they evolve over time, and failure to record the stable, immutable HGNC gene identifier hinders reliable genelevel searches. Similarly, failing to capture the full genomiclevel variant description results in the loss of the original variant discovery and, in some cases, referencesequence drift makes the originating variant impossible to reproduce accurately, and therefore impossible to validate.

 

  1.         Equally problematic is the inconsistency in reference sequence usage by GLHs. Some use Ensembl transcripts, others RefSeq, and many use a mixture of both depending on the historical evolution of their pipelines. In some workflows, including tools historically used by Genomics England, RefSeq transcripts are mistreated as pseudogenomic resources, collapsing transcriptspecific context and producing ambiguous or incorrect variant descriptions. Because these tools continue to assign the unique RefSeq identifiers to these pseudogenomic representations, they introduce positional errors and inaccurate mappings when variants are projected from the genome reference sequence onto the true RefSeq transcript coordinates. This results in variants that cannot be reliably matched across centres, and since RefSeq is used more prevalently on a global scale, these inaccuracies break interoperability with international databases and disrupt access to the wider evidence base.

 

 

  1.         In the most severe cases, transcript processing distorts gene structures entirely, excluding sequences from analysis or introducing artificial content. For example, excluding bonafide coding exons from routine analyses, as seen in clinically relevant genes such as SHANK3 (HGNC:14294), or introducing artificial intronic and exonic content, as observed in RYBP (HGNC:10480). More subtle discrepancies occur as well, for example, on GRCh37[1], the gene NR2E3 (HGNC:7974) exhibits a onebasepair alignment mismatch between transcript models that shifts downstream genomic coordinates. Although this was corrected in GRCh38, historical use of GRCh37 across the NHS means that disease-causing frameshift variants at this position may not have been scrutinised correctly and could have been missed. Likewise, small exonboundary mismatches in the RYR1 (HGNC:10483) gene on GRCh38 continue to generate inconsistent genomic to transcript projections across NHS laboratories.

 

  1.         For the NHS to deliver personalised medicine and AI at a national scale, it must establish a central, standardsaligned genomic database capable of receiving and integrating variant data from all GLHs. To populate such a system, historical variant data must be extracted from local databases, normalised, corrected, and updated. However, this work is currently underfunded and underresourced. Critically, the use of disparate tools to cleanse legacy data introduces further ambiguity rather than resolving it. Outsourced or generalpurpose tools frequently misapply HGVS, mishandle RefSeq transcripts, or collapse transcript structures, and because each GLH relies on different solutions, their outputs diverge instead of converging. As a result, NHS healthcare scientists are often left to manually repair errors, frequently using specialist platforms such as VariantValidator to restore correct HGVS expressions, resolve Ensembl–RefSeq discrepancies, and normalise variant representations post hoc.

 

  1.         This is neither efficient nor sustainable. A coherent national approach requires a validated, specialist egress tool capable of enforcing HGVS correctness, performing accurate mapping between Ensembl and RefSeq transcript sets, and providing robust variant normalisation. While VariantValidator does not yet generate VRS or SPDI identifiers, these capabilities are close to completion, and with formal collaboration between the NHS, Genomics England, the tool’s developers, and working with the professional societies with which VariantValidator is embedded, VRS and SPDI support can be integrated directly and to a globally accepted standard. By contrast, outsourcing to generic tools that replicate the same inaccuracies observed in Genomics England software risks embedding those errors into the national database. This is not a criticism of Genomics England’s efforts, but a recognition that every representation must be accurate at all levels if the foundation of the national dataset is to be stable, and effort is required to align Genomics England and NHS tooling with gold standard platforms and societies tackling these critical issues, rather than ignoring them. This ethos also applies to the generation of new data currently passing through the NHS Genomics Medicine Service and Genomics England pipelines, which must also be aligned to ensure accuracy going forward.

 

  1.         Finally, this work requires sustained investment in the healthcare genomics bioinformatics workforce. Data harmonisation and standards compliance are currently underresourced, despite being essential to the success of a central NHS genomics database. Without specialist tools, unified standards, and adequately supported expert staff, the national dataset will remain fractured, and the promise of personalised medicine and AI will remain out of reach.

 

  1.         A central NHS database must adhere explicitly to the emerging global technical standard for variant representation being developed collaboratively by the organisations listed in the introduction. Internally, the system may adopt optimal computational descriptors such as SPDI or VRS, but the underlying gene, transcript, and variant information must be robust, authoritative, and fully concordant with clinical HGVS expressions. Only then can computable IDs remain synchronised with the humanreadable nomenclature clinicians use for diagnosis.

 

  1.         If implemented correctly, the benefits are substantial.
  1.  
  2.  
  3.  
  4.  
  5.  
  6.  
  7.  
  8.  
  9.  
  10.          
  11.          
  12.          
  13.          
  14.          
  15.          
  16.          
  17.          
  18.          
  19.          
  20.          
  21.          
  22.          
  23.          
  24.         The National Institutes of Health (NIH) in the United States is already working directly with professional societies and specialist developers, including VariantValidator, to harmonise ClinVar submissions and their supporting tools. The UK must not fall behind. If the NHS continues to use tools that mis-handle HGVS, incorrectly map reference sequences, or apply inconsistent transcript logic, the result will be continued fragmentation, reduced diagnostic accuracy, and loss of global leadership.
  25.         In summary, the essential research infrastructure required is a standardsaligned national genomic database with validated dataegress pipelines, strict nomenclature enforcement, and robust mapping between humanreadable and computable variant descriptors, supported by investment in the healthcare scientist workforce. Without this, the promises of personalised medicine and AI will remain unrealised.

Question 3a: Recommendations for HDR UK, Genomics England, and the Genomics AI Network

 

  1.  
  2.  
  3.  
  4.  
  5.  
  6.  
  7.  
  8.  
  9.  
  10.          
  11.          
  12.          
  13.          
  14.          
  15.          
  16.          
  17.          
  18.          
  19.          
  20.          
  21.          
  22.          
  23.          
  24.          
  25.          
  26.         HDR UK, Genomics England, and the Genomics AI Network should jointly mandate and enforce a national standard for variant representation that aligns with international best practice and professional technical standards. Central to this is the adoption of authoritative HGVS, HGNC, and stable transcript references, supported by bi-directional mappings to SPDI and VRS.

 

  1.         These bodies should cocommission a validated national software platform for data egress and harmonisation, developed in partnership with specialist platforms already recognised by leading global societies and quality bodies, including ACGS and EMQN. Investment in healthcare scientists with expertise in variant curation, bioinformatics, and standards compliance is essential.

 

  1.         Finally, these bodies should collaborate actively with the international community via the professional societies to demonstrate that UK systems remain focused on adherence to standards-driven data in genomics and encourage others to contribute equally robust and interoperable data into global databases such as ClinVar and LOVD so that they enhance the UKs own diagnostic dataset.

Question 3b: Government progress in linking health data

 

  1.  
  2.  
  3.  
  4.  
  5.  
  6.  
  7.  
  8.  
  9.  
  10.          
  11.          
  12.          
  13.          
  14.          
  15.          
  16.          
  17.          
  18.          
  19.          
  20.          
  21.          
  22.          
  23.          
  24.          
  25.          
  26.          
  27.          
  28.          
  29.         Government initiatives such as the 100,000 Genomes Project, the establishment of the NHS Genomic Medicine Service, and the development of Trusted Research Environments represent important progress. However, they have not yet delivered national-scale data linkage. Variant data across GLHs remains discordant, inconsistent, and in many cases noncomputable. There is no unified schema for variant representation, no enforced nomenclature standard, no reliable linkage between GLH variant data and EHR systems, and no national reanalysis service. As a result, much of the evidence required for personalised medicine remains inaccessible or unlinked.

Question 3c: NHS digital and IT infrastructure barriers

 

  1.         The NHS digital estate remains a major barrier to personalised medicine. EHR systems are fragmented, lack interoperability, and are generally unable to ingest structured genomic data, let alone computable variant representations such as HGVS, VRS, or SPDI. Many NHS Trusts operate legacy systems that cannot communicate with genomic laboratories or national platforms. Existing government interoperability initiatives do not address the specific technical requirements of genomic data, which include strict standardisation, transcriptlevel resolution, and variantlevel precision. Without investment and enforcement at this level, personalised medicine pathways cannot scale.

Question 3d: Public trust and data governance

 

  1.         Public trust in genomic data sharing depends on transparency, clear governance, visibility of data flows, and confidence that data are used safely and appropriately. A national genomic database must therefore include transparent audit trails, publicly visible dataaccess logs, and clear explanations of how data are used to improve patient care. Strong oversight of industrial partnerships is essential to ensure that commercial use of NHS genomic data aligns with public expectations. Crucially, by enforcing internationally recognised standards and ensuring variant data is accurate, traceable, and clinically interpretable, the NHS strengthens public confidence in the safety and integrity of the genomic infrastructure.

 

References

Freeman PJ, Wagstaff JF, Fokkema IFAC, et al. Standardizing variant naming in literature with VariantValidator to increase diagnostic rates. Nature Genetics. 2024;56:2284–2286. https://doi.org/10.1038/s41588-024-01938-w

 

Higgins J, Dalgleish R, den Dunnen JT, et al. Verifying nomenclature of DNA variants in submitted manuscripts: Guidance for journals. Human Mutation. 2021;42:3–7. https://doi.org/10.1002/humu.24144

 

Lansdon LA, Porath B, Mori M, Miller DT, Dunham Drexler D, Wattenberg C, Richardson M, Steiner RD, Freeman PJ. Universal Presence of Gene/Variant Nomenclature Errors in Journal Manuscript Submissions. Clinical Chemistry. 2026;hvag010. https://doi.org/10.1093/clinchem/hvag010

 

Goar W, Babb L, Chamala S, Cline M, Freimuth RR, et al. Development and application of a computable genotype model in the GA4GH Variation Representation Specification. Genomics, Proteomics & Bioinformatics. 2022. https://doi.org/10.1142/9789811270611_0035

 

Danecek P, Auton A, Abecasis G, Albers CA, Banks E, et al. The variant call format and VCFtools. Bioinformatics. 2011;27(15):2156–2158. https://doi.org/10.1093/bioinformatics/btr330

 

Holmes JB, Moyer E, Phan L, Maglott D, Kattman B. SPDI: data model for variants and applications at NCBI. Bioinformatics. 2020;36(6):1902–1907. https://doi.org/10.1093/bioinformatics/btz856

 

Richards S, Aziz N, Bale S, Bick D, Das S, GastierFoster J, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of ACMG and AMP. Genetics in Medicine. 2015;17(5):405–424. https://doi.org/10.1038/gim.2015.30

 

den Dunnen JT, Dalgleish R, Maglott DR, Hart RK, Greenblatt MS, et al. HGVS Recommendations for the Description of Sequence Variants: 2016 Update. Human Mutation. 2016;37:564–569. https://doi.org/10.1002/humu.22981

 

Köhler S, Carmody L, Vasilevsky N, Jacobsen JOB, Danis D, et al. Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources. Nucleic Acids Research. 2019;47(D1):D1018–D1027. https://doi.org/10.1093/nar/gky1105

 

Landrum MJ, Lee JM, Benson M, Brown G, Chao C, et al. ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Research. 2016;44(D1):D862–D868. https://doi.org/10.1093/nar/gkv1222

 

 


[1] Genome Reference Consortium Human Build 37; a major human reference genome assembly released in 2009, widely used for clinical sequencing and genetic studies.