Written evidence from Mendelian (PMA0094)

 

Submitted by: Dr Peter Fish, CEO, Mendelian

This submission was drafted by a human with AI in the loop

Executive Summary

A critical bottleneck in the UK's ambition to lead in personalised medicine is the systemic inability to efficiently identify the precise patient cohorts who stand to benefit from approved treatments, clinical trials, and next-generation therapies. Patient discovery is currently governed by clinical serendipity, the chance encounter between a patient and the right specialist, rather than by systematic, data-driven capability.

This identification gap creates an R&D barrier, a commercialisation risk, and means patients never benefit from the enhanced management that personalisation promises. Many treatments are already MHRA-approved and NICE-recommended yet cannot reach their intended patients at scale. Without the ability to surface hidden cohorts, the UK is increasingly perceived as an unviable launch market for precision innovators. The NHS must transition from passive discovery to a proactive, AI-enabled, systematic eligibility assessment capability.

Three Recommendations

 

  1. Fix the data issue

         Resolve the data controller fragmentation that prevents GP, hospital, and genomic data from being linked and acted upon.

         Mandate the adoption of common data models (like OMOP) across national data infrastructure to enable validated tools to operate consistently.

         Require richer, multi-ontology data capture in secondary care aligned with FHIR and OMOP standards.

         Establish a governed pathway for AI access to unstructured clinical data under appropriate privacy-preserving conditions.

         Ensure structured omics data linking to patient records in near-real time.

 

  1. Fix the workforce issue

         Establish nationally and regionally commissioned Proactive Clinical Delivery Teams (potentially within the Genomic Medicine Service (GMS) and / or Highly Specialised Services (HSS)).

         These teams must hold patient-level clinical data access rights and clinical decision-making authority to take asynchronous, proactive action on AI-generated insights, operating independently of the frontline NHS workforce and without creating additional burden on GPs or hospital clinicians.

         Commission these as funded NHS services with a direct patient contact mandate; they cannot remain unfunded activities absorbed by an already overstretched care system.

 

  1. Adopt and deploy AI technology nationally

         Commission MendelScan (and others like it) as a nationally deployed NHS intelligence layer.

         Fund a Direct Care mandate enabling validated algorithms to flag at-risk patients through pre-approved, nationalised pathways, superseding research-only data access restrictions.

About Mendelian & MendelScan

Mendelian is a UK health data analytics company founded in 2015. Our product, MendelScan, is an MHRA-registered Class I MDD software medical device functioning as a clinical decision support system for rare disease case-finding. It analyses electronic health record (EHR) data across entire patient populations, identifying potential clinical patterns of hundreds of rare diseases and flagging patients for clinical human-in-the-loop review before patient recontact. The technology has already been used in the NHS to apply case-finding algorithms on over 10-million patient records, resulting in 1,500 patients being identified for further review.

As common diseases are further stratified by genomic, molecular, and phenotypic sub-classification, many resulting subgroups will themselves fall below the EU rare disease prevalence threshold of 1 in 2,000. The precision medicine of tomorrow is, in this sense, the rare disease landscape of today. The infrastructure built now for rare disease patient-finding will underpin personalised oncology, stratified cardiovascular care, and pharmacogenomics.

In 2023, Mendelian was selected as one of the NHS AI in Health and Care Award winners (Round 3, NIHR ref: AI_AWARD02579), run jointly by the NHS AI Lab, the Accelerated Access Collaborative, and NIHR. The award funded a large-scale study structured around the NICE Evidence Standards Framework for digital health technologies. The study ran from April 2023 to November 2024.

MendelScan was an exemplar project within the NHS Genomic AI Network (GAIN, Work package 3b), operating in partnership with the NHS Genomic Medicine Service (NHS GMS) across multiple clinical areas including: inherited retinal disease, inflammatory bowel disease, epilepsy, and cancer predisposition syndromes. GAIN was a 2-year project funded until March 2026. It has not been formally continued in its current structure.

Submission

Personalised Medicine and AI: the science and technology

 

Personalised medicine and AI: the scientific background

Advanced therapies: The science underpinning personalised medicine has reached clinical viability. Gene therapies like Luxturna are already MHRA-approved and NICE-recommended. At the frontier, antisense oligonucleotide (ASO) therapies represent personalised medicine at its most literal; a therapy designed for a single patient or a small cohort sharing the same mutation. EveryONE Medicines (now sadly defunct) was building a scalable platform for such therapies across various rare diseases, and was an active participant in the UK Rare Therapies Launchpad.

AI-driven case-finding: As part of Mendelian’s AI Award Study, 42 older MendelScan algorithms were validated on 23 million NHS primary care EHRs. These almost all achieved a specificity above 99.95%. For most diseases studied, 50–75% of patients were flaggable before the diagnosis code appeared, demonstrating that AI can surface clinical signals well ahead of current practice.

As an example, the algorithm for DiGeorge syndrome (22q11.2 deletion syndrome), demonstrated a sensitivity of 25.0%, successfully flagging one in every four known cases that are often overlooked in routine clinical practice. Its specificity was remarkably high at over 99.99%, which is vital for preventing alert fatigue by ensuring that healthy patients are rarely misidentified. With an adjusted positive predictive value of 13.3%, the algorithm suggests that approximately one in eight flagged individuals will have a confirmed diagnosis upon further clinical review. Strategically, the algorithm avoids focusing on "florid" manifestations like Tetralogy of Fallot, which are usually diagnosed at birth, and instead targets "non-obvious" combinations of features. It scans for primary indicators such as palatal defects, laryngomalacia, and immunodeficiencies, while also considering supportive features like short stature and neuropsychiatric symptoms. By systematically analysing EHR data, the platform acts as a critical safety net for patients who have missed standard paediatric screenings. Ultimately, this algorithmic approach provides a scalable and efficient method for reducing the profound diagnostic delays typically associated with the syndrome.

AI genomics analysis: Expanding on the EveryONE Medicines example, the identification of patients with deep intronic splice site mutations amenable to ASO therapy is not routinely conducted in the NHS genomics pipeline. Patients with treatable conditions go undiagnosed not because the science is inadequate, but because the analytical infrastructure fails to surface the relevant variants. AI tools, such as SpliceAI, can predict the pathogenicity of such variants with high confidence.

Significant near-term opportunities:

  1. Ramp up systematic patient identification for treatments that already exist by centrally commissioning and mandating the use of tools like MendelScan in proactive care projects. The Luxturna case is the clearest illustration: at the point of NICE approval in 2019, fewer than half of the estimated 86 eligible patients in England had been identified and when it came to treating patients, only a fraction were diagnosed to the right level or offered the option. The regulatory and commissioning system worked; the identification system did not. This pattern will repeat across every new precision therapy that reaches the market.
  2. Commission regional and / or national dedicated proactive precision medicine clinical delivery teams (potentially under the Genomic Medicine Service), with the authority and patient-level data access rights to act on AI-generated insights and connect identified patients to treatment pathways. Shift the burden off the NHS frontline.
  3. Ensure structured genomic (and other omics) data is made routinely available at a clinical level. In the near-term, this can be facilitated for early clinical (not research) projects, via linkage between siloes.

Currently Genomics England facilitates most whole genome and whole exome sequencing testing in England. Once sequencing is complete, results are returned to Genomic Laboratory Hubs (GLHs) in PDF format, as static, non-structured documents, with very limited information, and variant calling files. The GLHs only pass the PDF report back to the test requesting clinician. Raw data (and significant analysis power) is siloed within Genomics England, variant calling files are siloed in GLHs and, unless a very clear pathogenic diagnosis was made, very limited information is returned to the clinician and into the patient EHR.

This creates multiple levels of failure, all of which should be addressed:

    1. Variants below standard reporting thresholds are not flagged even when present in raw data or clinically relevant
    2. Test data is not routinely re-analysed
    3. When results are returned to referring clinicians and hospital EHR systems, they are stored as unstructured text, making it nearly impossible to systematically identify patients retrospectively based on their genetic findings.

These are foundational infrastructure problems. Investment in structured, machine-readable genomic result formats, and in the linkage of genomic records to EHRs, is a precondition for almost all downstream personalised medicine research.

Health data research infrastructure

 

The main barriers are structural and political, and, to a lesser degree, technical. The components needed are largely understood; what is missing is the governance will, commissioning decisions, and standardisation mandates to implement them coherently.

Research vs clinical utility: Much of the data aggregation efforts underway in the NHS aim to make data available for research purposes. This is conducted under research ethics, and generally all data is de-identified. Personalised medicine requires re-contactable patient data for clinical trial recruitment and direct patient care. We need to prioritise transitioning from a research-output philosophy to a clinical-intervention philosophy. Secure Data Environments (SDEs) have been built overarchingly for research purposes, but SDEs suffer from inconsistent datasets and a design philosophy focused on research output rather than clinical intervention.

Data controller fragmentation: The most significant blocker in the NHS data landscape is the tension between data controllers at different levels of the system. GP practices are independent data controllers. NHS Trusts hold their own data under separate governance. ICBs have regional aggregation functions but limited ability to compel cooperation. NHS England holds national mandates, but cannot easily operationalise them at the point of care. This fragmentation keeps patient data siloed across GP, hospital, and genomic records, governed by separate controllers with different risk appetites, capabilities, resources and goals.

Record depth, granularity, recency and completeness: AI case-finding is only as good as the data it operates on. This statement is so obvious it risks being overlooked in policy discussions that focus on algorithm performance, governance frameworks, and commissioning models. The practical consequences of poor data recency, incomplete records, and insufficient granularity are severe enough to undermine even the most technically sophisticated AI deployment. A system that flags patients based on incomplete, outdated, or coarsely coded records will generate outputs that are wrong, wasteful, and, critically, corrosive to the clinician trust that any successful AI deployment depends upon.

Recency: The moving target problem: For personalised medicine to function efficiently, the records supplied to any AI tool must reflect the current clinical state of the patient, not a historical snapshot. A flagged patient who has already received the recommended test, treatment, or referral, but whose record does not reflect this because the relevant entry sits in an unstructured clinical note, a discharge letter, or a secondary care system not yet reconciled with the primary care record, represents not just a wasted clinical interaction but an active failure of the system. If a clinician reviewing an AI-generated flag must spend time establishing that the recommended action was already taken months ago in a different care setting, the AI has created work rather than reducing it. Near-real-time data linkage across care settings is therefore not a nice-to-have feature of the AI infrastructure, it is a precondition for the AI being useful at scale.

Completeness: Data access is often only granted to setting-specific EHRs and often only the structured aspect of the EHRs. This creates a dangerous illusion of completeness where an algorithm may analyse what it believes is the patient's entire record and find no evidence of a relevant test or diagnosis, when in fact that test was performed, the result was significant, and the finding was documented in a clinic letter that no structured field ever captured.

NHS EHR systems capture a fraction of clinically relevant activity in structured, queryable form. The majority of clinical documentation, consultation notes, referral letters, discharge summaries, clinic correspondence, imaging reports, and genomic results, exists as unstructured free text or PDF attachments that are invisible to AI tools with access to structured data alone.

Granularity: The ICD-10 ceiling: Even where structured data exists and is current, the granularity of that data in NHS secondary care is frequently insufficient to support the discriminatory power that personalised medicine requires. Secondary care structured EHR data generally uses ICD-10 for diagnoses and OPCS-4 for procedures. ICD-10 was designed for epidemiological reporting and administrative billing, not for the precise phenotypic characterisation that rare disease identification or genomic stratification demands. The UK's continued reliance on the international base version, with approximately 14,000 codes, compares poorly with the ICD-10-CM Clinical Modification used in the United States, which contains over 70,000 codes including far more granular classifications for disease subtypes, hereditary conditions, and genetically defined disorders. This is not a technical inevitability; it is a policy choice that the NHS has made by default rather than by design.

The practical consequence of this coding ceiling is that secondary care data tells an AI what broad diagnostic category a patient falls into, but rarely the specific subtype, causative mechanism, or severity that personalised medicine decisions depend upon. A patient coded as having epilepsy may have a developmental epileptic encephalopathy caused by a specific gene variant that renders them eligible for a targeted therapy, but ICD-10 cannot express that distinction. A patient coded with a leukodystrophy may have a specific subtype amenable to gene therapy, but the code applied captures only the broad disease family. The AI is therefore working with data that has had much of its clinical signal compressed out of it before the algorithm ever runs.

The multi-ontology solution. Personalised medicine ultimately requires a plurality of structured ontologies working in concert. Mondo provides a rare disease classification framework. SNOMED-CT offers clinical terminology at a level of specificity far beyond ICD-10 (currently used in NHS primary care). LOINC provides standardised coding for laboratory results. Both FHIR and OMOP explicitly call for data expressed across multiple ontologies, precisely because no single coding system can capture the clinical reality that personalised medicine needs to interrogate. The NHS's continued reliance on ICD-10 alone in secondary care is not just a coding inconvenience, it is a structural constraint on what AI can find in NHS data, however sophisticated the algorithm.

Note: ICD-11 offers a meaningful improvement, with post-coordination capabilities that allow combinations of codes to express clinical detail no single code could capture, anatomical site, causative gene, mode of , and severity together. Active planning for ICD-11 adoption, with explicit engagement with its post-coordination capabilities for rare disease and genomic coding, should be a national commitment. But ICD-11 alone is not sufficient. It is a better administrative coding system, not a clinical phenotyping framework. The NHS must plan for ICD-11 adoption as one component of a broader shift toward multi-ontology data capture in secondary care, mandating richer structured data aligned with FHIR and OMOP requirements, and establishing a governed pathway for AI access to unstructured clinical data under appropriate privacy-preserving conditions, so that the clinical signal currently buried in free text can be made available to the tools designed to act on it.

Taken together, the recency, completeness, and granularity problems mean that NHS data, as currently structured and supplied, systematically understates the clinical picture of every patient an AI tool analyses. Flags will be generated for patients who have already been managed. Patients will be missed because their most relevant clinical information exists only in unstructured text. And the phenotypic signal available to discriminate between patients who do and do not have a condition of interest will be coarser than the clinical reality warrants. Each of these problems is individually addressable. Together, they represent a data quality infrastructure challenge that is as consequential for the success of AI-enabled personalised medicine as any question of algorithm design, governance framework, or commissioning model, and one that deserves equivalent policy attention.

The Single Patient Record (SPR) and Unified Genomic Record (UGR) programmes represent the NHS's most significant attempts to address some of these issues. The SPR's ambition, a single, secure, authoritative record owned by the patient, accessible across care settings, directly supports the infrastructure personalised medicine requires. These should both be implemented in a way that enables AI-driven clinical tools to act on linked genomic and phenotypic data.

Omics data: De-silo omics data and allow raw, or at least complete test data, to be linked to patient records. See the genomic recommendation above.

The hidden danger in self-sustaining data entities: When the Government establishes data organisations within the NHS ecosystem, whether national bodies like Genomics England or regional SDEs, it typically does so with a dual mandate; serve the public good, and eventually become financially self-sustaining. These two objectives are presented as compatible. In practice, they are frequently in direct tension. An organisation that must generate its own revenue from data assets will, rationally and inevitably, begin to treat those assets as proprietary. Data that should flow freely back into patient records, clinical care pathways, and AI-enabled services instead becomes a commercial resource to be packaged, licensed, and sold. The public good mandate does not disappear, but it competes, on unequal terms, with the institutional survival imperative.

The organisations created for patient benefit and to solve the NHS data problem become, over time, part of the data problem. Genomics England is an instructive example. Established to sequence and interpret the genomes of NHS patients and return clinically actionable findings to the healthcare system, it has also developed significant commercial data licensing activities. SDEs present a parallel problem at regional level. Built partly with NHS infrastructure funding and partly with ambitions toward financial sustainability, SDEs have a structural interest in controlling access to the data they hold, because access is what they have to sell. This creates friction at precisely the point where friction is most damaging, the interface between research-grade data insight and clinical intervention. Part of this is genuine governance complexity, but part of it reflects an institutional boundary that serves the SDE's own operational model as much as it serves patient safety. An SDE that freely enabled validated AI tools to generate clinically actionable outputs and route them directly into care pathways would be undermining its own position as the necessary intermediary between data and clinical use.

The problem extends further down the system than national bodies and regional infrastructure. Some NHS Trusts, GP supergroups, and primary care networks have developed their own data assets, aggregated patient records, research databases, data lakes, and commercial arrangements with technology partners. Data sharing agreements that should be straightforward become protracted negotiations. Access terms are structured to preserve institutional leverage. The cumulative effect is an NHS data landscape in which the entities nominally responsible for enabling data-driven care have become, in significant part, obstacles to it through the entirely predictable consequence of asking public bodies to survive commercially on assets that should, by rights, belong to the patients who generated them.

The solution is not to abandon the data organisation model; it is to be explicit about the conflict of interest it creates and to design funding models that remove it. The principle should be simple - any data derived from NHS patient care should flow freely back into NHS patient care.

Public trust: As part of Mendelian’s AI Award Study, we conducted a public sentiment survey (n=254), several important points stood out:

       89% endorsed use within their own GP practice or hospital, 93% supported use of this technology across the NHS

       Over 90% said they would feel grateful if proactively contacted based on AI analysis of their records, even without prior explicit consent for that specific search

       Positive endorsement was consistent across age groups, educational levels, and across those both familiar and unfamiliar with rare disease

       Respondents' primary concerns were not about AI itself, but about NHS capacity to act on findings. This is a significant finding, the public is not asking to slow down data access for AI, it is asking to ensure the NHS has the workforce to do something useful with what AI finds. This directly reinforces the case for properly resourced proactive clinical delivery teams, not for data access restrictions.

       Patients who had experienced the diagnostic odyssey were unambiguous in their support

Innovation in the NHS: deployment

 

Deployment in practice

 

Patient-doctor vs patient-NHS relationship: A fundamental tension runs through the ambition to deploy AI-enabled personalised medicine at population level: the entity with the ambition to act is not the entity with the clinical relationship, and the entity with the clinical relationship lacks the capacity and resources to act. If data can be accessed and relevant patients identified at a population-level, under current governance frameworks, the commissioning body cannot take direct clinical action. GP practices and Trusts hold the clinical relationship, but their goals and resources are misaligned with proactive, population-level action.

This tension is compounded by a further reality: the patient-doctor relationship in the NHS is largely a fiction. Clinicians rotate, retire, and move between practices and Trusts. The meaningful clinical relationship is not patient-doctor but patient-primary care network in primary care and patient-hospital in secondary care. The institution holds the relationship, not the individual clinician, and institutions have no inherent motivation to act proactively on behalf of patients they have not yet encountered. Recognising this reframes the governance problem: if continuity is already institutional rather than individual, the case for allowing a nationally commissioned clinical team to act on AI-generated insights, rather than routing every action through a GP or consultant with no meaningful prior relationship with the patient, becomes considerably stronger.

Three routes exist for resolving this tension:

  1. Mandate action through contractual obligations or QOF-style incentives, preserving existing structures but adding burden to an already overstretched workforce
  2. Use financial and other incentives to align motivation, but this does not resolve the underlying capacity problem
  3. The operational team model recommended above. This creates a separate function whose entire mandate is converting AI insight into clinical action, without routing that action through primary or secondary care. This model is not unprecedented. It mirrors the structure of existing nationally commissioned services, cancer screening programmes, for example, where a central function acts proactively on behalf of a defined patient population independently of whether any individual GP has identified a need. The infrastructure and governance precedent exists. What is missing is the commissioning will to apply it to AI-enabled personalised medicine case-finding, and the explicit Direct Care mandate that would allow validated algorithms to trigger clinical action through a pre-approved, nationalised pathway rather than being filtered through every separate ICB and an overstretched frontline care system.

Cancer screening as a structural analogy: National cancer screening programmes demonstrate that the NHS can commission proactive, population-level clinical action at scale, independently of whether any individual GP or hospital consultant has identified a need. These programmes have a central function, a funded workforce, a direct patient contact mandate, and governance frameworks that operate across data controller boundaries. The same model applied to AI-enabled personalised medicine case-finding would resolve the patient-NHS relationship tension, the data controller deadlock, and the workforce gap simultaneously.

MendelScan’s Genomic AI Network scaling fail: The GAIN project demonstrated that MendelScan is technically scalable across NHS infrastructure, but technical readiness and operational deployment are not the same thing, and the gap between them is where GAIN stalled.

Optum & TPP: The two dominant primary care EHR providers in England, Optum (previously EMIS) and TPP, present distinct but equally significant barriers to scalable AI deployment. Mendelian has explored deploying MendelScan directly inside the EMIS platform, a technically feasible proposition that has so far proven prohibitively expensive, though recent indications from Optum suggest a more accessible payment model may be in development. Mendelian is accredited to extract EMIS data via the NHS IM1 mechanism. TPP, which operates the SystmOne platform, presents a different kind of obstacle - a bulk extract mechanism exists and is sufficient for MendelScan's purposes, but TPP's institutional culture of limited external engagement makes this a clunky and unreliable route to scale. In both cases, the deeper problem is the same, accessing primary care EHR data at scale currently requires engagement with commercial platform providers whose pricing models, contractual terms, and governance appetites were not designed with population-level clinical AI deployment in mind. The logical resolution, and the one this submission advocates, is to bypass practice-level and platform-level engagement entirely, deploying MendelScan through larger regional or national data aggregators operating under a nationally commissioned framework, rather than negotiating access provider by provider and practice by practice.

SDEs: The project's experience within NHS SDEs illustrates this with precision. A dedicated task team worked for ten months to complete the technical mapping and documentation required for SDE integration, successfully completing the preparation for Phase 1 (generate aggregate and individual reports on undiagnosed rare disease patients using de-identified data within the SDE). Phase 2, the ability to convert those insights into clinical intervention by contacting GP practices about identified patients, proved a different matter entirely.

The fundamental structural obstacle was that SDEs are designed to authorise data access for research, not to authorise interventional clinical studies. Phase 2 would move the project into direct care; it fell outside the SDE's governance remit. Moving from de-identified research data back to identifiable patient data for GP intervention required a level of information governance clearance, specifically CAG251 designation, that the SDE structure was not positioned to facilitate for this use case.

Beyond governance, GAIN encountered a clinical buy-in problem that no amount of technical sophistication could resolve. Engagement with Local Medical Committees (LMCs) revealed resistance from GP leaders, centred on three concerns: ugh

  1. Insufficiently value for individual busy practices
  2. Uncertainty about the resource implications of proposed interventions for an already overstretched primary care workforce
  3. Concern that re-identifying patients via data originally contributed to an SDE under a "de-identified for research" promise would undermine the trust relationship between practices and the SDE. This last concern reflects the patient-NHS data relationship problem in microcosm, data shared for one purpose cannot straightforwardly be repurposed for another, even where the clinical case for doing so is strong and patient and public support is overwhelming.

GAIN showed that infrastructure gaps that must be filled before the next equivalent programme can succeed.

Conclusion

The Government has articulated a clear ambition: an NHS at the front of the global genomics revolution, the most AI-enabled care system in the world, with precision medicine made a reality for patients.

The Topol Review (2019) warned that 'uneven NHS data quality, gaps in information governance and lack of expertise remain major barriers' to AI adoption. Seven years on, those barriers remain.

The Luxturna example should serve as a call to action. If the NHS cannot reliably find patients for a NICE-approved, life-changing therapy, it is not ready for the next generation of precision treatments. The EOM case study demonstrates that this challenge will only intensify as individualised ASO therapies, gene therapies, and other precision treatments come to market. Getting patient identification right, systematically, equitably, at scale, is the foundational challenge on which the UK's personalised medicine ambitions stand or fall.

The primary barrier to the UK's personalised medicine ambition is not scientific excellence or public resistance, it is the absence of a systematic action layer to find and reach patients who can benefit, and the commissioning will to sustain it.

Supporting Evidence

Supporting evidence available on request:

        Boardman-Pretty F et al. (2026). 'Evaluating algorithmic approaches to rare disease case-finding: a retrospective validation study using electronic health records.' Orphanet Journal of Rare Diseases, 21:120. doi:10.1186/s13023-026-04240-6

        Marchini E et al. (2026, preprint). 'Public perception of the use of clinical decision support tools within the NHS for rare disease case finding.' Research Square. doi:10.21203/rs.3.rs-6304847/v1

        Mendelian NHS AI in Health and Care Award Final Report (December 2024), NIHR ref AI_AWARD02579

        NHS Genomic AI Network (GAIN) Report (genomicainetwork.nhs.uk).

 

Further information: www.mendelian.co/publications