Written evidence by Educate Ventures Research Limited (AIE0181)

 

Education committee - The use of artificial intelligence and EdTech in education

 

Submitted by: Professor Rose Luckin and Educate Ventures Research (EVR)

April 2026

 

Summary

The most important question in this call for evidence is whether Government is ensuring that AI in education increases skills rather than producing deskilling and over-reliance. The evidence shows this risk is real, measurable, and design-dependent: the same underlying AI model, implemented with different constraints, can either support learning or actively harm it. The label “AI tutor” currently covers products with fundamentally different capabilities and very different implications for learning outcomes. Procurement standards that do not distinguish between them will not protect learners.

The Government’s ‘Every Child Achieving and Thriving’ white paper commits to setting safety and efficacy standards for AI tools in schools, but the infrastructure to operationalise that commitment does not yet exist. There is no agreed definition of what counts as a learning-efficacy study, no requirement for vendors to publish subgroup data by attainment, EAL status, or SEND, and no mechanism for updating evidence thresholds as tools change. The NEU survey (February 2026, n=9,408) found that 49% of schools have no AI policy at all, and that only 14% of teachers support the Government’s planned AI tutoring programme.

A further structural problem is that traditional research timelines are incompatible with the pace at which AI tools change. Evidence collected in 2021–22 is published in 2024–25, by which point the products studied may no longer exist in that form. Dynamic evidence-gathering methods that combine expert consultation with ongoing practitioner observation offer one practical response.

The submission also addresses: the absence of subgroup evidence for AI tutoring tools across SEND, EAL, and lower prior attainment groups; the gap between independent and state schools in AI adoption and leadership capacity; the failure of the Curriculum and Assessment Review to address the metacognitive capacities most at risk from deskilling; the distinction between AI that supports teacher judgement in marking and AI that automates it; and the shortage of structured, subject-specific CPD, including for early career teachers in initial teacher training.

Nine recommendations are set out. The most time-sensitive concern the evaluation design of the AI tutoring trial: it should specify what category of product is being tested and what retention evidence it will require before the programme scales.

 

Introduction

Professor Rose Luckin is a Professor of Learner Centred Design at University College London and founder of Educate Ventures Research (EVR), an independent research consultancy specialising in the application of artificial intelligence in education. Her book Machine Learning and Human Intelligence (MIT Press) sets out a theoretical framework for understanding what human intelligence contributes that AI systems currently cannot replicate.

EVR submits this evidence because the inquiry falls directly within our research and expertise. The submission is focused on areas where we have direct evidence to offer. It does not address every term of reference: questions about children’s digital rights law, data protection regulation, and home education are better addressed by organisations with specific expertise in those areas.

 

1. The deskilling question

Not all AI tools carry the same risk

The Committee asks whether Government is ensuring that AI in education increases skills and knowledge rather than producing deskilling and over-reliance on technological support. There is a specific, empirically grounded answer to that question.

For example, a study by Bastani et al. (2025), published in the Proceedings of the National Academy of Sciences (volume 122, issue 26, DOI 10.1073/pnas.2422633122), found that students who used an unconstrained AI tutor for mathematics saw their test scores fall by approximately 17% after the tool was removed, compared to a group that had never had AI access. The study is a randomised field experiment conducted with approximately 1,000 students at a high school in Turkey during the 2023-24 academic year. It is peer-reviewed and directly responsive to the Committee's concern. A caveat applies: it is a single-site study in a specific national and curricular context, using GPT-4 as it existed in 2023, and its generalisability to English schools has not been tested. It is a good example of published field evidence on this question to date, not the final word on it. The same underlying model, implemented with pedagogical guardrails that gave hints rather than direct answers, produced a 127% improvement in session performance with no comparable deskilling effect. The two conditions in the Bastani study were not testing different AI systems. They were testing different implementation choices applied to the same model. That distinction matters for policy.

We suggest that it is useful to distinguish three categories of product.

        Category A tools support teachers and practitioners in tutoring tasks: planning, resourcing, and formative assessment. Their primary purpose is to reduce teacher workload, and their efficacy evidence concerns output quality rather than student learning gains.

        Category B systems are marketed as AI tutors and present some features associated with effective tutoring, such as adaptive feedback. However, their published evidence does not include tests of learning retention after tool withdrawal.

        Category C systems are those for which there is published evidence of learning gains that persist after tool removal. Very few commercial products currently meet this standard.

 

The label ‘AI tutor’ currently covers products from all three categories, and this is generating procurement confusion in schools. A Category A tool that produces high satisfaction ratings is not evidence of a Category C effect. Procurement standards that do not distinguish between these categories will not protect against the deskilling risk.

Category matters, but implementation is equally important. Our analysis of AI tutoring products finds that even a well-designed Category B or C system can fail to produce learning gains if it is deployed without teacher oversight, structured usage time, and monitoring of whether individual learners are actually engaging at the required level. Research on AI tutoring systems shows consistently that only a small proportion of students, sometimes as few as 5%, achieve sufficient sustained use to benefit. This group tends to be students who would succeed through other means in any case. A school that acquires a suitable tool but does not build in goal-setting structures, teacher dashboards, and scheduled practice time is unlikely to see the equity gains the tool was procured to deliver. The Bastani study illustrates the same point from the design side: the implementation constraints on GPT Tutor, which required the AI to give hints rather than answers, were what produced the learning effect. Procurement and implementation cannot be treated as separate decisions.

The scale of the deskilling concern in practice is visible in recent survey evidence. The National Education Union survey conducted in February 2026, drawing on responses from 9,408 teacher members in English state schools, found that 66% of secondary teachers agreed that pupils’ critical thinking had declined as a result of AI use. The figure among primary teachers was 28%. The concern is strongest among teachers in their twenties, who reported greater familiarity with and usage of AI than older colleagues: 57% of that group agreed, compared to 39% of those aged 50 and over. This is observed classroom experience, not generational scepticism about technology.

What is at risk: the theoretical account

Professor Luckin's model of interwoven intelligence, developed in Machine Learning and Human Intelligence, provides a principled account of what is at stake. The model holds that AI systems currently excel at cognitive tasks: retrieving, processing, and generating information. Human intelligence is distinguished by metacognition (knowing what one knows and does not know), meta-emotion (understanding one's own emotional responses), metacontextual awareness (understanding the social and situational context of a problem), and perceived self-efficacy (the belief that one can succeed through effort). Current AI systems cannot replicate these capacities.

When students outsource cognitive work to AI tools without developing these meta-level capacities, the cognitive development that education is designed to produce does not occur. Dependency on the tool is the lesser concern. Education systems rarely assess these meta-level capacities explicitly, which means the deskilling risk can accumulate without being detected by standard performance metrics.

 

2. The Government's framework for quality assurance

A commitment without the infrastructure to support it

The Committee asks whether Government has an appropriately well-defined framework for steering and regulating AI in education, and whether there is a consistent approach to quality assurance.

The Government's 'Every Child Achieving and Thriving' white paper commits to setting safety standards and efficacy assessments for AI tools used in schools, and to updating product safety standards to account for risks around mental health and cognitive development. The gap is in the infrastructure to operationalise that commitment. There is currently: no agreed definition of what constitutes a learning-efficacy study for AI in education; no requirement for vendors to publish subgroup analyses by attainment level, English as an additional language (EAL) status, or SEND; and no mechanism for revising evidence thresholds as AI tools change.

The absence of institutional guardrails is visible in schools themselves. The NEU survey published on 2 April 2026 found that 49% of teachers reported their school had no policy whatsoever on AI, either for staff or students, and that 66% had no policy specific to student use. These figures had not changed materially from the previous year, suggesting that the rate of policy development is not keeping pace with the rate of AI adoption in schools. Three quarters of teacher respondents (76%) reported using AI tools in their day-to-day work, up from 53% the year before.

The Government’s proposed AI tutoring programme for disadvantaged pupils addresses a genuine and urgent problem. Access to high-quality one-to-one tutoring has long been unequally distributed, and the pandemic widened that gap further. If AI tutoring can provide something approaching effective tutoring at scale for pupils who would otherwise receive none, the potential benefit is substantial. There is a wide body of research on Intelligent Tutoring Systems that is relevant and offers some grounds for cautious optimism. Design matters: well-designed AI tutoring tools can support learning rather than displace it. Whether the programme as designed will realise that potential depends on what it tests and how. The evaluation will need to specify what category of product is being used, and what evidence of learning retention after tool withdrawal it will require. Teacher support for the programme is low, and opposition is concentrated among those working in special schools and pupil referral units. That concentration matters: if the programme is to benefit the most disadvantaged learners, it is precisely in those settings that the evaluation evidence needs to be strongest, and at present it is thinnest. Getting the design right is in the programme’s own interest. It is also vital to engage these teachers in the project, because we know that the quality and nature of the implementation of an AI tutor is fundamental to success.

The ‘Every Child Achieving and Thriving’ white paper also commits to updating product safety standards to account for mental health risks. That commitment currently lacks a mechanism to fulfil it. The concern is not hypothetical: the interwoven intelligence model identifies meta-emotion (awareness of one’s own emotional responses and how they affect learning) and perceived self-efficacy (belief in one’s capacity to succeed through effort) as elements of human intelligence that AI systems struggle to replicate and that education systems rarely assess explicitly. Both are implicated in adolescent wellbeing. If students develop habitual reliance on AI tools for reassurance, feedback, or social interaction during a period when they would otherwise be developing these capacities through human relationships and effortful learning, the developmental effect may be significant and may not appear in academic performance data until considerably later. The published evidence on AI companion tools and adolescent wellbeing is currently too thin to quantify this risk, and that thinness is itself informative: the Committee should be cautious about allowing a product category to scale ahead of any evidence framework for assessing its developmental effects.

 

3. The evidence methodology problem

Research timelines are structurally incompatible with the pace of AI change

The Committee asks whether the Government's framework is adequate. The structural problem is one that better-funded research alone cannot solve: the timelines of traditional educational research are incompatible with the rate at which AI tools change.

Evidence collected in 2021 or 2022 is typically published in 2024 or 2025. By that point, the products studied may have been substantially updated, withdrawn, or superseded. A randomised controlled trial that takes three years from design to publication cannot provide the evidence base that school procurement decisions require in a market where AI tools change on monthly timescales. No government has solved this problem: Singapore has greater visibility over AI use in schools partly because its network infrastructure gives it direct data access, not because it has developed a methodology that resolves the timeline issue.

One practical direction is a Delphi-plus-citizen-science methodology: structured expert consultation combined with systematic, ongoing collection of practitioner observations, which allows evidence to be updated as tools evolve rather than only at the conclusion of long research cycles. This approach is being explored in advisory work with the DfE and warrants wider consideration as a complement to conventional research. It is not a proven solution, but it addresses the structural problem in a way that simply commissioning more conventional research does not.

This point is a diagnosis of why the current evidence infrastructure is inadequate to the task Government has set for itself. The submission does not argue against rigorous research; it argues that commissioning more of the same kind will not solve the problem. Naming that gap is more useful to the Committee than restating the commitment.

 

4. The disadvantage evidence gap

The absence of subgroup data is itself a finding

The Committee asks whether socio-economic and demographic factors affect access to and benefit from AI-enabled education. The honest answer is that the evidence to answer that question does not yet exist in adequate form.

The absence of subgroup data in AI tutoring efficacy studies means it is not currently possible to say whether AI tutoring tools reduce or widen attainment gaps for lower prior attainment students, EAL learners, or children with SEND. That absence should inform the Committee’s conclusions: procurement decisions made without it are being made without evidence about the groups who most need effective support. This is a different question from whether AI tools can support SEND learners in other ways. Speech-to-text tools, communication aids, and reading assistance tools have a distinct evidence base and serve different functions from AI tutoring systems. The concern raised here is specific to AI tools marketed as tutoring or learning interventions, where the efficacy claims rest on evidence that does not include the learners most likely to be targeted by Government equity programmes.

The NEU survey data reinforces this concern. Opposition to the Government’s AI tutoring trial was strongest among teachers in special schools and pupil referral units, where 49% of all respondents overall were opposed. This is a concrete signal that practitioner concern is concentrated precisely where the evidence is thinnest. The profession is being asked to adopt tools for the most vulnerable learners on the basis of evidence that does not include those learners.

AI companion products for young children represent a category where the absence of an evidence framework is particularly acute. Metacognitive and social-emotional development in early childhood depends on sustained, contingent interaction with human caregivers and peers. These are precisely the capacities that the interwoven intelligence model identifies as distinguishing human intelligence from AI, and they are at their most formative in the earliest years. There is no substantial published evidence base on the developmental effects of sustained interaction with AI companions in early childhood, and no regulatory framework that would require such evidence before a product reaches the market. The Committee may wish to consider whether this product category warrants specific regulatory attention, independently of the broader AI in education framework.

On parental digital literacy, the Committee asks whether disparities among adults perpetuate inequalities for children. The evidence on AI use outside school is limited, and we note that gap without overstating what is known. Children educated at home represent a further group the call raises and on which EVR has no specific evidence base; the Committee will be better served by organisations with direct expertise in home education on that question.

A further dimension of the access and inequality question is the gap between state-funded and independent schools. Independent schools have, in many cases, moved faster to adopt AI tools, to invest in staff training, and to develop institutional policies. Within the state sector, variation is also substantial. EVR’s advisory work with schools indicates wide differences in the level of confidence and expertise that senior leaders bring to AI adoption decisions, between schools that are actively and thoughtfully engaging with the question and those that are not yet engaging at all. This variation is not randomly distributed: it tends to correlate with existing resource and capacity differences, meaning that AI in education risks becoming another dimension along which advantage compounds rather than narrows.

The Committee asks specifically about digital infrastructure. EVR’s advisory work does not provide a systematic audit of infrastructure across schools and settings, but the variation described above, between sectors and between schools within the state sector, applies here too. Schools without adequate connectivity, device provision, or technical support cannot meaningfully adopt AI tools regardless of procurement decisions made elsewhere. The inequalities of access that already characterise digital provision in schools will shape AI adoption in the same way, and the Committee may wish to request specific evidence on this from organisations better placed to quantify it.

School and trust leaders face a particular difficulty. They are being asked to make consequential decisions about AI procurement, data governance, and cyber security without a central framework to guide them. The DfE has published CPD resources for leaders, but the expectation remains that each school or multi-academy trust navigates these questions independently. In the absence of agreed national standards, leaders are making judgements about data handling, third-party contracts, and the relative merits of competing products on the basis of whatever expertise they happen to have access to locally. The Committee may wish to consider whether this is an adequate basis for decisions that affect large numbers of children.

 

5. Curriculum and assessment

The Curriculum and Assessment Review did not address the central question

The Committee asks whether the Curriculum and Assessment Review adequately anticipated AI adoption. The interwoven intelligence model provides the basis for a specific answer: it did not.

If metacognition, meta-emotion, and epistemic cognition are the capacities that distinguish human intelligence from what current AI systems can deliver, then the curriculum question is not primarily about adding an AI literacy module or updating vocational content to include AI tools. It is about whether education systems are explicitly developing and assessing the meta-level capacities that matter most, and whether assessments are designed to detect whether those capacities have been developed independently of AI support. The Committee specifically asks about oracy, problem solving, and creativity. Each depends on capacities the interwoven intelligence model places in the human column. Oracy requires metacontextual awareness: reading a room, responding to an interlocutor in real time, adjusting register. Problem solving produces learning through effortful engagement that cannot be offloaded without undermining its purpose. Creativity, as distinct from content generation, draws on meta-emotional and metacognitive capacities that AI systems cannot replicate. A curriculum that does not explicitly protect and develop these capacities is not AI-ready: it is vulnerable.

The assessment validity problem the Committee raises about coursework completed outside the classroom is a symptom of this deeper issue. Research published by Wonkhe in March 2026 (Dickinson and Marshall, ‘Trained to stop learning?’), drawing on a survey of 1,055 students across 52 higher education providers and focus groups conducted in February and March 2026, found that 38% of students admitted submitting work they could not fully explain without going back to their sources, and that 47% worried their grades did not reflect what they actually knew. The strongest predictor of this pattern was assessment design, specifically whether a visible accountability moment existed at which students would need to demonstrate understanding in person. AI use as such was not the primary driver. When such a moment was present, students used AI differently, prompting it more carefully and testing their own reasoning rather than generating output. This is a higher education finding, and its direct applicability to schools requires care, but the underlying dynamic mirrors what the Committee is asking about: whether assessment systems can currently detect the difference between a student who has learned and one who has produced something plausible.

If a student completes assessed work with AI assistance, the mark may not reflect the cognitive capacities the assessment was designed to measure. Those capacities may not have been developed at all. That is a problem of educational design, not only of academic integrity, and it is not one that AI detection tools can resolve.

 

6. AI use in teacher assessment and marking

The Committee asks about AI use in teacher-facing assessment and marking processes.

Explainable AI models can increase the consistency and transparency of expert human assessment decisions without replacing professional judgement. The distinction matters: AI that surfaces patterns for a teacher to interpret and act on is functionally different from AI that generates a mark or written feedback autonomously. The former can support professional development and reduce unconscious inconsistency in formative assessment; the latter raises questions about accountability and about whether the feedback a student receives reflects a professional judgement at all.

Government guidance on AI use in teacher assessment should distinguish between these two functions. Tools that assist teachers in identifying patterns across student work, flagging potential mismatches between predicted and actual performance, or supporting consistency in formative feedback represent a different category from tools that generate feedback or marks autonomously. Procurement standards and professional guidance should reflect that distinction.

 

7. Teacher confidence and professional development

The Committee asks about teacher confidence in using AI and whether practitioners have adequate support and training. Recent survey data shows the scale of that gap directly.

A TeacherTapp survey conducted for the Royal Society on 20 March 2026, drawing on responses from approximately 9,250 teachers in England weighted to reflect national demographics, found that only 14% of teachers reported feeling confident across practical, technical, and human aspects of AI literacy. A further 33% described themselves as mostly confident in the practical use of AI tools. However, 34% reported limited confidence in most aspects of AI literacy, and 11% were unsure. The confidence gap was largest among older teachers: 47% of those aged 50 and over reported limited confidence, compared to 19% of those in their twenties.

On AI literacy in teaching, 37% of respondents said it was not currently addressed in their school's teaching at all, and only 2% reported that it was covered across most subjects. Regarding the support schools provide to teachers, 32% reported receiving only informal guidance such as sharing practice, 23% reported no support at all, and only 18% had access to structured professional development such as CPD sessions or training courses.

When asked what free resource they would most want to improve their AI literacy, the most common choices were ready-to-use lesson plans and classroom activities (20%) and curated lists of approved AI tools (19%). Teachers want practical, subject-specific resources they can use immediately, not generic conceptual training. EVR’s advisory work with schools points in the same direction: effective CPD requires sustained engagement within subject-specific contexts rather than one-off sessions. The appropriate use of AI tools in creative writing teaching differs materially from its use in mathematics, and CPD that treats AI literacy as a generic skill will not build the professional judgement that subject teaching requires.

The Royal Society's rapid review of AI literacy frameworks, published alongside the TeacherTapp survey results, found that no agreed national approach currently exists for ensuring all young people are equipped to use AI effectively and appropriately, and that most existing AI education efforts focus heavily on technical skills with less emphasis on broader societal and ethical dimensions. The review identified a risk of inequity: without a shared baseline, AI literacy is becoming dependent on local initiatives or individual teacher expertise, which tends to widen rather than narrow existing educational inequalities.

 

Recommendations

The Committee may wish to consider the following.

        Procurement standards for AI tools in schools should distinguish between Category A tools (teacher-support functions) and Category B/C systems (claims to deliver learning gains). Standards appropriate to the former are not adequate for the latter.

        Any Government-sponsored AI tutoring trial should specify, in its evaluation design, what evidence of learning retention after tool withdrawal it will require. Satisfaction ratings and engagement metrics are not substitutes for retention evidence.

        The Government's commitment in 'Every Child Achieving and Thriving' to set safety and efficacy standards requires, as a minimum, an agreed definition of what counts as a learning-efficacy study for AI in education.

        Vendors of AI tutoring systems that make learning-efficacy claims should be required to publish subgroup analyses by prior attainment, EAL status, and SEND. It is not currently possible to say whether AI tutors reduce or widen attainment gaps for these groups because that data has not been collected or published.

        The DfE should invest in developing dynamic evidence-gathering methods alongside conventional research to reduce the lag between tool deployment and evidence availability. Approaches that combine structured expert consultation with systematic, ongoing collection of practitioner observations offer one practical direction, enabling evidence to be updated as tools evolve rather than only at the end of lengthy research cycles.

        A review of curriculum and assessment should consider explicitly whether current programmes of study and assessment instruments develop and detect the metacognitive and meta-emotional capacities that are most at risk from deskilling through AI use. The Curriculum and Assessment Review did not address this.

        Government should publish a national AI literacy framework for schools that addresses not only technical understanding but ethical and societal dimensions, and that provides the agreed baseline currently missing. Without it, AI literacy provision will remain dependent on individual teacher initiative and will widen existing inequalities.

        The DfE should publish minimum data governance standards for AI tools used in schools, covering data retention, third-party processing agreements, and the use of student data for model training. These standards should be distinct from, and complementary to, the learning-efficacy standards proposed above. School and trust leaders are currently making decisions about third-party AI contracts without a national framework to guide them. The consequences of those decisions, particularly where student data is processed outside the school’s direct control, may not become visible until after contracts are signed and tools are embedded in practice.

        Structured, subject-specific CPD on AI in education should be made available to all teachers, not only those in schools that can resource it independently. The TeacherTapp data indicates that 34% of teachers currently have limited confidence in most aspects of AI literacy, and that 23% receive no school support at all. Initial teacher training requires specific attention: early career teachers who become reliant on AI tools before developing their own pedagogical judgement face a particular form of professional deskilling, and ITT providers should be supported to develop approaches in which AI is used critically and selectively. The lived experience of planning, assessing, and responding to learners is how teachers build the professional knowledge that AI use should complement rather than replace.

 

 

Professor Rose Luckin | Educate Ventures Research | April 2026

 

May 2026

10