Written Evidence by Dr Alexander Lees (AIE0041)
Education Committee
The use of Artificial Intelligence and EdTech in Education
I am a biodiversity scientist working on global macroecological patterns and conservation challenges based at a very large UK university (student population ~44,000) where my teaching remit is broadly reflective of my research brief (e.g. units on wildlife conservation, landscape ecology, ornithology etc). My research group also uses artificial intelligence (AI), for example in leveraging convolutional neural networks in ecoacoustic research or large language models (LLMs) for generative AI image classification. Despite engaging with AI for research, I am very concerned about generative AI subverting learning in higher education - and deskilling future cohorts of scientists in writing, analysis and critical thinking. This concern is justified given emerging evidence for cognitive and neural impacts associated with sustained AI exposure which include diminished critical thinking and superficial learning (Lee et al. 2025, Yan et al. 2025) ((
Along with academics at other universities, we have had to iteratively revisit our assessment portfolios to try and make them robust to students ‘cheating’ with generative AI. Many assessments expressly forbid the use of generative AI, although some invite its use and ask students to critique its content (ostensibly to demonstrate model fallibility). My university (as is the sector norm) considers the unauthorised use of generative AI in assessments to be academic misconduct https://www.mmu.ac.uk/student-life/course/assessments/academic-integrity. The university does not however (as also seems to be the norm in the sector based on conversations with peers) permit academics to use any software to ‘detect’ the use of generative AI. This is justified based both on the slim risk of false positives resulting in students being erroneously penalised (especially students for whom English is not their first language), and the rise of ‘AI humanizers’ and paraphrasing tools which may be able to bypass detection software altogether, along with potential violation of student's intellectual property and privacy rights by uploading work to 3rd party software.
I have been teaching at the same university for nearly ten years and am familiar with the standard of coursework submissions from the pre and post large language model eras. We know that usage of generative AI is near ubiquitous among student cohorts, with for example 94% using generative AI to help with assessed work in one recent national survey. Worryingly, approximately 12% of respondents admitting to direct use of AI-generated text in assessments (HEPI, 2026). Based on my own marking samples, student work is trending towards homogenisation with substantial unsanctioned usage of generative AI text and figures being used in assessments. This is obvious from student submission ‘markers’ like generic syntax and a high frequency of erroneous content. For example, students citing academic papers which do not exist or do not support the points in their work. These ‘hallucinated references’ are the most obvious markers of content that has been produced by generative AI and it my understanding based on the academic literature that these models will always be imperfect (e.g. Banerjee et a. 2025). The tacit admission of academic misconduct by 12% of survey respondents (also likely to be a self-selecting cohort) will be a big underestimate of the true figure. I would estimate closer to 30-40% based on my own experiences of marking student work since Chat GPT’s release on November 30, 2022 (after which we can no longer assume that work was written by a human). This misconduct is prevalent in undergraduate and postgraduate cohorts, but is also an emerging problem in postgraduate research too.
The Impact on Teaching
LLMs are fundamentally reshaping teaching in higher education as they now limit the diversity of assessments we can offer to students (Newton 2025). Almost all assessments which permit students to work outside of exam, laboratory of remote field course conditions are at least partially vulnerable but some assessment types such as a short-format questions are now regarded as being especially problematic and we now avoid these sorts of assessments. This cohort of increasingly problematic assessments may disproportionately impact ‘authentic assessments’ which mimic real-world work (e.g. in my domain consultancy reports, writing R code, or press releases) which are increasingly considered the gold standard of student assessment. However, because generative AI is now a staple of real-world professional environments, such authentic tasks are exactly what generative AI is likely to be most proficient at completing given the ubiquity of training data.
The biggest impact on workloads is the increased amount of time spent marking assessed student work - at the moment we are faced with having to critically appraise long-form student work with rigour that was never previously required. Prior to the release of generative AI I never encountered students citing fabricated references – they would often inappropriate ‘soft’ ones from websites or grey literature to support statements – but any peer-reviewed academic literature they included would be real papers. Such ‘hallucinated’ references are now common in submissions, but not always readily identifiable and may even be accompanied by broken or erroneous digital object identifiers (DOIs) which would suggest authenticity. Importantly, hallucination rates vary between LLMs which risks furthering digital inequity depending on what tools, sanctioned or otherwise, students use. More common that the citation of hallucinated references is the citation of references which do not support the statements and are very tangentially related to the subject matter (e.g. via the odd keyword) and may not even be from the correct academic discipline. Again, this was rarely an issue historically – when students provided references, they normally supported statements they were making as they had taken the time to read them. As this information can no longer be taken at face value then to assess student submissions rigorously you have to interrogate all the references to check they are relevant, or even exist at all. Differentiating between original analytical or synthesis work and superficially well-written but potentially erroneous generic responses provided by a computer model has become a major burden on my time. Such due diligence is basically unsustainable for very large student cohorts – especially against a backdrop of an increase and diversification of demands on staff time coupled with performance tracking and reductions in academic staff cohorts at many institutions (Leeming 2024). This increased marking burden is also associated with another time burden involving investigating and sanctioning cases of academic misconduct which is necessary to maintain the quality and integrity of degree programmes (Nowak 2026) at both undergraduate and postgraduate level.
Failure to differentiate between pieces of work that are superficially similar but differ fundamentally in evidence of critical thinking (or arguably more importantly accurate content) is a huge risk for both the student experience with students who cheat the system potentially rewarded with good grades whilst others feel a profound sense of injustice (personal observations). It ought also to be a huge worry for grade inflation – if massive unsanctioned LLM usage leads to homogenisation of submissions and hence grades.
There is a pressing need to for universities to be able to field assessments which are robust to cheating with generative AI and the effective deskilling of student cohorts while simultaneously making sure that students are AI literate and understanding the value and limitations of AI workflows. This will likely require a greater emphasis in my disciplinary area on assessments under exam conditions, practical and field course settings and perhaps even time-intensive oral examinations (viva voce). Some of these changes will necessarily conflict with pedagogical goals of authentic assessments and inclusivity which represent a ‘wicked problem’ without easy solutions (Corbin et al. 2025). Mitigating impacts on disadvantaged cohorts of students can be hopefully achieved through specific provisioning (e.g. extra time and support) and through a diverse portfolio of assessments that may be resistant to large language models in different ways. However, given that student cohorts often pick courses based on the types of assessments offered, it is imperative that regulators and accrediting bodies ensure that all universities are required to ensure that assessment portfolios are sufficiently diverse and robust to cheating with generative AI to avoid the current status quo, or even worse a race-to-the-bottom which will devalue degrees if assessments fail to separate students who have and have not done original work.
To quote Leaton Gray et al. (2025): “If universities cannot differentiate between digital and human responses, they risk accusations of awarding degrees earned via AI algorithms rather than through student effort. Should this become universal, the credentialised higher education (HE) model as we know it faces the risk of collapse without appropriate anticipation of any useful alternative.”
References
Banerjee, S., Agarwal, A., Singla, S. 2025. LLMs Will Always Hallucinate, and We Need to Live with This. In: Arai, K. (eds) Intelligent Systems and Applications. IntelliSys 2025. Lecture Notes in Networks and Systems, vol 1554. Springer, Cham. https://doi.org/10.1007/978-3-031-99965-9_39
Corbin, T., Bearman, M., Boud, D. and Dawson, P. 2025. The wicked problem of AI and assessment. Assessment & Evaluation in Higher Education, https://doi.org/10.1080/02602938.2025.2553340
Leaton Gray, S., Edsall, D. and Parapadakis, D. 2025. AI-based digital cheating at university, and the case for new ethical pedagogies. Journal of Academic Ethics, 23: 2069-2086. https://doi.org/10.1007/s10805-025-09642-y
HEPI (2026). Student Generative AI Survey 2026. Higher Education Policy Institute. https://www.hepi.ac.uk/reports/student-generative-ai-survey-2026/ Accessed 18 March 2026
Lee, H.-P. H., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. 2025. The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. Proceedings of the 2025 CHI conference on human factors in computing systems, Association for Computing Machinery, New York, NY, USA (2025), pp. 1-22. https://doi.org/10.1145/3706598.3713778
Leeming 2024. UK university departments on the brink as higher-education funding crisis deepens. Nature 633: 969-971 https://doi.org/10.1038/d41586-024-03079-w.
Newton, P.M. 2025. How vulnerable are UK universities to cheating with new GenAI tools? A pragmatic risk assessment. Assessment & Evaluation in Higher Education, 50: 1332-1343 https://doi.org/10.1080/02602938.2025.2511794.
Nowak, R., 2026. The procedures of investigating academic misconduct: Perspectives from academic staff at a UK Russell Group University. Journal of Academic Ethics, 24: 55 https://doi.org/10.1007/s10805-026-09734-3.
Yan, L., Pammer‐Schindler, V., Mills, C., Nguyen, A. and Gašević, D. 2025. Beyond efficiency: Empirical insights on generative AI's impact on cognition, metacognition and epistemic agency in learning. British Journal of Educational Technology, 56: 1675-1685 https://doi.org/10.1111/bjet.70000.
May 2026
4