Written evidence submitted by Thorney Isle Research (UAIG0016)
Evidence to the Public Accounts Committee
- Thorney Isle Research is a research and advisory body specialising in the interactions between technology, public policy, governance, and public administration. Our recent focus has been on the regulation of algorithms and artificial intelligence (“AI”), and the use of such methods in government, public administration, and the public sector generally. This contribution updates evidence to previous parliaments, taking account of new Bills and policy statements such as the AI Opportunities Action Plan (AIOAP) of 13 January 2025. It will focus on the risks and opportunities of AI adoption in government.
- We begin by observing that we see little evidence to support claims of “great potential” for AI in Government and public services (the AIOAP offers none), but considerable evidence of risks and problems. In regard to the Government’s stated expectations for its use of generative AI in particular and AI in general, it is remarkable that any democratic government would suggest that it puts at the heart of public services a technology that is built on the exploitation of low-paid labour[1], consumes unsustainably massive amounts of power[2] and water[3], is hugely expensive to develop and operate[4], and in use is likely to be inaccurate and in conflict with the rule of law and human rights[5].
- Currently, the use of “AI” and other algorithmic methods by the UK government does not meet the standards expected in a democracy working under the rule of law. We believe that there is a need for a stronger governance framework that reflects principles of natural justice, fairness and proper administrative procedure[6], and our recommendations suggest key elements of such a framework.
Terminology
- We use the term “algorithm” or “algorithmic decision making” in preference to “AI”, as “AI” is an imprecise term — a myriad of attempts to define and scope “AI” have struggled (and the AIOAP does not try). Also, from 2023, it became almost synonymous with generative AI such as Large Language Models. In its increased popular use, it has become to sound like a singular thing, and even a singular thing that has agency of its own: both these impressions are false, and unhelpful for critical analysis. Discussion of opportunities and risks must not be bound by ill-defined scoping constraints or false perceptions.
- Key government actors already use terminology based on “algorithm”. The Central Digital and Data Office’s (CDDO’s) Ethics, Transparency and Accountability Framework for Automated Decision-Making says, “This 7-point framework will help government departments with the safe, sustainable and ethical use of automated or algorithmic decision-making systems”. The Responsible Technology Adoption Unit (RTA) in DSIT and CDDO develop and publish the Algorithmic Transparency Recording Standard (ATRS) to “provide the framework for accessible, open and proactive information sharing about the use of algorithmic tools across the public sector”.
- Further, in government, high-impact algorithms may be used that would not necessarily fit a definition of “AI”. Even simple methods e.g. for assessing welfare eligibility or fraud risk can have huge effects on people's lives. There have been many cases reported worldwide where innocent people have been punished by false results from a computer program — something brought to public attention by the Post Office Horizon scandal. The Committee will be aware of the debate on 13 December 2024 in the House of Lords on the Public Authority Algorithmic and Automated Decision-Making Systems Bill[7] that covered this and the lack of transparency about the use of such methods.
- In public administration there is a world of difference between using established analytical data science methods; prediction and classification algorithms; and large language models or chatbots, but these specific processes, and others, get bundled under the label “AI”.
- Using data analytics to find important patterns in data to inform policy analysis is a valuable thing to do, as the pandemic illustrated, and is a matter for professional statisticians and others suitably qualified in the methods. But making predictions about individual people, like whether they might commit fraud or not be eligible for a visa, is something entirely different. This, and the use of chatbots, is covered in more detail below.
Recommendation 1: References to “AI” or “artificial intelligence” in relation to applications within government should be dropped and more precise terms used instead, or at least the general term “algorithms” as has already been adopted, to avoid confusion, gaming, and unnecessary constraints on governance and assurance.
Transparency
- The current and previous Governments mandated the use of the ATRS by UK government departments, with an intent to extend that requirement to the broader public sector. However, the repository for ATRS records on gov.uk[8] listed on 8 January 2025 only 23 records of which 14 were added in December 2024. Many of the new additions were applications of text processing, presumably inspired by the emergence of Large Language Models. The well-known and controversial applications of predictive systems by the DWP, the Home Office, etc[9][10], still do not appear[11].
- Transparency about the development and operation of algorithmic methods by central government is essential to establish legitimacy, accountability, accuracy, and absence of discrimination but is currently far from adequate.
Recommendation 2: The Algorithmic Transparency Recording Standard (ATRS), now mandated, should be used to immediately catalogue all current instances across government where an algorithmic method is used to inform decisions relating to legal persons.
The Rule of Law
- Government functions must obey the law. Relevant ones here are the Data Protection Act 2018 (which has specific clauses on profiling and automated decision making), the Human Rights Act 1998, the Equality Act 2010 including its Public Sector Equality Duty (s149), and the Public Records Act 1958. Principally, however, any public sector function must obey the fundamental laws concerning the functions being exercised — administrative law that determines what a public sector body is and the way it works.
- The rule of law also requires that government functions follow the principles of accountability, equity, transparency, explainability, predictability, and consistency, most especially in carrying out public administrative functions such as executing the administration of statutory entitlements (like benefits) and obligations (like taxation and licensing). This is especially important when making any calculation, estimate, classification, profile or prediction of any characteristics, circumstances, activities, opinions or behaviour of legal persons, individually or collectively.
- Further, government functions must be able to satisfy legal and democratic demands for reason-giving and principled justification for decisions affecting individuals[12]. This goes beyond explaining a decision to justifying why a particular outcome is considered normatively acceptable by reference to law or some underlying theory of moral or social acceptability.
- Numerous reports have suggested that there are governmental uses of algorithmic tools that do not satisfy these fundamental requirements, but due to the lack of transparency there is little evidence to prove or rebut this.
Recommendation 3: All systems catalogued as above must: fully describe how they meet all applicable legal and administrative requirements; demonstrate that they are accurate, consistent and unbiased; and show that deployers can provide meaningful explanations for their outputs and that outcomes can in turn be normatively justified.
Problems with stochastic methods
- Most current discussions about the potential for “AI” in government functions implicitly relate to “generative AI” (such as chatbots), predictive systems, or facial recognition (also diagnostic systems in health, but we will not discuss that large sector here). Except for some simpler predictive tools, such systems are typically constructed using machine learning (ML). ML methods create a model purporting to represent some aspect of reality based on statistical correlations between measured features of cases such as sources of text, faces, or historic records on people. It is well-known that such systems can be biased. However, there are more fundamental problems, particularly that they are not accurate enough for use in government functions under the rule of law.
- When such systems produce an output, e.g. new text, a match of a face, or a prediction about someone, it is based on the model’s calculation of relationships between the input features: These are probabilistic (stochastic), being based on statistical estimates of correlation (not causality). Depending on the ML method used, it may not be possible to explain how the output was produced from the inputs (the system is a “black box”). This is usually the case with neural networks and transformer-based methods as in the Large Language Models underlying chatbots.
Chatbots
- Such chatbots produce an estimate of what is the most likely text to follow the preceding words (such as a prompt). Outputs vary for each prompt, so they are inconsistent and unpredictable, and there is no way of explaining what led to a particular output or — without additional work — knowing whether it is accurate. These problems cannot be fixed[13]. However, these tools are currently proposed to be useful in accessing online government information, simplifying interactions with the public in administrative procedures, or summarising policy consultation responses.
- The Government Digital Service experimented with a chatbot interface to gov.uk, finding “Overall, answers did not reach the highest level of accuracy demanded for a site like GOV.UK, where factual accuracy is crucial. We also observed a few cases of hallucination — where the system generated responses containing incorrect information presented as fact”[14]. These are inherent problems with such tools — making them unsuitable for this purpose and for the same reasons (plus the lack of explainability and consistency) for use in statutory administrative procedures.
- The claim that such tools could summarise policy consultations is correct, but their use for this purpose is wrong. The benefit of a consultation to policy makers is to find the things they had not thought of, identified by people they hadn’t spoken to — the needles in the haystack. A chatbot summary will not necessarily pick them out. It will not enable the nuanced positions on the policy by stakeholder groups to be ascertained. Further, even without the potential “hallucinations”, an automated summarisation does not fulfil the democratic function of a consultation, to allow all voices to be heard and be seen to be heard. It falls far short of this, and in the process devalues and insults the considerable effort respondents make to be helpful in policy development[15].
Predictive systems
- A study co-authored by this respondent identified three dozen sources of risk when public bodies use algorithmic methods to make predictions about people or their circumstances, or to influence their behaviour[16]. Worldwide, instances of bias and inaccuracy are evident. Recent guidance from DfE and RTA on the use of data analytic tools in children’s social care[17] specifically warns against predictive uses, citing findings that they do not work.
- Prediction in relation to a person from a machine-learning algorithm essentially projects the statistical average result from a group of people onto an individual. This is fundamentally in conflict with the rule of law in a democracy where an individual’s case should be judged on their individual circumstances, not relative to statistically-derived parameters from a more-or-less representative group. The Committee will be aware of the challenges to provisions in the Data (Use and Access) Bill relating to automated decision-making[18]. Indeed, a strong case can be made for the practice to be banned[19].
- Generally, any form of binary classifier (i.e. where an algorithm flags a case as either positive or negative e.g. yes or no, unsafe or safe, fraud or not fraud, etc) is very difficult to properly calibrate as adequate data is needed to represent true positives, true negatives, false positives and false negatives. In many public sector instances reported, this may well not have existed as detailed data may only be collected on cases that are investigated rather than on a group representative of the population. If the data is not representative, the classifier may be inaccurate to the extent of being useless.
- Further, interpreting a positive output from such a classifier is not straightforward. Usually, the probability that the classifier will give a positive output if the case is truly positive may be known from development and testing. However, that is not the same as the probability of a new case being truly positive when the classifier flags it as positive. To work that out you need to know the proportion of positive cases in the population and apply Bayes’ theorem[20]. Again, mostly that population proportion is unknown making it impossible to interpret the result. Further, the mathematics here shows that if the proportion of positive cases in the population is in fact small, the likelihood of a false positive can become significant (e.g. two-thirds of housing benefit claims marked as high risk actually being legitimate[21]). Consequently, these kinds of predictors are extremely dangerous to use for public sector decisions.
Facial recognition
- There are many varieties of potential applications of facial and other biometric recognition in administration, e.g. in immigration and policing, and we do not propose to rehearse the extensive debate here. However, they are again stochastic in nature, based on a non-explainable ML process, and have been shown to demonstrate bias.
Recommendation 4: Any use of generative AI must be reported as part of the cataloguing as above, shown to be compliant with the Government’s own guidance on its use, justified in relation to cost, security, benefit and legality (including regarding intellectual property), and evidenced as sufficiently accurate for the purpose intended.
Recommendation 5: Any practices within government where algorithmic methods are used to identify or make predictions about individuals, including biometric recognition, are immediately suspended until a full examination is made of their accuracy and compliance with the rule of law and principles of good public administration.
Conclusion
- Basing analysis, policy and action on an arbitrary scope for “AI” as a thing or loosely-defined set of things is not useful in this context. What matters is how government departments are using or intending to use any algorithmic methods in the execution of their functions, particularly those that directly affect the lives of people. This is currently not known, and the Government’s own guidance and tools need to be enforced in the cause of transparency and good governance.
- Much of the speculation about the potential application of “AI” in government relates to stochastic systems. We propose that any such system that produces statistically-derived outputs will fail to meet the criteria set by the rule of law and principles of public administration, and the standards of our democracy. Any decision or policy proposal based on one is unlikely to be able to satisfy a public or parliamentary demand to hear a principled justification for the decision or proposal.
- The current flurry of excitement and grand claims for how “AI” could transform government and the public sector is therefore unsupportable.
January 2025
[1] https://www.bloomsbury.com/us/feeding-the-machine-9781639734979/
[2] https://www.goldmansachs.com/insights/articles/AI-poised-to-drive-160-increase-in-power-demand
[3] https://www.nature.com/articles/d41586-024-00478-x
[4] https://time.com/6984292/cost-artificial-intelligence-compute-epoch-report/
[5] https://tilburglawreview.com/articles/10.5334/tilr.303
[6] Andrew Murray. Automated Public Decision Making and the Need for Regulation. LSE Public Policy Review. 2024; 3(3): 3. DOI: https://doi.org/10.31389/lseppr.110
[7] https://hansard.parliament.uk/lords/2024-12-13/debates/AA0C1C17-11FA-410E-A394-846703400F55/PublicAuthorityAlgorithmicAndAutomatedDecision-MakingSystemsBill(HL)
[8] https://www.gov.uk/algorithmic-transparency-records
[9] Andrew Murray. Automated Public Decision Making and the Need for Regulation. LSE Public Policy Review. 2024; 3(3): 3. DOI: https://doi.org/10.31389/lseppr.110
[10] https://amp.theguardian.com/society/2024/dec/06/revealed-bias-found-in-ai-system-used-to-detect-uk-benefits
[11] https://medium.com/@imogen-parker/a-window-into-the-black-box-new-entries-on-the-uk-governments-ai-register-20ef267b5377
[12] Karen Yeung and Adam Harkens. 2023. How do "technical" design-choices made when building algorithmic decision-making tools for criminal justice authorities create constitutional dangers? Part I. arXiv preprint (2023). https://arxiv.org/abs/2301.04715
[13] https://joanna-bryson.blogspot.com/2024/02/2024-brief-overview-on-llm-foundation.html
[14] https://insidegovuk.blog.gov.uk/2024/01/18/the-findings-of-our-first-generative-ai-experiment-gov-uk-chat/
[15] https://www.theguardian.com/commentisfree/2024/apr/25/ai-public-voice-ministers-large-language-model-chatgpt
[16] https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3716166
[17] https://www.gov.uk/guidance/develop-and-use-data-analytics-tools-in-childrens-social-care
[18] See e.g. https://www.openrightsgroup.org/app/uploads/2024/12/ADM-joint-letter.pdf
[19] Tafani, 2024 https://zenodo.org/records/10866778
[20] The procedure for inverting conditional probabilities to find the probability of a cause given its effect https://en.wikipedia.org/wiki/Bayes%27_theorem
[21] https://www.theguardian.com/society/article/2024/jun/23/dwp-algorithm-wrongly-flags-200000-people-possible-fraud-error