AIG0005

 

Written evidence submitted by Thorney Isle Research

 

Evidence to the Public Accounts Committee

  1. Thorney Isle Research is a research and advisory body specialising in the interactions between technology, public policy, governance, and public administration. Our recent focus has been on the regulation of algorithms and artificial intelligence (“AI”), and the use of such methods in government, public administration, and the public sector generally. This contribution will focus on the risks and opportunities of AI adoption in government. We begin by observing that we see little evidence to support claims of “great potential”, but considerable evidence of risks and problems.

Terminology

  1. We use the term “algorithm” or “algorithmic decision making” in preference to “AI”, as “AI” is imprecise term — a myriad of attempts to define and scope “AI” have struggled. Also, in 2023, it became almost synonymous with generative AI such as large language models. In its increased popular use, it has become to sound like a singular thing, and even a singular thing that has agency of its own: both these are false, and unhelpful for critical analysis.
  2. Key government actors follow a similar approach. The Central Digital and Data Office’s (CDDO’s) Ethics, Transparency and Accountability Framework for Automated Decision-Making says, “This 7-point framework will help government departments with the safe, sustainable and ethical use of automated or algorithmic decision-making systems”. The Responsible Technology Adoption Unit (RTA) in DSIT and CDDO develop and publish the Algorithmic Transparency Recording Standard (ATRS) to “provide the framework for accessible, open and proactive information sharing about the use of algorithmic tools across the public sector”. Other guidance addresses specific processes that might get called “AI”, such as chatbots, data analytics, or predictive analytics.
  3. Further, in government, high-impact algorithms may be used that would not necessarily fit a definition of “AI”. Even simple methods e.g. for assessing welfare eligibility or fraud risk can have huge effects on people's lives. Discussion of opportunities and risks must not be bound by ill-defined scoping constraints.

Recommendation 1: References to “AI” or “artificial intelligence” in relation to applications within government should be dropped and more precise terms used instead, or at least the general term “algorithms” as has already been adopted, to avoid confusion, gaming, and unnecessary constraints on governance and assurance.

Transparency

  1. On 29 March 2024, the Government’s Response to the Consultation on the AI White Paper announced that “the use of the ATRS will become a requirement for UK government departments, with an intent to extend to the broader public sector over time”. However, the repository for ATRS records on gov.uk[1] listed, on 29 April 2024, eight records of which only three were from central government departments (two from the Cabinet Office itself).
  2. The UN Special Rapporteur on extreme poverty and human rights found that “the existence, purpose and basic function of automated government systems [in DWP] remains a mystery in many cases fuelling misconceptions and anxiety about them”. Transparency about the development and operation of automated systems is essential to establish legitimacy, accountability, accuracy, and absence of discrimination.
  3. Transparency about the use of algorithmic methods by central government is almost totally absent.

Recommendation 2: The Algorithmic Transparency Recording Standard (ATRS), now required, should be mandated absolutely, and used to immediately catalogue all current instances across government where an algorithmic method is used to inform decisions relating to legal persons.

The Rule of Law

  1. Government functions must obey the law. Relevant ones here are the Data Protection Act 2018 (which has specific clauses on profiling and automated decision making), the Human Rights Act 1998, the Equality Act 2010 including its Public Sector Equality Duty (s149), and the Public Records Act 1958. Principally, however, any public sector function must obey the fundamental laws concerning the functions being exercised — administrative law that determines what a public sector body is and the way it works.
  2. The rule of law also requires that government functions follow the principles of accountability, equity, transparency, explainability, predictability, and consistency, most especially in carrying out public administrative functions such as executing the administration of statutory entitlements (like benefits) and obligations (like taxation and licensing). This is especially important when making any calculation, estimate, classification, profile or prediction of any characteristics, circumstances, activities, opinions or behaviour of legal persons, individually or collectively.
  3. Numerous reports have suggested that some governmental use of algorithmic tools do not satisfy these fundamental requirements, but due to the lack of transparency there is little evidence to prove or rebut this.

Recommendation 3: All systems catalogued as above must fully justify how they meet all applicable legal and administrative requirements, and demonstrate that they are accurate, consistent and unbiassed and can provide meaningful explanations for their outputs.

Problems with stochastic methods

  1. Most current discussions about the potential for “AI” in government functions implicitly relate to “generative AI” such as chatbots, predictive systems, or facial recognition (also diagnostic systems in health, but we will not discuss that large sector here). Except for some simpler predictive tools, such systems are typically constructed using machine learning (ML). ML methods create a model of reality based on statistical correlations between measured features of cases such as sources of text, faces, or historic records on people. It is well-known that such systems can be biassed. However, there are more fundamental problems, particularly that they are not accurate enough for use in government functions under the rule of law.
  2. When such systems produce an output, e.g. new text, a match of a face, or a prediction about someone, it is based on the model’s calculation of relationships between the input features: These are probabilistic (stochastic), being based on statistical estimates of correlation (not causality). Depending on the ML method used, it may not be possible to explain how the output was produced from the inputs (the system is a “black box”). This is usually the case with neural networks and transformer-based methods as in the large language models underlying chatbots.

Chatbots

  1. Chatbots produce an estimate of what is the most likely text to follow the preceding words (such as a prompt). Outputs vary for each prompt, so they are inconsistent and unpredictable, and there is no way of explaining what led to a particular output or — without additional work — knowing whether it is accurate. These problems cannot be fixed[2]. However, they are currently proposed to be useful in accessing online government information, simplifying interactions with the public in administrative procedures, or summarising policy consultation responses.
  2. The Government Digital Service experimented with a chatbot interface to gov.uk, finding “Overall, answers did not reach the highest level of accuracy demanded for a site like GOV.UK, where factual accuracy is crucial. We also observed a few cases of hallucination — where the system generated responses containing incorrect information presented as fact”[3]. These are inherent problems with such tools — making them unsuitable for this purpose and for the same reasons (plus the lack of explainability and consistency) for use in statutory administrative procedures.
  3. The claim that such tools could summarise policy consultations is correct, but their use for this purpose is wrong. The benefit of a consultation to policy makers is to find the things they had not thought of, identified by people they hadn’t spoken to — the needles in the haystack. A chatbot summary will not pick them out. It will not enable the nuanced positions on the policy by stakeholder groups to be ascertained. Further, even without the potential “hallucinations”, an automated summarisation does not fulfil the democratic function of a consultation, to allow all voices to be heard and shown to be heard. It falls far short of this, and in the process devalues and insults the considerable effort respondents make to be helpful in policy development[4].

Predictive systems

  1. A study co-authored by this respondent identified three dozen sources of risk when public bodies use algorithmic methods to make predictions about people or their circumstances, or to influence their behaviour[5]. Worldwide, instances of bias and inaccuracy are evident. Recent guidance from DfE and RTA on the use of data analytic tools in children’s social care[6] specifically warns against predictive uses, citing findings that they do not work.
  2. Prediction in relation to a person from a machine-learning algorithm essentially projects the statistical average result from a group of people onto an individual. This is fundamentally in conflict with the rule of law in a democracy where an individual’s case should be judged on their individual circumstances, not relative to statistically-derived parameters from a more-or-less representative group. A strong case can be made for the practice to be banned[7].

Facial recognition

  1. There are many varieties of potential applications of facial and other biometric recognition in administration, e.g. in immigration, and policing, and we do not propose to rehearse the extensive debate here. However, they are again stochastic in nature, based on a non-explainable ML process, and have been shown to demonstrate bias.

Recommendation 4: Any use of generative AI must be reported as part of the cataloguing as above, shown to be compliant with the Government’s own guidance on its use, justified in relation to cost, security, benefit and legality (including regarding intellectual property), and evidenced as sufficiently accurate for the purpose intended.

Recommendation 5: Any practices within government where algorithmic methods are used to identify or make predictions about individuals, including biometric recognition, are immediately suspended until a full examination of their compliance with the rule of law and principles of good public administration is made.

Conclusion

  1. Basing analysis, policy and action on an arbitrary scope for “AI” as a thing or loosely-defined set of things is not useful in this context. What matters is how government departments are using or intending to use any algorithmic methods in the execution of their functions, particularly those that directly affect the lives of people. This is currently not known, and the Government’s own guidance and tools need to be enforced in the cause of transparency and good governance.
  2. Much of the speculation about the potential application of “AI” in government relates to stochastic systems. We propose that any such system that produces statistically-derived outputs will fail to meet the criteria set by the rule of law and principles of public administration, and the standards of our democracy. Any decision or policy proposal based on one is unlikely to be able to satisfy a public or parliamentary demand to hear a principled justification for the decision or proposal.
  3. The current flurry of excitement and grand claims for how “AI” could transform government and the public sector is therefore unsupportable.

April 2024


[1] https://www.gov.uk/algorithmic-transparency-records

[2] https://joanna-bryson.blogspot.com/2024/02/2024-brief-overview-on-llm-foundation.html

[3] https://insidegovuk.blog.gov.uk/2024/01/18/the-findings-of-our-first-generative-ai-experiment-gov-uk-chat/

[4] https://www.theguardian.com/commentisfree/2024/apr/25/ai-public-voice-ministers-large-language-model-chatgpt

[5] https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3716166

[6] https://www.gov.uk/guidance/develop-and-use-data-analytics-tools-in-childrens-social-care

[7] Tafani, 2024 https://zenodo.org/records/10866778