Written evidence submitted by Dr Julian Padget (UAIG0021)

Evidence in respect of “Use of AI in Government”

  1. This response primarily addresses “risks and opportunities of AI adoption in government” with secondary relevance for “departmental accountability on AI delivery, funding and implementation”.
  2. I write as an academic with 30 years' experience researching in the field of AI and 5 years' experience in AI-related standards-making at BSI, ISO, IEEE and CEN with a particular focus on algorithmic bias. At ISO I contributed to the writing of ISO 8183:2023 Data life cycle framework, at IEEE to IEEE Std 7003-2024 Algorithmic Bias Consideration and at CEN JTC21 I am contributing to the response to the European Commission’s Standardisation Request in respect of the EU’s AI Act.
  3. The central point of my evidence is that the effective management of bias is critical to the risk and opportunities offered by AI, while delivery of departmental accountability is a direct corollary of effective management. The biases that are present in data are essential to the correct functioning of AI systems but at the same time they are a constant source of risk. Accountability for the effective management of bias and the risks and opportunities it entails cannot be outsourced, which has implications for data and skills issues within Government departments
  4. I support this evidence through:
    1. Explaining why bias is both opportunity and risk, and requires robust management practices to realise the former and avoid the latter; see paragraphs 4, 5, 6
    2. Describing an approach to layered oversight within AI solution-owning departments; see paragraphs 7, 8, 9
    3. Arguing that AI should be a justified choice not a default choice; see paragraph 10
    4. References; see paragraph 11.
  5. Bias and fairness are interdependent: the conventional connotation of the word ‘bias’ is a negative one The concept of bias needs reframing for it to be fully understood in the context of its use in AI. Conversely, fairness is regarded positively and viewed as a universally satisfiable notion. It is not (Kleinberg et al., 2017). Fairness is relative (Floridi et al., 2020): what is fair for one may be unfair for another. Furthermore, delivery of “fair” outputs from AI systems depends on the ethical use of bias. In short, bias entails fairness.
  6. The ethical use of bias demands some explanation, while also being the controlling condition for the effective management of bias in AI development and operation. Unbiased data is of no more use than flipping a coin to decide. What matters is the presence of wanted bias that contributes to meeting business requirements and the minimisation of unwanted bias that either results in incorrect outputs now or is in effect a bug waiting to emerge later. In contrast to conventional programming, where a human writes the code, in (many) AI systems, it is the data that drives an algorithm that writes the code. What both resulting systems have in common however is that, in the first, testing only proves the presence of bugs, never their absence, while in the second, testing proves the presence of unwanted bias, not its absence. It is not possible to prove the absence of unwanted bias in an AI system. Bias is, in effect, what programs an AI system, and enables it to do something useful, yet it can easily also embed multiple bugs, thus its ethical use is central to the development and operation of AI systems that are safe.
  7. Management of bias is a whole life cycle responsibility. Bias in data is essential for building a useful AI system. Unfortunately, such a system is effectively obsolete from the day building is completed because the bias in future input data will steadily drift away from that in the original data. In short, previously wanted bias can become unwanted and unwanted bias can become wanted. Permanent oversight of such systems, to manage their ongoing use of bias is crucial to risk management and opportunity realisation in the use of AI and hence to departmental accountability.
  8. Oversight needs to take place at three levels: continuous evaluation, operational monitoring and organizational governance (IEEE Std 7003-2024).
    1. Continuous evaluation means that system input and outputs need to be assessed – at least in part, automatically – throughout the operational lifetime of the system. AI systems are not fit-and-forget pieces of software because, as explained above, their behaviour and efficacy change over time with respect to the data they are fed, and continuous evaluation is the canary in the coal mine that alerts us before hazards become accidents.
    2. Operations monitoring means that humans need to review the results of the continuous evaluation at an appropriate cadence. It puts the human on – not in – the loop because human-in-the-loop is not the panacea it is presented to be and does not absolve organizations of accountability for effective management. In general, a human embedded in an automated system, unless highly trained – like a pilot, despite Air France flight 447[1] – either lacks the situational awareness to be able to take over or becomes a victim of automation bias (Green, 2022). Whether true or not, cases such as those associated with the Home Office processes have created a public reputation for governmental use of AI that could be hard to counter and further similar cases could kill off the drive for AI adoption in government.
    3. Governance means that departmental processes need to consider operations reports as a standing item, which further implies competence to judge and act upon such information at all managerial levels (ISO 42001:2023).
  9. Operation and maintenance costs of software solutions containing AI may well be higher than for conventional software. An underlying premise to the use of trained models is that the future will look like the past. This can be largely true but depending on the problem – for example, revision is essential for a recommender system – the speed of divergence through drift varies (cf. weather forecasting, where tomorrow is frequently like today but next week is not). The two remedies for drift – retraining and continuous learning – both have drawbacks: determining the optimum frequency plus re-running all the bias metrics for the first, and risk of unintended or malicious manipulation for the second, but one or the other is inevitable. Policy and process for lifelong operations and maintenance, including audit, are part and parcel of AI-solution ownership.
  10. Procurement of software systems always exposes the buyer, even those with good tech knowledge, to risk (IEEE P3119). Enthusiasm for AI solutions needs to include awareness that AI is not the default solution. If a simpler, non-AI, solution is possible, or meets most of the business requirements, it then has the potential to reduce development time and cost (solutions with AI-inside may well be more costly simply because of that) as well as reducing cost of ownership through simpler oversight and operations and maintenance requirements (paragraphs 6, 7, 8). If there is a strong case for AI to be part of the solution, a department procuring such a system needs to develop both the skills to evaluate proposals fully and ensure they have the resources in place to carry out the oversight that implements the effective management of bias through to retirement and decommissioning.
  11. Bibliography
    1. Floridi, Luciano, Josh Cowls, Thomas C. King, and Mariarosaria Taddeo (2020). “How to Design AI for Social Good: Seven Essential Factors”. In: Science and Engineering Ethics 26.3, pp. 1771–1796. DOI: 10.1007/s11948-020-00213-5.
    2. Green, Ben, The Flaws of Policies Requiring Human Oversight of Government Algorithms (April 26, 2022). Computer Law & Security Review, Volume 45, 2022, URL: http://dx.doi.org/10.2139/ssrn.3921216 
    3. IEEE P3119, IEEE Draft Standard for the Procurement of Artificial Intelligence and Automated Decision Systems. Unpublished. URL: https://standards.ieee.org/ieee/3119/10729/ 
    4. IEEE Std 7003-2024, IEEE Standard for Algorithmic Bias Consideration, 2025. URL: https://standards.ieee.org/ieee/7003/11357/ 
    5. ISO/IEC 8183:2023 Information technology — Artificial intelligence — Data life cycle framework, 2023. URL: https://www.iso.org/standard/83002.html 
    6. ISO/IEC 42001:2023 Information technology – Artificial intelligence – Management system. International Standards Organization, 2023. URL: https://www.iso.org/standard/81230.html 
    7. Kleinberg, J. M., S. Mullainathan, and M. Raghavan. “Inherent Trade-Offs in the Fair Determination of Risk Scores”. In: 8th Innovations in Theoretical Computer Science Conference. Ed. by C. H. Papadimitriou. Vol. 67. LIPIcs. Schloss Dagstuhl - Leibniz-Zentrum für Informatik, 2017, 43:1–43:23. DOI: 10.4230/LIPICS.ITCS.2017.43.

 


[1] https://en.wikipedia.org/wiki/Air_France_Flight_447