Open Data Institute—written evidence (LLM0083)

 

House of Lords Communications and Digital Select Committee inquiry: Large language models

 

 

About the ODI

The Open Data Institute (ODI) is an independent, non-partisan, not-for-profit organisation founded by Sir Nigel Shadbolt and Sir Tim Berners-Lee in 2012. We have a mixed funding model and have received funding from multiple commercial organisations, philanthropic organisations, governments and intergovernmental organisations to carry out our work since 2012.

 

The ODI wants data to work for everyone: for people, organisations and communities to use data to make better decisions and be protected from any harmful impacts. We work with companies and governments to build an open, trustworthy data ecosystem. Our work includes:

 

       consultancy: working with organisations in the public, private and third sectors, building capacity, supporting innovation and providing advice

 

       research and development: identifying good practices, building the evidence base and creating tools, products and guidance to support change

 

       policy and advocacy: supporting policymakers to create an environment that supports an open, trustworthy data ecosystem

 

Our 5 year strategy sets out what we think are the elements of an open and trustworthy data ecosystem for a world where data works for everyone. Our approach allows us to adjust our implementation and engagement as the world around us, and the organisations we work with, change. Our activities will be set out on an annual basis, mapped to the six principles that guide everything we do:

 

1.     We believe that a strong data infrastructure is the foundation for building an open, trustworthy data ecosystem on a global scale and that this can help address our most pressing challenges.

 

2.     Strong data infrastructure includes data across the spectrum, from open to shared to closed. But the best possible foundation is open data, supported and sustained as data infrastructure. Only with this foundation will people, businesses and governments be able to realise the potential of data infrastructure across society and the economy.

 

3.     For data to work for everyone, it needs to work across borders – geographic, organisational, economic, cultural and political. For this to happen ethically and sustainably, there needs to be trust – trust in data and trust in those who share it.

4.     There is greater need than ever for trusted, independent organisations to help people across all sectors, economies and societies to benefit from better data infrastructure.

 

5.     For data to work for everyone, those collecting and using it need to be highly alert to inequalities, biases and power asymmetries. All organisations working in data must take proactive steps to ensure that they contribute fully and consciously to creating a diverse, equitable and inclusive data ecosystem.

 

6.     The world needs a new cohort of data leaders individuals who have data knowledge and skills and are equipped to understand the value, limitations and opportunities offered by data, data practices and data sharing.

 

 


 

The ODI’s approach to this consultation

 

Since its inception, the Open Data Institute has been committed to our mission: to work with companies and governments to build an open, trustworthy data ecosystem. As our chair, Sir Nigel Shadbolt told the Commons Science and Technology Committee in February:[1] “Although we are talking a lot about AI, for the algorithms, their feedstock—their absolute requirement—is data.” This means that, as the Open Data Institute has argued since its creation, ‘the need to build a trustworthy data ecosystem is paramount’.

 

If we want to make the most of the opportunities presented by AI and LLMs, while ensuring that we do all we can to mitigate the risks, we need to think about data. We need to ensure that, as much as possible, the data used is open - accessible, available and assured. We need to ensure that the data that needs to be protected is protected; that we have the right data infrastructure to do that properly; and that we have the data literacy, the right governance (including participatory data governance) and a sufficient grasp of data ethics to build public trust in data use and prevent its abuse.

 

 

Capabilities and trends

 

1. How will large language models develop over the next three years?

 

a)  Given the inherent uncertainty of forecasts in this area, what can be done to improve understanding of and confidence in future trajectories?

 

In evidence to the House of Commons Science and Technology committee, chair of the ODI Sir Nigel Shadbolt noted that ‘these large models have a fundamentally statistical character, so there will always be some sense of indeterminacy in them’ – this requires some thinking about how best to look at the ways in which such models ‘come to their conclusions and where their frailties are’. Academia (and civil society) will need to be able to ‘kick the tyres on these systems just as hard… as the industrial researchers do in the big tech platforms.’

 

Large language models pose questions for regulators but also for others in the ecosystem. One point raised at the ODI workshop was ‘if I provide a dataset, am I legally responsible for the AI outcomes?’ What are current responsibilities or legal repercussions – and how might this change with new regulation?

 

This suggests a few ways that understanding and confidence could be improved. One would be to provide greater access to models (and the data behind them) to academics, civil society organisations and other experts – greater understanding of how they are currently being developed could aid future forecasting.

 

Another would be to build a horizon scanning function – the AI regulation white paper highlights this as one of the ‘central functions’ that should be developed to support individual regulators. Any horizon scanning function should learn from and build on existing government practice and institutions, including the Futures, Foresight and Emerging Technologies function in the Government Office for Science; a landscape review of where horizon scanning is already taking place in government, and historical examples of how it has been deployed in fast-moving fields, should be considered. Whilst the market for LLMs is dominated by other countries, greater focus should be placed on developing a deeper understanding of the barriers that are preventing the UK from developing its own LLM, and becoming a leader in the field.

 

 

The AI Index report at the Stanford Institute for Human-Centred Artificial Intelligence in the USA offers an example of an initiative currently doing well in supporting decision-makers to understand the landscape of AI and LLMs. The initiative covers ethics, regulatory issues, public opinion and economic impact amongst others. The UK should aspire to establish its own UK focused index and ensure that such an initiative is properly resourced in order to replicate the effectiveness of the AI Index Report, and to deliver true value to the UK as it works to improve its understanding and confidence in future trajectories of LLMs.

 

In the UK, the emphasis of such an initiative would be different than in the US, exploring UK-specific institutions and priority sectors. UK government practices, and the way data is used in large language models, will be different. For example, the UK needs to explore the role of LLMs for the National Health Service, with attention to how data is collected and governed. At the moment, public services such as the NHS do not have the capacity or developed capabilities to share data for research processes and there is work to do in understanding how this could be achieved. Other areas would include exploring the impact of foundation models in the UK. We cannot understand the future impact of foundation models without developing a horizon scanning initiative such as the AI Index Report work.

 

There may also be general points to be made about better understanding of what data is available on LLMs, and understanding what data is going into large language models. (There might be a link with wider data and evidence gaps around data and digital, as we touched upon in our blogpost on researcher data access and the Online Safety Bill.)

 

It is important to understand public perceptions of AI following the launch of open access tools based on LLMs such as ChatGPT. The most recent survey on public perception of AI was run by the Alan Turing Institute together with Ipsos and generated useful findings, but was run before this launch. It is important to understand public opinion and trust to steer future implementation.

 

 

2. What are the greatest opportunities and risks over the next three years?

 

a)  How should we think about risk in this context?

 

Opportunities include LLMs being developed that can augment existing jobs and support people to do their jobs better – for example, in allowing tools to be built that can help with drafting and structuring written work. We wrote about some of these examples in our recent response to the Generative AI in Education Consultation.

 

Risks are that they can be deployed, based on biased, incomplete or otherwise limited data, in sensitive scenarios where they should not be used; lead to jobs being cut where LLMs and tools built upon them are not capable of replacing roles; and not living up to expectations or causing harms which then damage trustworthiness for more sensible uses of LLM, generative AI and other AI-based applications.

 

There are also environmental risks, given the volume of compute required (in our response to the AI regulation white paper, we suggested measuring compute as a potential tool in the governance of foundation models and could help encourage more environmentally-friendly innovation).

 

Finally, we have seen that whilst the government has identified ‘AI safety’ as a critical theme for the AI Summit hosted in the UK this November, concerns are being raised by some that the agenda on ‘AI safety’ is being shaped by big tech organisations, and therefore do not acknowledge or recognise a number of risks. This could be a missed opportunity to make the summit relevant for, and driven by voices from researchers, civil rights groups, sector specialists and wider society. In order to most effectively think about risk, and to ensure we mitigate them, whilst leveraging opportunities, we must take care to include more voices.

 

 

Domestic regulation

 

3. How adequately does the AI White Paper (alongside other Government policy) deal with large language models? Is a tailored regulatory approach needed?

 

a)  What are the implications of open-source models proliferating?

 

In general, we agree with the government’s fundamental approach to AI regulation – sector-specific regulators being supported by central functions to apply cross-cutting principles in their domains but as stated in our response to the AI White Paper, we think the government should provide more information on funding, powers and the nature of the central functions and how they will operate.

 

Given their foundational nature and a particular need to look at the ways in which such models ‘come to their conclusions and where their frailties are’, we believe that a tailored regulatory approach is required for LLMs.

 

A key area where AI regulation should be tailored to account for implications of large language models is their relationship with copyrighted content and changes in the dynamics of public common goods like open data sets/ data sets that have been created from collaborations. Many have become concerned about data rights, biases and flawed datasets, and a lack of public trust in the data that feed these models. This concern applies for all generative models, not only for large language models.

 

For example, we are aware that there may soon be a precedent for regulations that address transparency about training datasets used in large language models. In the US, the Federal Trade Commission has recently ordered OpenAI to document all sources of data used to train its large language models (Zakrzewski, 2023). A group of large media organisations have published an open letter urging lawmakers around the world to introduce new regulations to require transparency of training datasets (Agence France-Presse et al, 2023). We should see demands for information about training data as but the latest wave in an ongoing push for corporate transparency. In the UK, laws around the mandatory registration and publication of information by companies go back to the 1800s, and over this time, regulators have developed standardised approaches to avoid each company choosing its own way to report on its finances and other activity. Perhaps we need the same for disclosures about the data that foundation AI models have been trained on (O’Reilly, 2023).

 

There is a need for the regulatory response to address the potential impact of generative AI in elections, and especially on social media. For now, if you share generated media from a different platform it is outside of the scope of moderation. As such, at the moment social media content policy deals with generative AI content in very limited ways.

 

On the proliferation of open source models, there may be some issues around data/AI literacy – the question of how we equip people to understand which models to choose, so they ask the right questions in deciding which models are most effective and ethical for their purposes is one that we need to pay more attention to.

 

For the most part, large language models are currently not open source. While some examples have emerged from Facebook, these are still not fully open source. To generate open source LLMs, this area needs to be resourced, with attention to which data would be used and the extent of the compute required. Big tech organisations are now investing in this area - however, in the case of OpenAI for example, they had stated an intention for the technology to be open source, but this did not materialise.

 

 

4. Do the UK’s regulators have sufficient expertise and resources to respond to large language models? If not, what should be done to address this?

 

Given this is an emerging and quickly evolving field, it is unlikely that even the UK’s specialist information and digital regulators (the first mention of ‘large language model’ on the ICO website appears to have been in April 2023) currently have sufficient expertise and resources, let alone the sectoral regulators who may expect to regulate LLMs in their domains.

 

The Digital Regulation Cooperation Forum could start some work to understand what will be required of regulators around LLMs, how other countries are adapting (and any comparisons with fast-moving technological change in other sectors), and what training may be required. (Note, though, discussions around whether the DRCF as currently constituted is effective - it might need more funding/resource/reform to do even this.)

 

There are also broader questions of how data and AI literacy should be taught in schools and to the wider population.

 

Finally, where other countries are working with their own large language models, the UK has currently not allocated the resources to keep up with this. The latest models cost in the region of the full budget that the UK’s AI Taskforce has been allocated.

We would suggest that regulators need significantly larger funding allocations to effectively respond.

 

 

5. What are the non-regulatory and regulatory options to address risks and capitalise on opportunities?

 

a)  How would such options work in practice and what are the barriers to implementing them?

 

b) At what stage of the AI life cycle will interventions be most effective?

 

Our belief is that interventions should take place at every stage of the AI life cycle in order to be most effective in addressing the potential risks and capitalising on opportunities.

 

c) How can the risk of unintended consequences be addressed?

 

When considering unintended consequences, we should differentiate between unintended consequences with low level impacts, and those with significant and / or discriminatory unintended consequences that can harm or negatively impact users or populations.

 

On the regulatory side, there are various approaches that may be supportive to addressing and mitigating unintended consequences, alongside the horizon scanning we wrote about earlier in this consultation response. Options include mandating greater transparency around and/or researcher access to data about models. Any transparency measures should consider the intended audience and the desired effect and operators of AI systems should consider the explainability of what they are publishing to non-expert audiences – as the research informing the government’s own algorithmic transparency reporting standard shows, there are different transparency requirements for different audiences (simple and understandable for non-expert members of the public, much more detailed for independent experts in academia and civil society). Greater public awareness and AI literacy will be required to help people - from members of the public to senior decision makers - understand the details being released. Further, the government should consider mandating that particular models, for example those used in government and high stakes situations, are independently evaluated. In addition, it may be useful to explore solutions such as international bodies who can review models.

 

On the non-regulatory side, options include building UK sovereign capability in LLMs to support wider innovation and research, something Sir Nigel Shadbolt (and Professor Dame Wendy Hall) spoke about in front of the Science, Innovation and Technology select committee. Sovereignty will also allow researchers fuller access to ‘kick the tyres’ of the models, as we would expect industrial researchers to do at the big tech companies.

 

There is something to be said about the data going into models here, and understanding bias, limitations etc.

 

In terms of the stage of the life cycle, we have noted that while the AI White Paper stipulates that compliance requirements should be determined before systems are designed, it is not clear how and when monitoring of these requirements will happen. It is, of course, important that later and continuous monitoring should be implemented and we would advise generating terms of use of UK data that has been collected and used to train existing models so as to learn from their outputs and any unintended consequences.

 

 

International context

 

6.  How does the UK’s approach compare with that of other jurisdictions, notably the EU, US and China?

 

a)  To what extent does wider strategic international competition affect the way large language models should be regulated?

 

b) What is the likelihood of regulatory divergence? What would be its consequences?

 

Broadly speaking, the UK and US are currently pursuing similar approaches, issuing guidance and non-binding principles with no new central regulator being created that regulate outcomes; the new EU regime is statutory and more concerned with regulating AI systems, rather than outcomes; and China’s approach is more concerned with controlling the flow of information, and focuses on recommendation algorithms and synthetically generated images and chatbots.

 

We anticipate that regulatory divergence is highly likely, and already underway. However, international harmonisation and collaboration will obviously be useful for example, the UK’s foreign policy aims (in the integrated review) around ‘regulatory diplomacy’, and AI being an area mentioned as one of the UK’s strengths in the AI White Paper. Regulatory divergence could pose challenges for companies looking to operate in multiple jurisdictions, and this is likely to lead to less collaboration and fewer opportunities to work on global challenges.

 

Another related concern is the potential formation of conflicting ‘blocs’ taking very different approaches to LLMs. The nature of this challenge is not specific to AI and LLMs - this challenge is linked to the current geopolitical environment in general, and the fight for dominance, being played out through the competition over development of chips and an AI “arms race” among other areas.

 

Further, international collaboration should be used to ensure that models are produced in equitable ways and work for underrepresented groups and populations. Examples include making LLMs more environmentally friendly (helping them to achieve more with less data, and constraining levels of compute required), ensuring they represent diverse populations and minority or underrepresented groups where there is little or no data, and representing marginalised voices in processes. As part of this, countries can exchange best practices in involving the public and incorporating feedback.

 

This could also be an opportunity to pitch for more participatory data and AI governance, and the possibility of the UK becoming a world leader on this specifically.

 

 

August 2023

9

 


[1]              https://www.theodi.org/article/ai-excitement-is-everywhere-but-we-need-to-talk-about-data/#:~:text=Co mmittee%20in%20February-,(see%20video%20here),-%E2%80%98although%20we%20are