Dr Mercedes Bunzwritten evidence (ACT0009)

 

House of Lords Communications and Digital Select Committee inquiry:

Scaling Up: AI and creative tech

 

Addressing the barriers to scaling creative AI in the UK: a focus on issues and solutions around data

 

My background and expertise

I am Professor in Digital Culture and Society at the Department of Digital Humanities, King’s College London, where I explore AI’s recent, groundbreaking ability to ‘calculate meaning’. I analyse computational capabilities and study their effect on societies through analysing public and commercial AI projects and the role data plays.[1] I am particularly invested in the creative sector as the co-founder of the Creative AI Lab, a collaboration with Serpentine Galleries; the lab works with artists, creatives and tech entrepreneurs. As the co-founder of two publishing companies, I have significant knowledge of creative tech and the start-up culture, including through my former career as a digital journalist and technology reporter of The Guardian.

 

Summary of evidence

This evidence focusses on issues creative industries face in the UK when working with AI, which makes foreign buyouts attractive: the legal and economic uncertainty of training data. For most creative industries it is sufficient to finetune existing models, and there is no need to build foundational models from the ground up. To finetune models, creative industries rely on sector specific training data. I draw my insights from my research including two cases my lab workshopped: tests of using generative AI in a UK fashion house, and the creation of an AI model generating choral music trained on recordings of 15 UK choirs. This evidence details current difficulties creative tech faces in using AI in the UK; it does not address labour shifts within the creative industries.

 

1.              To develop efficient and successful AI tools and services, UK businesses need to train or finetune their models using large datasets. These datasets are often re-purposed or scraped from the internet; they are rarely newly collected with consent of the data subjects. Given this, the situation around training data is still one of economic and legal uncertainty. To avoid those uncertainties, which can be financially painful, businesses move under the protective wing of bigger tech corporations usually based in the US. To keep businesses in the UK and allow them to scale up, the following barriers need attention:

 

1.1              The legal uncertainties around training data used by creative tech to train their models result in a significant financial burden due to the need of legal advice and the threat of potential prosecution. For example, one of the projects I studied spent £14,000 on legal fees; even though their data was taken with consent to train an AI model, there was no clarity about the status of data, model, and model output. Businesses also may shy away from using the right and relevant data for effective training, because this data has an unclear legal status. The result will be services and tools functioning insufficiently. In international comparison this will lead to a disadvantage.

 

1.2              Moreover, the different areas of law applicable to training AI systems and using their outputs make the legal situation cost intense – privacy and UK GDPR, intellectual property and copyright each applying differently to data, models, as well as model outputs. Start-ups and SMEs do usually not have an in-house legal team. This is one of the reasons why they prefer to embrace the opportunity to work under the wing of a bigger player such as Alphabet/Google, Amazon/AWS, Microsoft, Apple, Facebook, etc. usually based in the US. Thus, legal uncertainty is a particular disadvantage for the UK lacking big players, less so for the US.

 

Possible actions: To create helpful guidance about the current legal situation and its uncertainties and establish AI/creative tech surgeries. To advance AI-specific legislation that clarifies uncertainties. To support the creation and platforming of data in the creative industries and beyond.

 

1.3              Companies who use data without consent/knowledge of data subjects suffer from the looming danger regarding their reputation putting client and customer adoption at risk, and with it their business. A YouGov survey from 2021 found that adults in the UK (75%), but also the US (72%), are worried over control of their personal data; for the UK, this worry was again confirmed 2023 through a survey by the Ada Lovelace Institute.[2] Even if AI businesses keep the training material a business secret such as Open AI, they can be found out. For example, when prompted generative AI can create mimetic images or near-verbatim reproductions of written works they have been trained on. This backdoor allows to inquire training material, because AI models can only generate what they have seen. The court case New York Times vs. Open AI launched in 2023, uses ‘memorised’ training data as evidence.

 

1.4              Creative tech is furthermore facing an ethical challenge: operating under fair principles is particularly important for the creative industries. The creative market will avoid AI companies developing creative tech and tools that take unfair advantage of their peers. To counter such an impression and demonstrate their good will, Stability AI, the London-based AI developer of generative AI systems such as ‘Stable Diffusion’, allowed artists to remove their work from the two billion images in the LAION-dataset used to train its Stable Diffusion 3 model. They received 40,000 requests by artists removing 80 million images, according to the AI start-up Spawning that provided the opt-out technology ‘Have I been trained?’. In the UK, Stability AI faces a court case for infringing on Getty’s copyrights to train its models to be heard in summer 2025; the legal proceedings in the English High Court commenced in January 2023. The best solution, however, is to collect data with consent hosted in ‘data trusts’ steered by a data steward.

 

Possible actions: to fund and showcase flagship examples of best data practice.

 

1.5              In short, innovation in the UK economy is being stifled due to looming costs caused by legal uncertainty about the status of data. Addressing these barriers is essential. The issues that come to light in the creative sector wanting to use AI are shared problems applying to all UK companies that aim to work with training data.

 

Data = UK’s asset

2.              For the UK this is more problematic than for other nations, because data is one of our nation’s strengths and assets. This has been seen early – the Science and Technology Committee report ‘The Big Data Dilemma’[3] spoke of the UK already in 2016 as a ‘world leader’ noting that ‘data has huge potential value to the UK, both as a driver of productivity and as a way of offering better products and services to citizens’; the committee’s 2013 report declared data one of eight great technologies for future growth. Data gives the UK an advantage, as other countries race ahead to build publicly available foundational models, for example the supercomputer of the Swiss National AI Institute available to public, private, and non-profit organizations, likewise Denmark with its cutting-edge Danish language model available for commercial and public organisations or Sweden’s large language model aimed to support its public sector. The UK has the Isambard AI, but maybe even more valuable: data.

 

2.1              UK’s data landscape is unique offering a vibrant ecosystem of commercial and public organisations that specialize in data collection, analysis, and services across various sectors. This spans from the UK Biobank, Health Data Research UK and NHS Digital to YouGov, Ipsos UK, Acorn/CACI as well as Experian UK or the Office for National Statistics, UK Data Service, Ofcom, and Transport for London, among others. The culture and conversation about data is enriched by the Alan Turing Institute, Open Data Institute and Open Knowledge Foundation, and supported by the British Library’s Data Services among others.

 

Possible action: To create a stronger link among existing players across the public and private divide could surface new opportunities.

 

2.2              Moreover, data as a resource is particularly strong regarding cultural collections and creative industries including big data players such as the British Library Digital Collections, National Archives, BBC’s digital archives and the British Film Institute as well as British museums such as Tate, V&A, National Gallery but also Art UK. UKRI’s push towards a digital infrastructure for its natural science collections under the name of DISSCO UK (Distributed System of Scientific Collections) will add to this.[4]

 

2.3              Transforming UK’s data wealth into the high-quality datasets which companies and public institutions need to innovate should be a goal, the more as data can be used by the UK to actively address five of the twelve challenges of AI governance, which the Science, Innovation and Technology Committee noted in its 2024 report.[5] For example, the creation of datasets can be used to balance the danger of bias, to reduce the need for processing personal data, to introduce transparency into AI’s black box, to avoid IP and copyright infringement, through providing access to data.

 

2.4              Given that data is one of UK’s central assets and steering mechanisms, the conversations and visions regarding data need profound reframing: the government’s own 2024 survey into the public’s attitudes to data surfaced that as much as one third expect that only certain privileged groups will benefit from data.[6] Data deserves to be seen as something that is creating opportunities for society as much as for businesses. The government’s actions will be decisive for a much-needed reframing of data from a culture of fear to a culture of care and opportunities.[7] This can only be achieved through thinking public investment and industry opportunities together.

 

‘Regulate to innovate’ - empower and enable creative innovation through AI

3.              While it is clear that ‘regulation represents the missing link in the UK’s overall AI strategy’,[8] more can be done. The list below focuses on actions that a) make the data ecosystem more productive and b) support businesses to overcome existing barriers.

 

3.1            Actions to make the wider ecosystem of data more productive:

 

3.1.1     Short-term: The UK should update its outdated 2020 National Data Strategy[9] in light of AI and embrace data as an asset for public projects and industry, besides strengthening access to hardware (cloud, GPUs).[10]

 

3.1.2     Long-term: The UK should create a British Library of Data reflecting the fact that data is a new form of knowledge bringing in societal and economic opportunities.

 

3.1.3     Mid-term: Flagship projects that celebrate the profound benefit of data for the public and the immense opportunities for businesses should be showcased to highlight UK’s strong position in the data economy and to reframe the public’s and businesses’ view on data. The UK should strategically support data projects that generate trust and showcase best practice.

 

3.2            Actions to help businesses overcome issues with data allowing them to scale:

 

3.2.1     Address the information needs of creative tech regarding legal uncertainties now: provide guidance for legal issues as well as orientation and insights through a data surgery for businesses assisting with a constantly changing landscape.

 

3.2.2     Help businesses to find data by creating a central platform hosting data equivalent to Health Data Research UK.

 

3.2.3     Create examples of best data practice i.e. collect and curate data sets with consent of data subjects or allow the possibility to opt-out.

 

3.2.4     Incentivise and advertise the role of a data steward who ensures integrity, governance, security and acts as an ombudsman for data subject inquiries.

 

3.2.5     Support the creation of publicly available data tools assisting businesses in allowing data subjects to query data sets such as the tool ‘Have I been trained’.

 

Addressing legal uncertainties around data in AI and building a strong data ecosystem consisting of available data collections showcasing best data practices will remove an existing barrier to scaling up in the UK.

 

 

14 October 2024

5


 


[1]              See for example the publication Bunz, Mercedes and Photini Vrikki, ‘From Big to democratic data: why the rise of AI needs data solidarity’, Democratic Frontiers edited by M. Filimowicz. Taylor & Francis, 2022.

[2]              Ada Lovelace Institute, ‘Evidence review: What do the public think about AI?’ [report] 2023.

[3]              Science and Technology Committee, House of Commons, ‘The big data dilemma: Fourth Report of Session 2015–16’, February 2016, p. 7.

[4]              UKRI, Infrastructure Fund Projects [website announcement], August 2024.

[5]              Science, Innovation and Technology Committee, Governance of Artificial Intelligence (AI) [report], May 2024.

[6]              Department for Science, Innovation, Technology, ‘Public attitudes to data and AI: Tracker survey (Wave 3)’, February 2024.

[7]              Pasquale, Frank, ‘Data-Informed Duties in AI Development’, Columbia Law Review, vol. 119, no. 7, November 2019, pp. 1917-1940.

[8]              Ada Lovelace Institute, ‘Regulate to innovate: the future of regulation. A route to regulation that reflects the ambition of the UK AI Strategy’ [report], November 2021.

[9]              Department for Digital, Culture, Media & Sport, National Data Strategy [policy paper], December 2020.

[10]              The UK should embrace a similar strategy as the EU who plans to make High-Performance Computing available to startups, industry, and research.