The Copyright Licensing Agency Ltdwritten evidence (LLM0026)

 

House of Lords Communications and Digital Select Committee inquiry: Large language models

 

 

Introduction

 

About The Copyright Licensing Agency Ltd (CLA)

The Copyright Licensing Agency Ltd (CLA) is a collective management organisation as defined in The Collective Management of Copyright (EU Directive) Regulations 2016.

 

CLA is the recognised UK collective rights management organisation for collective licensing of extracts from text and images from book, journal and magazine content (including some websites) to the education, business and public sectors. CLA exists to simplify copyright for content users and copyright owners. Our mission is to help customers legally access, copy and share the published content they need, while also making sure that copyright owners are paid for the use of their work. We have been providing licences as well as a growing range of related services, that simplify copyright and make it easier to access content for 40 years.

 

Collective licensing is a cost-effective blanket licensing solution and offers a practical alternative where it is not easy to license on an individual basis for specific uses due to the volume of rightsholders/users and the scale of use.  Since it is not possible to take account of the exact rights ownership of each extract which may be copied or used, the licence fees are shared between all the relevant rightsholders.

 

CLA is a not-for-profit organisation.  It has four members: Authors Licensing and Collecting Society Ltd (ALCS), Design and Artists Copyright Society (DACS), PICSEL Ltd (Picture Industry Collecting Society for Effective Licensing) and Publishers’ Licensing Services Ltd (PLS) and distributes the revenue it collects to its members, who in turn distribute to authors, publishers and visual artists.

 

CLA is a member of the British Copyright Council (BCC) and supports the separate response to this call for evidence made by the BCC.

 

Questions

 

Capabilities and trends

1.              How will large language models develop over the next three years?

a)              Given the inherent uncertainty of forecasts in this area, what can be done to improve understanding of and confidence in future trajectories?

 

To improve understanding and build confidence, the UK Government must establish a robust regulatory framework. It should ensure that each of the cross-sectoral regulators (or a dedicated single regulator for AI, should one be established) are proactive and continuous in their engagement with developers of LLMs to ensure they are well informed of current activities - including an understanding of how the LLMs are trained, the processes involved and what data they are trained on and technological developments in this area. Regulators must be equipped to identify and address the risks that may arise from these developments to minimise potential for harm to individuals and to uphold the law. For example, in ensuring that the rights of individuals (including the rightsholders that CLA and its members represent in the creative industries) are preserved and maintained.

 

 

2.              What are the greatest opportunities and risks over the next three years? a) How should we think about risk in this context?

 

There are serious concerns for the rightsholders CLA represents around the unauthorised ingestion of copyright protected works, author recognition, fair remuneration and compensation and trustworthiness of outputs, amongst other issues. These are current issues and present significant risks to the creative industries.

 

The House of Lords report ‘Arts and the Creative Industries: The Case for a Strategy’, recognised that the creative industries are worth £109bn annually to the UK economy. However, failure to adequately address the ethical and legal risks posed by LLMs will severely undermine not only the economic value of the creative industries but the UK’s internationally respected ‘gold-standard’ copyright framework.

 

We further detail these risks as follows:

 

 

 

 

 

 

 

With regards to opportunities, collective licensing in particular, where rightsholders can choose to be included (or not), presents an opportunity to address some of the risks posed by LLMs. The collective licensing, as operated by CLA is a cost-effective blanket licensing solution and offers a practical alternative where it is not easy to licence on an individual basis for specific uses. CLA currently offers a licence which includes text and data mining to Media Monitoring Organisations (MMOs). It is worth noting that the MMO industry has grown significantly in the last fifteen years and is now worth £2.8 billion. MMOs generate revenue from a suite of products and services all of which are based, to a varying degree, on published content. MMOs licensed by CLA generate significant and growing revenue from the analysis and insight provided by mining published content. The licences offered by CLA are a flexible and practical solution to a global technology-focussed sector, reliant on utilising rightsholders works.

 

Moreover, CLA offers licences to all schools and universities. The increasing use of AI and LLMs in educational settings raises issues related to copyright infringement and, without adequate regulation, risks undermining the existing licensing framework. The licences offered by CLA, as well as other CMOs and rightsholders, ensure that education users have access to high quality education material to utilise in teaching and learning activities, which ultimately benefits educational outcomes, and rightsholders are fairly remunerated for the use of their work by educational institutions.

Domestic regulation

 

3.              How adequately does the AI White Paper (alongside other Government policy) deal with large language models? Is a tailored regulatory approach needed?

a)              What are the implications of open-source models proliferating?

 

The implications of open-source models proliferating without adequate and tailored regulation are the risks highlighted in our response to Q2 above – overall, potentially very damaging to the UK’s creative industries if rightsholders are not fairly compensated nor given a choice as to how and when their works are used. CLA believes a tailored regulatory approach is needed to address the rights of ‘content producers’. The AI White Paper sought to exclude the ‘balancing the rights of content producers and AI developers’ from the scope of the proposal for an AI regulatory framework and we feel that this is a flawed approach.  By failing to adequately consider the rights of creators and rightsholders in regulation, this will undermine not only the principles outlined in the White Paper (particularly transparency and explainability, accountability and governance and fairness) but also the economic value the creative industries can continue to deliver to the UK economy in the face of the challenges of open-source model proliferation.

 

Any principles for regulation must explicitly address the requirement for LLMs to be transparent as to the provenance of training data and labelling of outputs. Regulation must also require developers of LLMs to operate within the UK’s copyright framework and to have appropriate licences in place, with the opportunity for rightsholders to choose whether they want their works included in the LLM or not.

 

CLA welcomes the Government’s five principles to facilitate the safe and innovative use of AI. We have yet to see the detail around what each principle entails, and how these would each be implemented and monitored but look forward to further engagement with the UK Government on this in due course. 

 

 

4.              Do the UK’s regulators have sufficient expertise and resources to respond to large language models?[5] If not, what should be done to address this?

 

The UK Intellectual Property Office (IPO) regulates Collective Management Organisations (CMOs) under The Collective Management of Copyright Regulations 2016. It does not currently have a broad regulatory authority and will require resource and knowledge to effectively regulate LLMs. Cross-sectoral regulators must work with rightsholders, their representatives and the IPO to develop a means to regulate LLMs effectively, uphold and oversee compliance with the UK copyright legal framework.

 


5.              What are the non-regulatory and regulatory options to address risks and capitalise on opportunities?

 

The scope of the regulatory framework as outlined in the AI White Paper must be expanded upon to include the ‘rights of content producers’. CLA believes that broadening the scope will go some way to address the risk to the creative industries and will help to ensure opportunities for LLMs will be developed safely, ethically and legally.

 

With regards to a non-regulatory approach, CLA is involved in the working group for the Code of Practice for copyright and AI, being led by the IPO. We are supportive of this work and welcome the opportunity to contribute. The work being done to develop the voluntary code of practice is beneficial in helping to understand where there are areas of agreement between rightsholders and AI firms, and where there are areas that require further investigation. A voluntary code of practice on its own however, is unlikely to be sufficient and regulation is likely to be needed to protect the rights and interests of rightsholders and preserve the value to the UK economy of the copyright framework.

 

a)               How would such options work in practice and what are the barriers to implementing them?

n/a

 

b)              At what stage of the AI life cycle will interventions be most effective?

 

In relation to rightsholders, interventions will be most effective at the data ingestion and training stage. Developers of LLMs should be required to adopt principles similar to those required by data protection law in UK and Europe (‘privacy by design’) and evaluate at the outset the provenance of training data, if permission from rightsholders to use the data for the intended purposes has been secured and ensure rightsholders have the option to choose whether their works are included or not and include clear attribution where permission is granted.

 

c)              How can the risk of unintended consequences be addressed?

 

Robust regulation to address risks and to capitalise on opportunities is needed to ensure unintended consequences are minimised. Regulators will need to be aligned in their approach and flexible to ensure they can act quickly and the copyright legal framework in the UK must be upheld. Further, work needs to be done to develop industry-wide best practices in the development, deployment and operation of LLMs again as a measure to mitigate the risk of unintended consequences occurring.

 

 

International context

 

6.              How does the UK’s approach compare with that of other jurisdictions, notably the EU, US and China?

 

In other jurisdictions, we are seeing a divergent approach to regulating AI. The EU has adopted a pro-legislative approach with the EU AI Act. The UK Government, in contrast, announced in its AI Regulation Whitepaper the intention to deliver a “pro-innovation” regulatory framework, stating that it preferred not to legislate and “…avoid heavy-handed legislation which could stifle innovation and take an adaptable approach to regulating AI”, seeking instead to apply a principles-based framework for existing regulators to interpret and apply to AI. And as noted earlier in this response the ‘rights of content producers’ from the scope’ were expressly excluded from the scope. 

 

A failure to effectively regulate and legislate AI and protect the rights of individuals (including those content produces and rightsholders whose works have, and continue to be, accessed and used without permission) will undermine public confidence and will detrimentally impact the creative industries.

 

b)               To what extent does wider strategic international competition affect the way large language models should be regulated?

 

n/a

 

c)                What is the likelihood of regulatory divergence? What would be its consequences?

 

n/a

 

 

September 2023

6