Dr Hayleigh Bosherwritten evidence (AIC0025)

 

House of Lords Communications and Digital Select Committee inquiry: AI and copyright

 

 

Dr Hayleigh Bosher is a Reader (Associate Professor) in Intellectual Property Law at Brunel, University of London. She is particularly known for her expertise in copyright law and policy in the creative industries. Hayleigh is the author of research, policy and public engagement pieces on the topic of artificial intelligence including:

 

 

 

 

To what extent is it clear under existing UK copyright law whether and when the use of protected works for training AI models is lawful? Are there particular areas (for example around text and data mining or temporary copies) where you see significant uncertainty?

 

  1. The core principle of copyright is that it requires permission to use protected works. So, when AI models utilise data sets containing copyright protected materials, whether that be books, music, images or a database, the starting position of copyright law would be that this would require permission, unless an exception applies. The restricted act of reproduction under copyright law is a broad legal concept. The rightsholder has the exclusive right to authorise or prohibit direct or indirect, temporary or permanent reproduction by any means, in any form, in whole or in part.[1] It does not matter so much how a copy is made, it is simply enough that a protected work is used without permission. Copyright is not concerned with the regulating of technology, it is intended to encourage the creation and dissemination of creativity. It regulates copying, in any material form.[2]

 

  1. The current exception for text and data mining does not apply to AI models because it is only applicable for non-commercial purposes.[3]

 

  1. Whilst some AI firms argue that in their particular system the copy is temporary (because, they assert, once a copy is tokenised the original copy is then allegedly discarded), this is a misunderstanding of the extent of the legal definition of copying. The UK copyright exception for temporary copies does not apply to a database. Moreover, it requires that the use is transient or incidental and is for the sole purpose of enabling the transmission of the work or a lawful use of the work, and has no independent economic significance.[4] Therefore, in my view, it is does not apply to the use of copyright protected content for the purposes of generative AI.

 

How would you describe the main stages or processes through which generative AI systems use data (i.e. from initial data collection, through training, to the generation of outputs)? At which points in those processes do you think copyright is most clearly engaged, and how?

 

  1. There are four main points of the AI process that engage with copyright. These are 1) the sourcing and cleansing of the training data; 2) the training of the AI model; 3) the processing of the AI model to respond to a user prompt; and 4) the AI generated output.

 

4.1.         To collate, cleans, and organise the data, the content will need to be copied by the AI developer. This was confirmed by the UK Court in the Getty Images[5] case and in the German GEMA[6] cases. Copying at this first stage engages with primary copyright infringement. In the Getty Images case this claim dropped because it was accepted that the training of the AI model took place in the US, where companion proceedings are underway in the Northern District of California.[7]

 

4.2.         The training of the AI model uses copyright protected content. Some AI systems may store a copy of the copyright content in order to do this, which would require permission, even if it is transient. This includes where the work has been ‘tokenised’ and then deleted, this could still constitute a restricted act of copyright through copying and adaptation.[8]

 

4.3.         The process of activation; when a user selects a prompt and the AI model responds, on the basis of its training, by generating content that matches the request.

 

4.4.         The generated output engages with primary copyright infringement. The test would be whether the generated image takes a substantial part of the copyright protected image. To answer this question a court would have to consider firstly, was there copying and if so, were the parts copied substantial. This question was not considered in the Getty case, as it did not form part of the complaint. The GEMA case, however, found that Chat GPT did reproduce the relevant copyright works as outputs.

 

  1. Other types of copyright infringement might also be relevant. It is a criminal offence under UK law to make for sale or hire, an article which is an infringing copy of a copyright work.[9] It is also a criminal offence to specifically design or adapt something that makes copies of a copyright work, where the designer knows, or has reason to believe, that it would be used to make infringing copies.[10] On the question of ‘reason to believe’ the Magistrate’s Court has held that broad knowledge of the items complained of, and the nature of, the infringement is sufficient.[11]

 

To what degree is it accepted that some copying of data (including temporary and intermediate copies) is unavoidable in training and operating a generative AI model?

 

  1. In my view, whether copying for the purposes of training and operating AI models is avoidable, is not a relevant legal question. In the recent adaptation of Wuthering Heights, there was unavoidable copying from Emily Brontë’s book to make the film. The question is whether a licence was needed, which in this case there was not because the copyright had expired and the work was in the public domain. If AI models are copying copyright protected work, then the law applies whether the use is avoidable or not.

 

To what degree is it a settled view that, once trained, a model does not store “copies” of individual works but instead stores numerical parameters? Does this characterisation matter for copyright analysis, particularly where models can sometimes reproduce text or images from their training data verbatim or near-verbatim?  

 

  1. In both the recent UK Getty Images decision and Germany’s GEMA case, the courts stated that training of AI models involved copying that amounted to copyright infringement. This claim was not advanced by Getty in the UK case because the training took place outside of the UK. In the German case Open AI was found guilty of infringing the copyright of songwriters in Chat GPT. The decision is relevant to UK as both laws are based on the implementation of the EU Information Society Directive.[12]

 

  1. Instead, Getty contended that the model itself was in breach of secondary copyright infringement – because the making of the model would have constituted infringement of the copyright works had it been carried out in the UK, instead of the US. This is not settled law as the decision is currently under appeal on this very point of statutory interpretation.[13] The Judge stated in her initial decision that if her statutory interpretation on this point was wrong, then copyright infringement would be found. 
  2. In its reasoning, the UK Court in Getty spent much time on the technical operation of the AI training process and model, defining how this takes place, relying heavily on the expert evidence. On the other hand, the German Court simply decided that an embodiment of the copyright protected work must have been copied in one way or another to reproduce it in the AI generated output.

 

Some stakeholders argue that training AI models on copyrighted works is analogous to humans reading and learning from those works and therefore should not be treated as infringement. From the standpoint of UK copyright law, is this analogy persuasive or misleading?

 

  1. This metaphorical device is deliberately used to influence the way in which the functioning of AI is understood. The word ‘reading’ evokes the human experience of reading, which is not a restricted act of copying, whereas the technological specific understanding of the activity is what both the German and UK Courts have found to be a reproduction by storage and temporary copying through the training process.

 

  1. In Getty the UK Court determined that the memorisation did not amount to an infringing copy within the model since it did not hold a reproduction within it in the technical terms.

 

  1. It is my view that the use of memorisation as a metaphor for the technical function of the AI training process led to a misleading understanding of facts. Memorisation is an inappropriate metaphor in the context of AI training because it evokes the act of human memorisation that does not constitute a copyright infringement. This is because copyright does not protect an idea, only the expression of the idea. If a human memorised a copyright work and then used that memory to reproduce the work, the reproduction would, in certain circumstances, constitute an infringement. In any event, an AI training model is not equivalent to the human brain. Firstly, on the basis that the human brain would not be capable of memorising the huge data sets that AI models ingest. Secondly, when humans create, they may draw inspiration from previous works, but they are also able to draw from themselves and their life experiences. AI models reproduce from the source data alone.

 

  1. Suggesting that AI memory is like human memory is to subtly, but significantly, project the norms surrounding the regulation of human memory, which is of course not regulated, onto AI technology. This in turn treats the AI memory as epiphenomenal; largely divorced from the actual structure and function of the technology. The same logic applies to the use of the term ‘reading’.

 

  1. Whilst terms such as reading, memorisation, and hallucination might be accepted parlance in the AI sector, it should be avoided in the context of policy and regulation because it is inaccurate and misleading metaphorical rhetoric which inappropriately infers that an AI model should be regulated the same as a human brain.

 

Within UK law, how far can existing exceptions realistically be interpreted to cover large-scale training and inference-time retrieval?

 

  1. If any current copyright exception applied already, the UK Government would not have considered implementing a new copyright exception to cover AI activity. Over 11,500 responses to the consultation were received, the introduction of an exception for all text and data mining purposes with rights reservation was supported by only 3% of respondents; and the introduction of an exception to copyright for all text and data mining purposes with no rights reservation was supported by only 0.5% of respondents.[14]

 

  1. Copyright exceptions are intended to be narrow. Article 9(2) of the Berne Convention states that exceptions must be 1) limited to certain special cases 2) not conflict with the rightsholders normal exploitation of the work; and 3) not unreasonably prejudices the legitimate interests of the rightsholder.

 

  1. In the Getty case, before the primary infringement claims were dropped, Stability AI did not put forward any arguments that copyright exceptions applied to the training of the data, unlike in the companion case in the US, their defence is based on an argument of Fair Use, which does not apply under UK law. Their next line of defence was that if there was liability it should fall on the user and not them. The Court decision stated, however, that if liability was to be found it would fall on the AI developer, in this case Stability AI.

 

Why are outputs “in the style” of a named creator and digital replicas difficult to address under the current UK copyright regime?

 

  1. To bring a copyright infringement claim, the victim of an unauthorised deepfake must be the copyright holder.[15] The rightsholder is usually the person who took the photo or made the video.[16] This means that a copyright claim is only viable where the images or videos used in the deepfake were taken by the victim. Moreover, enforcing this would be a challenge since there are currently no requirements for transparency in the use of copyright protected materials in the training or processing of AI.[17]

 

  1. Another challenge is overcoming the test for infringement. Primary copyright infringement occurs when the whole, or a substantial part, of a copyright-protected work is copied without permission, or the benefit of a copyright exception. It seems probable that, under the circumstances where the deepfake is intended to mimic the person from the original image, there would be a substantial similarity between the two. However, it is not without obstacles. The test for substantial part relates to the quality, and not the quantity, of the parts taken. In this context, quality refers to the reason the work is protected by copyright, its originality.  It follows that a claim for copyright infringement can be rebutted where the defendant can demonstrate that the work copied is unoriginal and consequently unprotectable. In these circumstances, even where there was copying, it would not amount to copyright infringement.

 

  1. In relation to digital replicas that mimic style, the challenge is that copyright does not protect style. This is called the idea/expression dichotomy and is an important aspect of copyright which means that ideas, such as a theme are not monopolised; instead, only the individual creative expression is protected. This means, for example, the idea for a song that celebrates birthdays is unprotectable, but Happy Birthday by Stevie Wonder, Happy Birthday by Kygo featuring John Legend, Birthday by Anne-Marie, Birthday by Will.i.am featuring Cody Wise, Birthday by JP Cooper, and Birthday by The Beatles all enjoy copyright protection.

 

In the UK and in other jurisdictions, what is usually meant by “image rights” or “personality rights” and how do they interact with copyright and neighbouring rights?

 

  1. The UK does not currently have image rights or personality rights. Before deepfakes, similar situations would call on laws such as passing off. For example, this was used when the retailer Top Shop sold a T-shirt with singer Rhianna’s face on without permission.[18] (Passing off applies where a claimant can demonstrate goodwill, misrepresentation, and damage to goodwill.[19]  However, it is considered notoriously difficult to achieve and only applies in narrow circumstances.[20] It would also only apply to those trading from their image.)

 

  1. In other jurisdictions image rights, and publicity rights, are typically reserved for celebrities and those trading from their name, voice and likeness.

 

  1. Guernsey protects personality rights, including voice, likeness, appearance, silhouette, face and mannerisms. It requires registration, which grants a property right that can be protected, licensed and assigned. There are limitations to this right, such as fair dealing for the purposes of research, news reporting, of the arts, or for educational purposes.[21]

 

  1. Most US States recognise a right of publicity or image rights to address the use of an individual’s persona in commercial contexts.[22] Typically, these only apply in circumstances of advertising, merchandise, or for commercial purposes. The scope of the rights differs significantly between States, and many require commercial value in the individual’s identity. The US Copyright Office advocated that new federal legislation is urgently needed to address the harms that can be inflicted by non-commercial uses, proposing a new right that extends protection to all individuals regardless of their commercial value of their identities.[23]

 

  1. In my research[24] I argue that a personality right should be introduced into UK law to protect individuals from unauthorised digital replicas. This is preferable to deepfake specific legislation which will quickly become obsolete in light of technological developments. In this instance, the law should focus on what it is trying to protect, instead of what it is trying to protect against.

 

  1. Some argue personality rights could be a threat to creative and free expression.[25] Researchers also note the importance of not restricting biographical and historical information, such as for the purpose of creating a biopic film which requires a replica of an individual.[26] The new statutory right must be balanced with freedom of expression, which can be achieved through limitations, just as copyright provides the right to restrict the unauthorised use of a work[27] with specific exceptions.[28]

 

  1. Another concern is that individuals, particularly creators, could be pressured into signing away any personality right to a dominant player, due to inequality of bargaining power in the industry.[29] However, copyright is structured as a balancing tool between stakeholders for this precise purpose.[30] Utilising copyright principles, the law can enable the rightsholder to receive a reward that corresponds to a fair share of the revenue generated by the exploitation of their work.[31] Rights can be unwaivable and a statutory reward can be required in exchange for any transfer of rights.[32]

 

  1. Any legislative intervention must be centred around the consent of the victim, not the intent of the perpetrator, which can vary and be difficult to prove. Likewise, it must encapsulate all relevant parties, regulating those that generate the deepfake, those who possess or share it, as well as platforms, websites and applications that facilitate, generate, host or store deepfakes.

 

  1. Additionally, the harms of deepfakes permeate across industries, exacerbate systemic issues of sexism, racism and misogyny whilst escalating challenges of misinformation and threats to the creative industries, conceptualised as ‘an assemblage of differential tensions, comprised of interlinking and overlapping ruptures and continuities.’[33] The UK must adopt a holistic approach that includes legislative and technological interventions as well as educational and cultural actions that address the systemic issues that underpin and propel deepfake harms.

 

 

February 2026

8


 


[1] Copyright, Designs and Patents Act 1988, section 17 (hereafter CDPA 1988).

[2] For further analysis see; Bosher H., ‘Copyright Infringement in AI Models: Statutory interpretation, Memorisation and Metaphors in Getty Images and GEMA’ (2026) 48(4) European Intellectual Property Review, pp. 1-26.

[3] CDPA 1988, section 29A.

[4] CDPA 1988, section 28A.

[5] Getty Images v Stability AI [2025] EWHC 2863 (Ch).

[6] GEMA v Open AI, Case No. 42 O 14139/24.

[7] Getty’s US complaint focuses on copyright infringement, trade mark infringement, breach of its terms and conditions and circumvention: Getty Images (US), Inc. v. Stability AI, Inc. 1:23-cv-00135-UNA.

[8] Gervais D., Marmanis H., Shemtov N., and Rowland C. Z., ‘The Heart of the Matter: Copyright, AI Training, and LLMs (21 September 2024).

[9] CDPA 1988, s107(1)(e).

[10] CDPA 1988, s 107 (2).

[11] Kinnier-Wilson J., ‘Copyright: Criminal Copyright Infringement – Sections 107 and 110 CPDA 1988 – Copyright in a Photocopy’ (1994) 16(11) E.I.P.R., 217-248, 295.

[12] Directive 2001/29/EC.

[13] Getty Images (US) Inc & Ors v Stability AI Ltd (Re Form of Order) [2025] EWHC 3343 (Ch) (16 December 2025).

[14] 88% expressed support for option 1 - require licences in all cases. Making no changes to copyright law (option 0) was supported by 7% of respondents. https://www.gov.uk/government/publications/copyright-and-artificial-intelligence-progress-report/copyright-and-artificial-intelligence-statement-of-progress-under-section-137-data-use-and-access-act

[15] CDPA 1988, s96(1).

[16] CDPA 1988, s11(1).

[17] UK IPO, Copyright and Artificial Intelligence (UK IPO, 17 December 2024).

[18] Robyn Rihanna Fenty, Roraj Trade LLC, Combermere Entertainment Properties LLC v Arcadia Group Brands Limited (T/A Topshop), Top Shop/Top Man Limited [2013] EWHC 2310 (Ch) 55.

[19] Reckitt & Colman Products Ltd v Borden Inc (Jif Lemon) [1990] 1 WLR 481 HL.

[20] Emmanuel Oke, ‘Image Rights and Passing Off: Should Reputation be Enough for Celebrities to Succeed in English Courts?’ (2020) 15(1) JIPLP 49-54.

[21] Guernsey protects personality rights, under the Image Rights (Bailiwick of Guernsey) Ordinance 2012, s31.

[22] Thomas Mccarthy J and Roger E Schechter, The Rights of Publicity and Privacy (Thomson Reuters 2024) 5:62.

[23] US Copyright Office, ‘Copyright and Artificial Intelligence, Part 1: Digital Replicas’ (US Copyright Office July 2024) 29.

[24] ‘Do Deepfakes, Digital Replicas and Human Digital Twins Justify Personality Rights?’ (2026) Journal of World Intellectual Property, pp. 1-47 in press.

[25] E.g., Google Response to UK IPO AI Consultation, 13 <https://storage.googleapis.com/gweb-uniblog-publish-prod/documents/Google_response_to_UK_Copyright__AI_Consultation_February_2025_hLpZUuW.pdf> accessed 14 August 2025.

[26] CREATe, Copyright and AI: Response by the CREATe Centre to the UK Government’s Consultation (CREATe Working Paper 2025/2) 41.

[27] CDPA 1988, s16.

[28] CDPA 1988, Chapter III.

[29] CREATe, Copyright and AI: Response by the CREATe Centre to the UK Government’s Consultation (CREATe Working Paper 2025/2) 42.

[30] Abraham Drassinower, ‘From Distribution to Dialogue: Remarks on the Concept of Balance in Copyright Law’ (2009) 34 J. Corp. L. 991-1007, 992.

[31] Laddie, Prescott and Vitoria, The Modern Law of Copyright and Designs (Butterworths 2011) 2163.

[32] Hayleigh Bosher, ‘The UK Economics of Music Streaming Inquiry’ (2022) 33(2) Ent.L.Rev. 50-56.

[33] Benjamin Jacobsen and Jill Simpson, ‘The Tensions of Deepfakes’ (2023) 27(6) iCS 1095–1109, 1106.