Financial Timeswritten evidence (LLM0034)

 

House of Lords communications and digital committee inquiry - large language models


 

Introduction

 

  1. The Financial Times (FT) is one of the world’s leading news organisations, recognised internationally for its authority, integrity and accuracy. The FT Group, part of Nikkei Inc., provides a range of business information, news and services. It includes the Financial Times, FT Specialist and a number of services and joint ventures. The group employs more than 2300 people worldwide, including 700 journalists in 40 countries.

 

  1. The FT welcomes the opportunity to submit evidence to this inquiry. The advent and adoption of large language models (LLM) and their derivative products has the potential to substantially impact the media sector. Building a robust understanding of the likely effects and setting the appropriate regulatory guardrails is essential if we are to experience the wide-ranging benefits of LLMs without being subject to the potential harms. We consider that this inquiry presents an opportunity to advance the debate on this regulatory environment and the FT is responding with this objective in mind.

 

Capabilities and trends

 

  1. User-facing products that sit on top of LLMs, such as OpenAI’s ChatGPT and Google’s Bard, are being adopted by consumers and integrated into search platforms and wider product suites. We are not in a position to know, at this stage, how this technology will develop over the coming years and the role it will ultimately play in how information is sought and retrieved online. However, it is our view that there exists substantial potential for LLMs to erode traffic to, and engagement on, publisher-owned platforms as users substitute consumption on these sites - including FT.com and our other properties - with generative AI tools.

 

  1. If user behaviour develops in this way, this would have an impact on the capacity of media owners to monetise their intellectual property (which has been used by LLM developers without authorisation or a licence - see below). Almost all the FT’s revenue lines - including subscriptions, advertising and event sales - depend upon user engagement on FT-owned digital platforms. Engagement on our sites also fuels our ability to understand our users and in turn create products to better serve them.

 

  1. To gain a more complete view of this risk - which is essential to inform swift and proportionate policy responses - we consider it would be beneficial for Ofcom to conduct some analysis of the emerging impact of generative AI tools on media consumption habits and to put in place industry-level monitoring around this question.

 

Risks - intellectual property

 

  1. The developers of LLMs have, without authorisation, used the intellectual property of publishers as part of the corpus of data to train their systems and - unless technical measures are taken to prevent them from doing so - continue to use it to inform system outputs. The FT was not approached for prior permission to use our journalism in this way. Protected materials have therefore been used without authorisation or a licence; we believe this to be an infringement of our IP rights.

 

  1. There are legal routes to access our content which the developers of these systems have chosen not to take. The FT is in the market offering text and data mining licences to commercial partners and has been for many years. We have a track record of creating bespoke arrangements and had we been approached by LLM developers for access to our content then we would have entered commercial negotiations, setting the price and terms of a licence - covering both ingestion and usage outputs - according to the value delivered and the costs and risks to our business.

 

  1. The UK government must take steps now to support publishers to exercise their IP rights. Whilst there are ongoing legal proceedings which will establish case law in this area, this process will likely take years and much damage to publishers, and the broader creative sector, may be done in the interim. Although it is for the courts to interpret the law, a clear statement from government, articulating a policy position to the effect that developers require licences to ingest copyright protected materials for LLM training purposes would support the establishment of genuine, good faith licensing negotiations between rights-holders and LLM developers.

 

  1. We note that some LLM developers have attempted to create a distinction between freely-available and paywalled digital content; scraping materials without permission which sit outside publisher paywalls but not those which require a paid subscription. There is no legal foundation for such a distinction and a reinforcement of this point by the government would also be welcome.

 

  1. If users do, at scale, substitute the use of publisher properties with generative AI, media outlets will need to pivot their models such that a greater proportion of their revenue is derived from licensing content as an input to an end-product. The legal foundation for this must be established. Without such clarity there could be severe negative long-term implications for news publishers and the creative industries.

 

Risks competition

 

  1. Whilst the FT takes steps to keep web scraping bots off our digital properties, we have no choice but to let search crawlers access our content - both that which is in front of, and behind, the paywall - in order that the FT is indexed for, and appears within, search results. The FT, like all news publishers, is heavily dependent upon search engines for traffic. We are concerned that the largest tech businesses are able to secure an unfair advantage in the development of LLMs as a result of dominant positions held in adjacent markets, and particularly search.

 

  1. This use of data across purposes is the type of anti-competitive conduct that the new Digital Markets Unit regime is seeking to prevent. However, it is unclear whether the ‘strategic market status’ designation criteria is sufficiently broad to give the CMA powers to address these harms.

 

  1. Furthermore, as the Digital Markets, Consumers and Competition bill is still making its way through parliament, we are at least a year from the implementation of the new regime and its provisions benefiting consumers and businesses. Whilst the timeline for this much-delayed legislation cannot now be accelerated, we consider that the AI regulation white paper could have identified these competition issues, bringing them to fore and to the attention of the competition authorities.

 

Risks - misinformation and disinformation

 

  1. LLM chatbots respond to user prompts with plausible, semantically and syntactically correct outputs. However, they are predictive in nature and do not operate in such a way as to guarantee, or even check, the truthfulness of their outputs. They regularly generate false statements.

 

  1. Whilst these systems are presented by their developers as experiments, they are commercial offerings with current and future monetisation models. This technology is being integrated into search products, productivity applications, and beyond, as developers compete for share in various digital markets. We are concerned that the competitive pressure to roll-out these systems will result in consumers being exposed to misinformation. This is particularly concerning in relation to search which remains a pivotal and trusted access point to the internet for consumers. Beyond this, we are aware of, and concerned about, the risk of these systems being deliberately deployed by bad actors to generate, and disseminate, disinformation at scale.

 

Opportunities

 

  1. We expect there will be widespread adoption of AI business tools across the media sector. Promising applications exist across customer services, marketing, operations. These could lead to efficiencies and may, over time, reduce the cost base associated with certain functions. However it seems unlikely that this will result in a sustained competitive advantage being delivered to any particular firms.

 

  1. In respect of applications in the editorial function, the FT is proceeding with caution. In May this year Roula Khalaf, the FT’s editor, wrote to our readers setting out our approach. At the centre of this is the principle that ‘FT journalism in the new AI age will continue to be reported and written by humans who are the best in their fields and who are dedicated to reporting on and analysing the world as it is, accurately and fairly.’ We are aware that other media owners have adopted a less cautious approach, some of whom have had to roll-back on their use of this technology after publishing stories containing false information.

 

Domestic regulation

 

  1. The government’s AI regulation white paper places great emphasis on supporting AI innovation and investment in the UK. We are concerned that it does so at the expense of protecting the IP rights of creators. This could have a long-term corrosive effect on the sustainability of the media sector and the financial model for journalism. It could also undermine the creative industries - a sector in which the UK is an established global leader.

 

  1. The proposed approach to AI regulation is principles-based and initially non-statutory in nature. In the Secretary of State’s foreword to the white paper she outlines the risks of ‘rushing to legislate too early.. placing undue burdens on businesses’. We are concerned that due consideration has not been given to the risks of failing to legislate quickly enough. Even a narrow view of this technology taken from a news media perspective suggests that it is likely that new legislation is likely to be required to safeguard society from the potential harms of AI. For example, given generative AI outputs are not user-generated, many of the Online Safety Bill’s obligations would not apply to the developers and deployers of this technology. And yet it is already apparent that these applications have the potential to disseminate misinformation which is served in contexts that users have grown accustomed to trusting (i.e. online search results).

 

  1. If legislation is needed to address the issues arising from the development and deployment of AI, the UK government’s proposed ‘wait-and-see’ approach is likely to ultimately limit its scope to set the rules that will govern AI as other jurisdictions pass new laws and benefit from regulatory first-mover advantage.

 

  1. History tells us that speed is essential. The delay between the rise of today’s dominant digital platforms and the introduction of regulatory tools to curb anti-competitive conduct and address safety issues has resulted in poorly-functioning markets and substantial societal harms. The UK’s proposed non-statutory approach risks a similar set of circumstances emerging around AI.

 

 

August 2023