DMG Media—written evidence (LLM0068)
House of Lords Communications and Digital Select Committee inquiry: Large language models
- This response is made on behalf of DMG Media, publishers of the Daily Mail, Mail on Sunday, MailOnline, Metro and metro.co.uk and, through its sister company Harmsworth Media, the i, inews and New Scientist.
- We welcome the Communication and Digital Committee’s decision to examine the many difficult questions raised by the rapid development of large language models. Committee members will be familiar with the evidence we have presented to previous inquiries about the damage the unregulated growth of online platforms has caused to news publishers over the last 25 years.
Executive Summary
- The commercial exploitation of large language models represents a paradigm shift in how news content is likely to be distributed in the future, and who will reap the rewards. It is clear from the numerous inquiries dating back to the Cairncross Review in 2019, that the expropriation of news content by online platforms has left the news publishing industry in a parlous state. AI could make it impossible to produce independent, commercially funded journalism at all. The key points examined in this submission are:
- We now know that our content and that of other news publishers has been used to train large language models since at least 2015. In Google’s case, to begin with it appeared to be purely academic research (which, of course, can benefit from the relevant IP copyright exemption). It is only in the last year that it has become clear that the purpose was commercial.
- In addition to our content being part of internet-wide datasets such as Common Crawl, we know that Google made specific use of news content via the CNN/DailyMail dataset (CNN/DM), 70 per cent of which is MailOnline content. Our content was used to test the effectiveness of AI training, because of its uniform headline/bullet points/article text structure. Key words were removed from the bullet points and the AI model was tested by being asked to replace them from its reading of the text. OpenAI have also made use of our content. We know they trained their models on the Reddit TL;DR dataset, then tested them with CNN/DM to check their effectiveness.
- All of this work was done without our permission or any payment. We are actively seeking advice on potential legal action.
- There is also the question of the continuing use of our content going forward. Unlike traditional search, which answers search requests by presenting links which sometimes (though not always) deliver users to news publishers’ content, AI search gives what appears to be a complete response. In Google Bard’s case there are currently no links at all, so no revenue will return. Bing/OpenAI has links, but they are presented as footnotes, and one would expect the click-through rate to be very low. Quality news publishing is already seriously threatened by traditional search and social media, AI-powered search could destroy it altogether.
- For AI search to be commercially successful it will have to be constantly updated with fresh information, for which news sites are the obvious source. At the moment this is being done without permission or payment. That must end – establishing proper terms for copyright consent and payment should be a first priority for the Digital Markets Unit (DMU) once the Digital Markets Competition and Consumer (DMCC) Bill is passed.
- Even more serious than that, AI threatens to pollute society’s collective pool of knowledge altogether. Because AI works on establishing the probability of one word following another, rather than actually comprehending language as a human does, over time it removes all data which does not match the norm. When this happens on a large scale it causes large language models to collapse. ‘Pure’ datasets with no AI content will have to be preserved or created to prevent AI becoming a dog that eats its own tail. There is also emerging evidence that AI models exhibit political bias.
- Finally, there is the question of legal/regulatory responsibility for AI content. At the moment legal responsibility for news content presented in traditional search and social media lies with the news publisher, and platforms enjoy immunity. Can that continue when news content is generated by an LLM drawing on numerous sources and quite possibly reproducing none of them accurately nor in line with legal requirements?
Background – how digital platforms have cannibalised news publishers since the turn of the millennium
- For more than two decades digital platforms have been able to expropriate our intellectual property and exploit it in competition with our own websites without restraint. They have neither had to bear the cost of creating news content, nor, thanks to Section 230 of the US Communications Decency Act and equivalent legislation elsewhere, have they carried the cost of holding legal responsibility for the content they host. Until very recently the only regulation they have faced in the EU and UK has concerned user privacy and the use of cookies – which they have successfully exploited as a lever to reinforce their existing market dominance.
- It is true that the internet has given great opportunities to news publishers. In DMG Media’s case, we went from being a publisher of newspapers in one market, the UK, to running a global news site with full-scale editorial and commercial operations in the USA and Australia, as well as Britain. The business premise was that the huge numbers of new users supplied by the platforms would bring with them the advertising revenue that would fund this expansion – indeed, that is the justification given by the platforms for using news content without payment.
- However, economies of scale, very large upfront investments, minimal costs of adding new users, incumbent data advantages and network effects caused digital markets to tip with extraordinary speed in favour of single dominant players. These dominant players – Google, Facebook and Apple – exercise vastly more market power than news publishers, which unavoidably have had to become their business partners.
- Where payment for content has been offered, it has been on take-it-or leave-it terms, which represent a fraction of the revenue previously earned from print operations. Furthermore, when it has come to generating the advertising revenue all those millions of news users were supposed to deliver, the same platforms which dominate search and social media have leveraged that dominance into digital advertising markets, where self-preferencing by those platforms has sharply reduced the proportion of advertising spend which reaches publishers.
- The result has been a major disparity between the value news publishers supply to the digital platforms and the revenue platforms return to publishers. Matt Elliott, Professor of Economics at Cambridge University, has calculated the value of news in ensuring information supplied by platforms to their users is accurate, relevant and constantly refreshed is £1 billion a year to Google and Facebook in the UK alone.[1] In contrast Professor Annabelle Gawer of Surrey University has calculated that the revenue returned by Google to UK news publishers is less than £75 million a year.[2] The consequence has been that all news publishers – especially smaller, regional businesses – have struggled to turn hugely increased digital user numbers into sustainable revenue models.
- Encouraged by the powerful advocacy of the Communications and Digital Committee, the Digital Markets, Competition and Consumers Bill now being debated by Parliament should address those problems. It will give the Competition and Markets Authority’s Digital Markets Unit statutory powers to impose codes of conduct which should prevent advertising market-rigging and self-preferencing, and force platforms to agree fair, reasonable and non-discriminatory terms of payment for their use of news content.
- Unlike its more rigid European counterpart, the Digital Markets Act, the DMCC Bill has been deliberately drafted in a way which should allow the DMU to respond rapidly to changes in digital markets. Large language models and generative AI are almost certain to provide it with its first major challenge.
How large language models use news content
- News content plays a central role in the development and exploitation of AI. Large language models (LLMs) are a form of generative AI that are trained on large datasets; LLMs generate text and can be used to summarise, translate, answer questions and draft text-based content.
- In essence, to have any use, generative AI must mimic the work of a journalist – it must absorb large quantities of data, analyse its significance, and then present its conclusions to the user in simple but accurate language. To do this commercially, it must be able to carry out these tasks consistently and reliably – much has been made of the tendency of early AI models to ‘hallucinate’ - i.e. present obviously nonsensical answers when confronted with a user request it apparently does not properly comprehend. Of course, in reality the LLM does not ‘comprehend’ anything in the true sense, its responses are purely based on probability. While some responses may be clear hallucinations, the potentially more dangerous ones are those that appear plausible to the human user but in fact contain misinformation. The use of generative AI in search, one of the earliest commercial uses of generative AI, has highlighted the importance of accurate, sensible responses.
- A survey of academic papers has shown that MailOnline content, along with that of other major news publishers, forms a significant part of popular LLM training datasets, such as Common Crawl and WebText.
- Common Crawl consists of webpages from millions of web domains. Most modern LLMs, including OpenAI’s GPT-3/3.5, BloombergGPT and Google’s LaMDA used Common Crawl or its subsets for training. 540,000 MailOnline webpages were included in Common Crawl in June-August 2018, amounting to ~0.007% of all Common Crawl content for those months. In preparing Common Crawl to use for model training, it is typically filtered to reduce low-quality web content and increase the prevalence of high-quality web content, such as news.
- WebText consists of webpages mentioned in top-ranked Reddit posts, and was used to train OpenAI’s GPT-2, GPT-3, BloombergGPT and other models. MailOnline is the 24th top domain included in WebText (~119,000 webpages or 0.3% of the total).
- The following chart illustrates the use of our content and that of other major news publishers by Common Crawl and WebText, ranked by number of pages included in Common Crawl:

- However, to ensure reliability, generative AI must not only be trained but also be tested, and academic papers show that this is where MailOnline news content has played an even more significant role.
- To give an example, Teaching Machines to Read and Comprehend,[3] a paper by Hermann and others for Google DeepMind and Oxford University, describes how a dataset of around a million news stories purloined from the CNN and MailOnline websites was used to train machine reading systems and test their ability to answer questions posed on the contents of documents they have seen.
- MailOnline news stories were chosen because their consistent structure – headline, followed by bullet points summarising key points, then the full text of the story – lent itself to the task of testing. The LLM’s ability to accurately read a news story was tested by removing keywords from the bullet points, and then asking it to replace them from its understanding of the full text. MailOnline content makes up 70% of the CNN/Daily Mail dataset.
- Google’s own blogs confirm how the CNN/DailyMail dataset is used in training its AI models for the task if summarization – see the reference to ‘CNN/DM’ in the first column under ‘Summarization’ in the Google diagram below:[4]

- OpenAI has until recently been more circumspect about how its LLM was trained. However, a recent OpenAI blog revealed that the CNN/DailyMail dataset played a key role here as well.[5] Although OpenAI LLMs see news content during pre-training (the blog does not reveal which dataset was used at this stage), fine-tuning is done using the Reddit TL;DR dataset, which consists of posts from the Reddit social media site. This fine-tuning is then checked by asking the model to produce summarisations of content from the CNN/DailyMail dataset. OpenAI was very happy to report that its model generated excellent summaries without any further training.
- In all these cases our content was used to train and test AI without permission being sought or any licence agreed. Of course, in the near term, that is a potential issue for the courts. A number of major media businesses have already launched legal actions over the unauthorised use of their content in generative AI, including Getty Images[6] and Thomson Reuters.[7] Others are actively considering legal action: these include the New York Times[8] and DMG Media.
- The issue that must concern legislators and regulator is this: now that machines have been trained to absorb information gathered by other parties, and organise, present and monetise it as a news publisher would, what effect will this have on the incentive for media organizations to continue to create and publish high quality, reliable news content – and what will be the impact on society and democracy should news organizations be forced to reduce investment or cease to operate.
Why news content will remain crucial to the operation of LLMs going forward
- As the Government has acknowledged with the creation of the Digital Markets Unit, the first iteration of digital search already presents a serious threat to a competitive and pluralistic market in news. In theory, search should help competition by allowing all providers of news to reach far larger audiences without the heavy costs of printing and distributing a newsprint product. In return the search engine provides links which take users back to news publishers’ websites where their investment in journalism is rewarded by advertising clicks.
- In practice the overwhelming dominance of one search engine – Google – has placed it in a position where it can pick and choose which news sites it promotes and which it does not, and the terms of compensation, if any, it offers to news providers. Many users read the link but do not click through, generating no revenue for the news provider. Those who do click through may generate ad revenue, but as Google also owns most of the digital advertising ecosystem, and runs it for its own benefit, a disproportionate share of that revenue goes to Google.
- To put it another way, Google is already dominant in search, which is one of the main means of distributing news content, and in digital advertising, which is the main way of funding it online. Generative AI threatens to allow it to control content creation as well.
- Generative AI is in its very early days, and the long term associated business models and monetization required to fund it are not yet clear. It is even less clear how generative AI will return any value to news publishers. What is clear is that it is potentially a very serious threat to news organisations.
- From news content in their training and testing data, such as Common Crawl and the CNN/Daily Mail dataset, LLMs gain the ability to synthesize information and to ‘comprehend’ and generate text. The extent to which LLMs learn to imitate news content is described in an OpenAI research paper, which finds that “GPT-3 can generate synthetic news articles which human evaluators have difficulty distinguishing from human-generated articles”.[9] Clearly this threatens to reduce traffic to news publishers’ websites.
- However, this is not the only way applications using LLMs benefit from continuously updated news content online. The accuracy and reliability of generative AI output has been central to the discussion around its release and adoption. News provides high quality, continuously updated information, which is hugely important for the commercial exploitation of generative AI. Unless LLMs can incorporate current news information on an ongoing basis they will be unable to generate responses that refer to any events that occurred after the training dataset was created. If generative AI is going to be a key component of search, and a viable tool for users drafting documents, it is clear that the continuous ingestion of up-to-date news will be vital for its success.
- There is also worrying evidence that the more LLMs are trained on AI-generated data, the more likely they are to forget their original human-generated data, leading to catastrophic collapse. Again, the reason for this is that LLMs do not comprehend text as a human does; they merely calculate the probability of one word, or group of words, following another. This means, to take a hypothetical example, that if 90 per cent of the world’s cats are yellow and ten per cent are blue, an LLM will deduce that the normal colour of a cat is yellow. After a while it will start to depict blue cats as green, before eventually presenting them all as yellow.[10] This is the opposite to journalistic technique, which is to look for the exceptions to the norm, because that is where news is found. LLMs’ inability to cope with rare events means they rapidly forget the true distribution of information. Not only does this have worrying implications for minorities of all sorts, but it means LLMs will have to be continuously retrained on fresh, human-generated datasets.
- However, this may be increasingly difficult to maintain. Recently there has been clear evidence that LLM operators recognise they cannot continue to use huge quantities of copyright news content without paying for it and as a consequence there are moves to allow publishers to block the web crawlers employed to find and ingest it. When OpenAI recently revealed its GPTBot web crawler for ChatGPT it also announced it would respect robots.txt, the industry-wide standard tool for preventing crawlers from extracting data. It has been estimated that 10 per cent of the world top 1000 websites, including Amazon, Reuters and the New York Times, have already blocked GPTBot.[11] The Guardian has announced it is following suit,[12] and DMG Media has the issue under active consideration.
- The position with Google is more complex. While it has acknowledged the need for a debate around how publishers can control the use of their content,[13] at the moment it is extremely difficult for publishers to block Google crawlers. This is because Google operates a number of crawlers, and does not make it clear to what extent they feed search or AI, or both. Although search referral traffic is less valuable than direct traffic, it is still a significant contributor to publisher revenue, especially for smaller publishers. Therefore, very few publishers have as yet blocked Google’s crawlers. This in turn raises the possibility that Google could use its dominance in search to lever dominance over OpenAI in generative AI.
- And of course blocking only offers a solution to future unauthorised ingestion of news content. Although in time they are likely to be superseded by new models, today’s LLMs can’t unlearn the data on which they have already been trained.
- LLMs therefore present a disturbing conundrum: an LLM must have fresh supplies of human-generated news content if it is not to become a dog which eats its own tail, but where will that content come from if search responses generated by the LLM prevent news publishers from funding genuine, human-generated journalism?
How will operators of LLMs recompense publishers for their use of news content?
- There can be no doubt that LLMs can now produce content that in certain cases is not easily distinguishable from professional journalism. The prospect of competing with AI-generated content, originating from models trained using news content without compensation and produced at low marginal cost, threatens a grim future for news publishers.
- This risk is abundantly clear in generative AI responses which are complete in themselves and therefore compete directly with news websites, without necessarily providing any click-through.
- To give an example, below is the answer provided by New Bing, currently the most readily available generative AI search tool, to the request ‘What is the IMF’s latest forecast for the British economy?’:
The IMF has delivered a major U-turn on its forecast for the UK economy, saying it expected growth this year just one month after predicting a contraction. Britain's economy is expected to expand 0.4 percent in 2023, the International Monetary Fund said in its latest outlook document which partly cited weaker energy prices². The UK economy is expected to avoid a recession this year, the International Monetary Fund has said, after it sharply upgraded its growth forecast³.
- Beneath the answer are links to five news stories, content from which is presented in the main body of response without headlines or urls:
(1) IMF Forecasts UK Economic Growth In Major U-turn. https://www.ibtimes.com/imf-forecasts-uk-economic-growth-major-u-turn-3695137.
(2) IMF expects UK economy to avoid recession. https://www.bbc.co.uk/news/business-65669399?at_campaign_type=owned&at_campaign=Social_Flow&at_link_origin=BBCPolitics&at_format=link&at_bbc_team=editorial&at_link_type=web_link&at_link_id=9AA38FE6-F971-11ED-9CD1-BE96AD7C7D13&at_ptr_name=twitter&at_medium=social.
(3) British government welcomes IMF's U-turn on forecast for UK economy. https://www.timeturk.com/en/british-government-welcomes-imf-s-u-turn-on-forecast-for-uk-economy/news-75934.
(4) IMF dramatically upgrades outlook for UK economy - but says more cost of living pain to come. https://news.sky.com/story/imf-forecasts-dramatic-upgrade-for-uks-economic-outlook-but-more-cost-of-living-pain-to-come-12887402.
(5) IMF downgrades forecast for UK economic recovery. https://www.theguardian.com/business/2021/jan/26/imf-downgrades-forecast-for-uk-economic-recovery
- The choice of Timeturk, a Turkish news site with an English language edition, and a two-year-old article from the Guardian may suggest New Bing’s AI still needs some refinement. However, the answer given was accurate, easy to read, and probably contains enough information for the average reader. It is doubtful many would click on the links, particularly as the first one is to a US-headquartered news site with little profile in the UK.
- Therefore, while it happily uses journalism provided by the BBC and Sky News, it is very unlikely to return any significant traffic to them nor, in the case of Sky, any ad revenue.
- Similarly, on Google Bard – Google’s generative AI chat tool - the following response appeared in response to “What are the top financial news headlines today?”:
Sure, here are the top financial news headlines for today, June 1, 2023 in the UK:
- UK inflation hits 9%, highest level in 40 years.
- Sterling falls to two-year low against the dollar.
- Bank of England raises interest rates to 1.25%.
- House prices in the UK fall for the first time in 10 months.
- UK government borrowing hits record high.
UK inflation hits 9%, highest level in 40 years.
The UK's Consumer Price Index (CPI) rose to 9% in April, the highest level since 1982. The rise in inflation was driven by higher energy and food prices.
Sterling falls to two-year low against the dollar.
The value of the pound sterling fell to a two-year low against the US dollar on Thursday. The fall in the pound was due to a combination of factors, including the UK's higher inflation rate and the Bank of England's decision to raise interest rates.
Bank of England raises interest rates to 1.25%.
The Bank of England raised interest rates by 0.25 percentage points to 1.25% on Thursday. The rate hike was the fifth consecutive increase and the highest level since 2009. The Bank of England is trying to bring inflation under control, but it is also aware that higher interest rates could slow economic growth.
House prices in the UK fall for the first time in 10 months.
House prices in the UK fell by 0.3% in April, the first time they have fallen since June 2021. The fall in house prices was due to a combination of factors, including the rising cost of living and the Bank of England's decision to raise interest rates.
UK government borrowing hits record high.
The UK government's borrowing hit a record high of £232.5 billion in the financial year to March 2023. The rise in borrowing was due to the COVID-19 pandemic and the war in Ukraine.
In this Bard response, news organisations (clearly the source of the provided information) are neither credited nor linked to, so there is no prospect of any value being returned to the news publisher, nor does the user have any way of assessing the reliability or political affiliations of the source.
- One can also see the first clues as to how generative AI models will be funded. The opportunities are there – New Bing’s answer to the query “best shop for work shirts” is shown below; the first link shown when the user hovers over the response is an ad, linking to a clothing retailer.


- If generative AI is to be funded by selling links in answers to search queries, the trust that users place in the search responses must be high; the business model of search advertising has relied upon this user perception of reliable, high-quality results for decades. As in traditional search, news content will be of critical importance in developing user trust in the reliability and accuracy of generative AI search responses, and thus in the search platform’s ability to monetize such responses. And yet the shorter synthesized response on chat search as opposed to traditional search also raises the possibility that as AI develops, news publishers may be given no credits and therefore receive no referrals unless they pay for them.
- Unless the reforms promised by the DMCC Bill dramatically increase digital advertising revenue and ensure terms of payment for news content which properly reflect the cost of producing it, supplying content to AI platforms is highly unlikely to prove a viable business model for news publishers.
- We must therefore be aware that AI could destroy the economic foundation of journalism altogether, unless journalism is publicly funded. That is why it is so important that the competition implications of AI are addressed now, and not 25 years down the line, as is finally happening in the case of earlier versions of search, social media, and digital adtech.
AI is currently a competitive business. Can that be maintained, and what would be the benefits if it is?
- In these very early days, generative AI does not yet have a dominant player. Google is by all accounts the biggest investor in AI, until very recently running two separate AI research teams: Brain and London-based Deepmind, bought for $500m in 2014.[14] But the first AI search product available to the public was ChatGPT, developed by start-up OpenAI. Earlier this year Microsoft announced a £10 billion partnership with OpenAI,[15] giving its GPT-4 powered AI search product New Bing an edge over Google in the battle for AI users.[16]
- There are already signs that the existence of a viable competitor to Google in AI search is opening up more competition than existed before. Samsung has considered switching from Google to Bing as the default search option on its smartphones. Eventually it decided to stay with Google,[17] at least for the time being, but the fact that the review took place demonstrates that Google can no longer take market dominance for granted.
- Meanwhile OpenAI has shown the possibility of a very different approach to content creators. Google steadfastly refused to reward publishers for news content in traditional search until driven to it by the threat of legislation, first in Australia, then elsewhere. In contrast, OpenAI CEO Sam Altman said at a White House event in May this year that his company’s AI would respect copyright and was working to ensure content creators are rewarded.[18] Microsoft’s CEO Satya Nadella has also spoken about sharing ad revenue with publishers.[19]
- It remains to be seen how this might be achieved, but it cannot be any coincidence that when there are two players in a marketplace, the possibility of offering fair and reasonable terms to secure good relationships with third-party suppliers suddenly starts to make commercial sense, rather than the wholescale appropriation of copyright content which has been the practice in digital markets dominated by single players. We would very much welcome further entrants and more competition in the training and provision of generative AI models.
- This is why we believe it is very important that legislators and regulators consider the effects of LLMs on competition from the very beginning. It is vital that a competitive marketplace in generative AI is maintained, and tipping does not take place.
- We are fortunate in the UK that, after five years of intensive work, we have a regulator, the CMA, which has grasped the dynamics of digital markets, and Parliament is on the brink of giving statutory powers to the DMU, which has the flexibility to protect and succour competition as new markets are born and grow. Generative AI should be a No 1 priority for the DMU. The most effective way to achieve this would be for the CMA to launch a Market Investigation, which would give it the legal power to demand the internal analyses it needs to fully understand how AI will function and the effect it will have on competition in digital marketplaces.
Copyright, legal responsibility and media plurality – other AI risks which legislators should consider
- Great care must also be taken to preserve intellectual property rights. If, as appears to be the case, LLMs will collapse without fresh supplies of human-generated content, then financial incentives to produce and publish such content, including news content, must be maintained. Yet for some years digital players have been campaigning for an exemption to copyright law. They very nearly succeeded when the Intellectual Property Office proposed to introduce a new copyright and database exception which would allow text and data mining for any purpose, including commercial exploitation. An exception in place since 2014 already allows text and data mining for non-commercial purposes, which is doubtless why early work on LLM was presented as academic research, although it was funded by digital platforms.
- However, in February this year the Government announced[20] that it would not be proceeding with this proposal, and we believe the stage is now set for a much broader inter-disciplinary approach, in which competition, and therefore the CMA and DMU, must play a central role.
- Another question which has so far barely been addressed is who carries legal responsibility for AI generated news content? Owners of LLMs appear ready to claim copyright in their AI output to prevent it being used to train rival LLMs.[21] But can a business claim ownership of content without also accepting legal responsibility for it?
- At present digital platforms enjoy immunity from libel and privacy law under section 230 of the US Communications Decency Act, which is mirrored by similar legislation around the world. Legal responsibility lies with publishers to whom users are directed by the links presented in traditional search and social media. But if someone is libelled or their privacy invaded in an AI-generated response, who can they sue? In the case of Bard they would be unable to identify any source of the information other than Bard. With OpenAI there might be a source identified in a footnote, but OpenAI may well take information which is not libellous then combine it with information from other sources in a way which is libellous. The same would apply to factual accuracy and the host of other issues covered by IPSO and Ofcom regulation. Who will hold the ring?
- A worrying example of the problem was uncovered recently when and American law professor revealed that Chat GPT had falsely accused him of sex attacks on students during a trip to Alaska which never took place, while he was supposedly employed in a university department where he had never worked. ChatGPT cited as its source a Washington Post article which did not exist.[22]
- Finally, academic research is beginning to raise questions about potential threats LLMs pose to media plurality. Recently, widely-publicised[23] research by a team led by Dr Fabio Motoki at the University of East Anglia asked ChatGPT to answer questions about political beliefs in the USA, UK and Brazil, first in the persona of a left-leaning person, then in the persona of a right-leaning individual. The responses given were then compared with Chat GPT’s default answers. Their report concluded: ‘We find robust evidence that ChatGPT presents a significant and systematic political bias toward the Democrats in the US, Lula in Brazil, and the Labour Party in the UK.’[24]
- The researchers came to no firm conclusions as to what caused this apparent bias. They noted that ‘OpenAI declares it cleans the CommonCrawl dataset and adds information to it. Although the cleaning procedure is reasonably clear and apparently neutral, the selection of the added information is not. Therefore, there are two non-exclusive possibilities: (1) the original training dataset has biases and the cleaning procedure does not remove them, and (2) GPT-3 creators incorporate their own biases via the added information.’ Whatever the reason, if it is true that LLMs are displaying consistent but undeclared political bias, this should be a matter of deep concern. Given the expected uses of generative AI in assisting users in drafting reports, proposals and university essays, there is a serious risk that society comes to accept a Californian millennial worldview as the global norm, and regards any other point of view, however legitimate, as an aberration. This should be deeply worrying for those believe media plurality is the cornerstone of democracy.
- For all the reasons outlined in this submission it is absolutely essential that legislators and regulators work together to examine the many unaddressed questions surrounding AI, and not look at it solely from the point of view of one regulatory discipline, whether it is competition, copyright, or libel and privacy law. We recommend that the Committee calls on the Digital Markets Unit to work with the Ofcom, the Information Commissioner and the Intellectual Property Office to understand the training and function of LLMs and draw up conduct requirements as a matter of urgency.
September 2023
15
[1] https://newsmediauk.org/wp-content/uploads/2022/10/Value_of_UK_News_to_Digital_Platforms_-_Final.pdf
[2] https://newsmediauk.org/wp-content/uploads/2023/05/Gawer-Article_FINAL.pdf
[3] https://proceedings.neurips.cc/paper_files/paper/2015/file/ afdec7005cc9f14302cd0474fd0f3c96-Paper.pdf
[4] Introducing FLAN: More generalizable Language Models with Instruction Fine-Tuning – Google AI Blog (googleblog.com)
[5] https://openai.com/research/learning-to-summarize-with-human-feedback
[6] https://www.reuters.com/legal/getty-images-lawsuit-says-stability-ai-misused-photos-train-ai-2023-02-06/
[7] https://today.westlaw.com/Document/I720ff0cab79911ed8636e1a02dc72ff6/View/ FullText.html?transitionType=Default&contextData=(sc.Default)&firstPage=true
[8] https://www.npr.org/2023/08/16/1194202562/new-york-times-considers-legal-action-against-openai-as-copyright-tensions-swirl
[9] https://arxiv.org/pdf/2005.14165.pdf, Language Models are Few-Shot Learners
[10] The AI feedback loop: Researchers warn of 'model collapse' as AI trains on AI-generated content | VentureBeat
[11] The Major Companies and Media Outlets Blocking OpenAI's Crawler GPTBot (businessinsider.com)
[12] https://www.theguardian.com/technology/2023/sep/01/the-guardian-blocks-chatgpt-owner-openai-from-trawling-its-content?CMP=Share_AndroidApp_Other
[13] https://blog.google/technology/ai/ai-web-publisher-controls-sign-up/
[14] https://techcrunch.com/2014/01/26/google-deepmind/
[15] https://www.theverge.com/2023/1/23/23567448/microsoft-openai-partnership-extension-ai
[16] https://www.reuters.com/technology/openai-tech-gives-microsofts-bing-boost-search-battle-with-google-2023-03-22/
[17] https://www.wsj.com/articles/google-is-spared-a-search-engine-switch-by-a-major-partner-f06b734f
[18] OpenAI works on copyright solution for large AI models (the-decoder.com)
[19] Microsoft wants publishers to be part of its chatbot success story (the-decoder.com)
[20] https://techcrunch.com/2023/02/03/the-uk-rolls-back-controversial-plans-to-open-up-text-and-data-mining-regulations/?guccounter=1&guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&guce_referrer_sig=AQAAAN14dPc1kVOZS7z-bMQF2DewYjTaYKxaHEzk9ZrRUE17tUNjw1QlzP0CD_TFERj2NjYJ2xo5fjxH2zcFriiP3Joc3BxBkPO2TYadC8PNu3ejHfhWuOYfIhfI8hCcTixMQ49I7N1213_frS-SASM3nHzfo5DbC2nuupPmtxUjUY27
[21] https://www.businessinsider.com/openai-google-anthropic-ai-training-models-content-data-use-2023-6?r=US&IR=T
[22] https://www.dailymail.co.uk/sciencetech/article-11948855/ChatGPT-falsely-accuses-law-professor-SEX-ATTACK-against-students.html
[23] https://www.washingtonpost.com/technology/2023/08/16/chatgpt-ai-political-bias-research/
[24] https://link.springer.com/article/10.1007/s11127-023-01097-2#Sec16