87.A new frontier is fast emerging around news influenced by artificial intelligence (AI). Some issues extend existing challenges examined in the previous chapter. Others feel new.185 The full implications for news remain uncertain but some trends are becoming clearer.
88.AI has long been used in news and will likely shape parts of production further,186 for example automating rote tasks.187 AI-assisted journalists are already being used.188 Concerns about job cuts loom.189 Some publishers are also trialling chatbots which can be fine-tuned on inhouse news.190 These efficiencies alone are unlikely to transform the nature of news or media finances.191 However, we note that unresolved challenges around hallucinations, biases and audience mistrust may dampen enthusiasm for AI journalism in the short term.192
89.Public service media warrant particular scrutiny. Some industry experts have suggested that AI will help public service broadcasters (PSBs) better serve diverging audience needs; others have argued that over-personalisation of news content using AI might undermine the concept of delivering shared experiences that bring audiences together.193
90.During our visit to San Francisco we heard that the traditional concept of internet search is being disrupted by generative AI summaries which draw on multiple sources. Aravind Srinivas, Co-Founder of Perplexity AI, outlined the shift away from keyword searches and towards more extended questions with AI tools.194 Meta’s advances in wearable technologies suggest that audio and visual interactions with AI will become more common and seamless. At Google we discussed the proliferation of AI-summary tools, like those that can convert lengthy tomes into engaging AI-narrated podcasts.195 In future, these might be interruptible, enabling users to ask the AI narrators for more information.
91.This all suggests that the trend towards disintermediation of news by tech platforms will continue.196 Like search engines and social media, large language models (and the applications built on top of them) are likely to have increasing influence over the types of information people see. We cannot yet know what the most popular news applications of the future will look like; much also depends on how Google responds to the challenge to its search business. But some issues are already clear. Providing search users with much longer answers to questions may obviate the need for visiting the actual source of news. Whether the AI provides one answer to a question about current affairs, or a variety of views, may shape the average casual user’s views on a topic. Political risks for tech firms loom too, as the way they curate information attracts scrutiny—illustrated recently by the controversy over Google’s Gemini AI.197
92.There is a possibility that diverging regulations in the US, EU, UK and China will influence what types of AI tools (including news chatbots) are used in different jurisdictions—perhaps shaped by rules on using copyrighted or personal data for training. Meta for example did not release its latest AI model in the EU, reportedly citing an “unpredictable” regulatory environment.198
93.During our visit to San Francisco, we were told that up-to-date news will remain valuable to AI firms, as this is used to provide models with timely and accurate information. Whether, and how much, tech firms will pay remains unclear though.199 Media firms with business models based on users clicking through to a website might suffer if the AI summary is ‘good enough’ for the average reader.200 Some stakeholders worry that this could make it economically unviable for some news outlets to continue producing quality journalism.201
94.James Harding of Tortoise Media told us that generative AI would “increase the value of quality journalism because there will be so much stuff out there that is a mash-up of everything else that is out there”.202 Conversely, Professor Nielsen has warned that generative AI could make some publishers “more efficient at delivering something that audiences and advertisers increasingly do not value”.203 Generic content is likely to become less useful to AI firms and discerning readers alike.204
95.The AI landscape is fast becoming dominated by OpenAI (backed by Microsoft), Google, Anthropic (backed by Amazon) and Meta, followed by a series of smaller firms.205 A long tail of smaller (likely open source) providers may also emerge, but market consolidation and competition issues already loom.206 If the trend towards generative search continues, many smaller media outlets may struggle to make money or reach consumers—particularly if they do not have licensing deals granting them revenue streams and exposure.207 Elsewhere in the market, there is a risk that a few (larger) media outlets will experience growing dependency on a handful of AI firms. In future, it is conceivable that tech firms might exert extensive financial influence over some news outlets (or even buy them outright via affiliates) and secure a ready supply of timely journalism.
96.We found Ofcom’s media ownership and merger rules poorly suited to this evolving environment. Ofcom has a duty to “maintain a sufficient plurality of providers of television and radio services”, but its current frameworks focus largely on traditional print and broadcast outlets. Ofcom reviews the Media Ownership Rules every three years. The Secretary of State can request that Ofcom carry out a Public Interest Test on potential media mergers.208
97.In 2021 Ofcom called for updates to the Media Public Interest Test framework to include online “news creators”, which would cover online-only news websites.209 This modest change has remained outstanding for three years. The Government recently launched a consultation on implementing this recommendation by amending the definition of a “newspaper” in the Enterprise Act 2002 to include online news publications.210 The consultation set out a “policy intention to exclude aggregators … and online intermediaries” from the updated rules, citing the number of firms that could be caught in scope and Ofcom’s 2021 report as justification. The Government suggested that this approach
“reflects the way in which news is consumed in the modern day, whilst avoiding bringing into scope additional entities that are less likely to pose public interest concerns”.211
98.We were disappointed that the Government did not seek a wider update to the media plurality rules and we struggled to follow its logic in pursuing a limited approach via this consultation. Research by the regulator has highlighted the “significant role” that online news intermediaries play across the news value chain.212 The scope of an expanded regime could be limited to the largest news intermediaries without much difficulty. The risk of burden to business is less clear too: if the Secretary of State decides to investigate a merger then presumably there will be public interest grounds for doing so—and if there is no investigation then it is not obvious how significant the burden to business would be.
99.In our evidence session the Minister further suggested that the Digital Markets, Competition and Consumers Act 2024 would address media competition. While the Act may help,213 it is no panacea. The legislation is cross-sector, not specific to news. It remains unclear how far the provisions of that Act would affect tech firms’ practices in news, and how quickly. Experiences from the EU suggest that rapid and robust implementation will be key, but difficult.214 ITV warned that “the threat to news business models risks outpacing the speed of implementation”.215
100.Advances in generative AI are enabling tech firms to provide engaging and high quality news summaries. This suggests they are increasingly acting as publishers and may need to be regulated as such. Ofcom’s media plurality framework is rapidly becoming outdated, and the previous Government’s years-long timeline for implementing vital changes has been inadequate. The Government should commit to a 12 month deadline for responding to future Ofcom priority recommendations on media plurality.
101.The Government’s proposed amendments to the media mergers regime are a good start. But we are disappointed it has not sought a wider update to the media plurality regime. The decision to exclude online intermediaries looks oddly short sighted given the rapid advances in tech firms’ ability to produce news summaries. We appreciate that tech firms are not newspapers but this does not mean their evolving role in the news landscape should be overlooked. We recommend the Government works with Ofcom to set out plans and timelines for capturing online news intermediaries within the scope of the media ownership rules.
102.As generative AI progresses and tech platforms continue to dominate advertising and data flows, we anticipate growing convergence between regulatory remits and their impact on news publishers. Further co-ordination across regulators may be helpful. The News Media Association highlighted the focus of the Information Commissioner’s Office (ICO) on data privacy, saying that its proposals around cookies would undercut news media business models.216 DMG Media noted separate work by the Competition and Markets Authority (CMA) on Google’s proposed changes to third-party cookies, which would also damage publisher revenues.217 Ofcom has statutory duties to oversee media plurality but still lacks adequate means, while the CMA has powers but its priorities for news are less clear. The issues around the use of personal data for AI training are a further potential challenge.
103.The Digital Regulation Cooperation Forum, which brings together the CMA, ICO, Ofcom and the Financial Conduct Authority, has various related workstreams (for example AI and data protection) but does not appear to have dedicated projects addressing the impacts of regulation on news media.218
104.The Digital Regulation Cooperation Forum should establish a dedicated workstream examining areas of regulatory crossover, conflict and collaboration that will affect the news sector—focusing in particular on privacy, advertising and competition.
105.Our recent report on large language models examined the use of copyright materials for AI training.219 In brief, AI firms need significant amounts of data to train their models. This includes text but increasingly audio, video and other types of data. Under the UK’s current text and data mining rules, obtaining permission typically involves acquiring a licence or relying on an exception. Noncommercial research is however permitted.
106.Tech firms have said their use of data for AI training is legitimate—citing legal exceptions and arguing that allowing machines to ‘read’ and ‘learn’ from material should be permitted. Many copyright holders (such as news publishers, academics, performers and similar) have argued in contrast that the copyright law exceptions do not apply, and that tech firms should seek permission or provide renumeration for using their data.220
107.Our present inquiry examined three implications of this for journalism in more detail.
108.First is the difficulty of balancing competing strategic objectives. Many stakeholders favour an AI-friendly approach and looser rules on text and data mining—the process by which tech firms obtain the content needed to train their models.221 There are multiple reasons. Boosting productivity and improving public sector service efficiencies will increasingly require AI knowhow. Yet onerous regulation may make it harder to attract AI investment and talent, or encourage tech firms to withhold products from the UK market. Making AI firms pay more for data might hamper startups. Some argue that deepening geopolitical competition with China makes it important for national security reasons for the UK to ‘win the AI race’.222
109.On the other hand, the UK’s diversified economy and longstanding intellectual property rules benefit a range of people and businesses. As our reports on the creative industries and on large language models found, undermining copyright principles will likely damage the UK’s £100 billion creative industries (which the new Government has recognised as a strategic growth sector).223 We were told that AI frameworks that do not improve the economic outlook for quality journalism risk eroding the foundations of our democracy.224 Some suggest that the best response to data scraping by big tech firms should not involve making unethical behaviour cheaper and easier for everyone else,225 particularly as developers seem prepared to pay vast sums on compute.226
110.Favouring tech firms without supporting media organisations also risks further splintering the news landscape along lines identified in Chapter 3: wealthier outlets (who can strike licensing deals and may enjoy greater prominence) may serve a small demographic of well-informed news aficionados, while a much larger proportion of the population becomes increasingly poorly served as the economics of mass market journalism worsen, smaller outlets offering alternative perspectives struggle to gain prominence, and local news deserts expand.227 Baroness Jones of Whitchurch, Parliamentary Under-Secretary of State for the Future Digital Economy and Online Safety at the Department for Science, Innovation and Technology, said she was “trying to find a way through that is acceptable to all sides”.228
111.The principle of using real-time and archive news for the development of AI products remains contested.229 Lawsuits are multiplying on some fronts even as licensing deals between news organisations and tech firms emerge elsewhere.230 OpenAI told us that they were leading the way on establishing partnerships. Several stakeholders welcomed this move towards more partnership-based development.231 Sceptics described the deals as an insurance policy against litigation, suggesting it encourages future lawsuits to be directed against challenger AI firms (who may not be able to afford such deals) rather than wealthy incumbents.232
112.Whether generative AI licensing deals are one-time offers or long-term partnerships remains uncertain. Under the first scenario, AI firms would extract most of the value of news content from a news publisher upfront and then train new models (by reusing the tokens and ‘vector representations’) without having to relicense when the deal expires. Alternatively, deals might create long-term partnerships which envisage relicensing and align the interests of the AI developer and news publisher—perhaps with a particular focus on up-to-date news from reputable sources, which remains valuable for generative search.233
113.The nature of licensing deals made now will set precedents and may influence which firms survive into the future.234 The terms governing access to real-time news, royalties, opt-outs, anti-cloning protections, transparency and responsiveness to market conditions will be key.235 Smaller outlets with less data and little bargaining power risk being left out. Collective licensing (for example through organisations like the Copyright Licensing Agency) could provide a more even playing field, alongside a system for responsible data access—though progress, appetite and prospects are mixed.236 (We note that our report is focused on news: while the principles of intellectual property apply broadly, it is possible that the details of AI licensing agreements will differ across economic sectors).
114.The third implication of AI and copyright rules that we considered was the practical difficulties inherent in developing new text and data mining rules. In 2022 the Intellectual Property Office (IPO) proposed to change the UK’s rules to allow any form of commercial mining. This Committee examined the trade-offs and concluded the IPO had not considered sufficiently the impact on the creative sector. The Government subsequently confirmed it would no longer pursue a “broad copyright exception” and set up a working group to develop a new code of practice.237 We followed this work closely and were disappointed by the way the IPO-led roundtables were handled and the lack of progress. We wrote to the Government in May 2024, raising concerns that their
“record on copyright was inadequate and deteriorating … the Government has set up and subsequently disbanded a failed series of roundtables led by the Intellectual Property Office … The Government’s reticence to take meaningful action amounts to a de facto endorsement of tech firms’ practices”.238
115.Recent media reports suggest that the new Government has been exploring an opt-out approach,239 building on the previous administration’s consultation.240 This might involve specifying that tech firms may acquire data for non-research purposes unless rightsholders specifically decline. Anyone wishing to opt out might use tools like robots.txt to tell AI crawlers to exclude a site. A comparable regime is used in the EU—RELX, an information platform, previously said it worked “tolerably well”.241 Google advocates allowing mining for “both commercial and research purposes”.242
116.But adopting an EU-style opt-out scheme wholesale would be problematic. The Financial Times told us that there is no clear enforcement mechanism for infringements, short of costly and uncertain court cases that few can afford.243 The lack of transparency also makes it hard to prove illegal scraping anyway. DMG Media highlighted the stakes for those considering litigation in the UK: “If a news publisher loses it would then be open season for LLMs to use its copyright content without restriction.”244
117.An assessment by the European Publishers Council found that publishers cannot tell if a crawler is operating for research or commercial purposes, and it is technically difficult or impossible to block crawlers outright. Some third parties might appear to be crawling for academic research but then give or sell data to tech firms, who are one step removed from any abuses.245 A note on Google Search Central acknowledges that “while Googlebot and other respectable web crawlers obey the instructions in a robots.txt file, other crawlers might not”.246 Publishers also worry that blocking crawlers in other ways may affect whether they show up in other online search rankings, and have accused tech firms of exploiting dominance in internet search to gain advantages in obtaining AI training data.247
118.The complexity of the issues outlined above should not become an excuse for inertia, and we note the Government’s forthcoming plans in this space.248 We welcomed the Prime Minister’s comments recognising the “basic principle that publishers should have control over and seek payment for their work”.249 Jon Slade, Chief Operating Officer of the Financial Times, suggested that publishers need more clarity on copyright law.250 This could specify more clearly how copyright applies to text and data mining for large language models used for commercial purposes, and establish enforceable protocols and sanctions for the Government’s future text and data mining regime.
119.At a minimum, this would likely require a transparency mechanism enabling rightsholders to check if their data has been used. Original repositories of raw data might be too unwieldy, but lists of websites or metadata may be manageable.251 If tech firms have concerns about revealing commercially sensitive data, vetted researchers, the Government or a regulatory unit could be established to act as an ‘honest broker’ to carry out the checks.252
120.Any regime would need to require web crawlers to identify themselves, or else tech firms could remain immune from retribution where rightsholders’ data has been misused. Rules would also need to be flexible; it is possible that if internet search and generative AI services converge, web crawling activity may do the same. New rules would also need to be clear about how far protections extend to the real-time use of news to help generative AI tools answer questions—as opposed to simply using archive data to train base models.
121.Enforcement will need more work too, as copyright is typically treated as a private matter, and the UK lacks suitable institutions for addressing breaches. The Intellectual Property Office (IPO) does not have regulatory powers comparable to the Competition and Markets Authority (CMA), for example. DMG Media suggested that the Digital Markets Unit, which sits within the CMA, should address anti-competitive use of web crawlers.253 We noted however that such enforcement might still focus on the largest tech firms with Strategic Market Status, without sufficiently addressing the long tail of smaller AI firms that may also be breaching rules.254 A wider approach might involve the IPO referring cases to the relevant existing regulator (whose commensurate powers and remits may need reviewing to ensure meaningful action can be taken).
122.Aside from rule changes, the Government could champion AI firms acting responsibly. Start-up companies like UK-based Human Native AI indicate that there is an emerging market for providing licensed AI training content.255 The Government could encourage such moves to make the UK an attractive AI training destination—particularly around technical areas valuable to fine tune specialised AI models. Finally, following the failure of the IPO-led working group process, the Government should be cautious about the risks of discussion forums becoming protracted exercises in entrenching the status quo.256
123.Baroness Jones of Whitchurch acknowledged the need to encourage AI and also “protect the rightsholders”, including news media. She said the Government was “moving at pace” and noted that the prospect of voluntary agreements was “clearly not the case now”. She further suggested that a transparency mechanism was a “good idea”.257
124.The use of news content to train generative AI has the potential to reshape the economics of the media industry. The UK needs a better framework for governing how this works. There are arguments for and against tougher rules. On the one hand, the UK must remain competitive in AI development, or else lose any claim to international leadership. Economic prosperity, public sector efficiencies and national security all provide good arguments for establishing an AI-friendly training regime.
125.But that does not mean the Government should pursue rules that primarily benefit foreign tech firms (who seem prepared to pay vast sums on energy, computing facilities and staff—but not on data). Previous efforts to find a solution have been weak and ineffectual. The Government must aim for a robust framework that helps the creative industries strike mutually beneficial deals with tech firms, aligns incentives, respects intellectual property and champions responsible AI development in the UK. Media organisations, for their part, will need to continue to demonstrate their value—and be clear that their position is not about special pleading or propping up outlets for which there is limited demand.
126.While we welcome the new Government’s desire to make progress on this issue, we caution strongly against adopting a flawed opt-out regime comparable to the version operating in the EU. Much better means for ensuring technical viability, transparency, consent and enforcement are needed for a new text and data mining regime to work to UK advantage. If the Government gets this right, it can provide speedy regulatory certainty and encourage a new AI-licensing startup scene to flourish too.
127.Any proposal for a new text and data mining regime must include transparency mechanisms that enable rightsholders to check whether their data has been used. It must offer technical enforceability that goes beyond the likes of robots.txt indicators, which remain inadequate. Meaningful sanctions for non-compliance are essential and the Government’s anticipated IP consultation should explore the options for independent regulatory enforcement. Requirements for web crawlers to identify their purpose are needed too. The Government should encourage good practice by championing an emerging market for licensed AI data training providers. We urge the Government to dedicate significant technical, policy and political resource to address these challenges at pace. The Department for Science, Innovation and Technology should outline its plans in response to this report.
128.The Competition and Markets Authority should investigate and address tech firms leveraging dominance in one domain, notably internet search, to secure anti-competitive advantages in obtaining data for generative AI training. We suggest this should be an immediate priority given the pace of market developments and impacts on news media business models.
185 Appendix of Committee visit to San Francisco
186 See for example: Washington Post, Press Release: The Washington Post leverages automated storytelling to cover high school football on 1 September 2017: https://www.washingtonpost.com/pr/wp/2017/09/01/the-washington-post-leverages-heliograf-to-cover-high-school-football/ [accessed 15 November 2024]; Reuters, Press Release: Reuters News Tracer on 15 May 2017: https://www.reutersagency.com/en/reuters-community/reuters-news-tracer-filtering-through-the-noise-of-social-media/ [accessed 15 November 2024]; BBC Research & Development, ‘Natural language processing’: https://www.bbc.co.uk/rd/projects/natural-language-processing [accessed on 2 August 2024]
190 Washington Post, ‘Climate answers’: https://www.washingtonpost.com/climate-environment/climate-answers/ [accessed 16 October 2024]
192 Written evidence from DMG Media (FON0030), Professor Rafael Calvo (FON0047), BBC (FON0059), The Bristol Cable (FON0008), Q 145 (Robert Colvile), Q 10 (Paul Lee), written evidence from Dominic Young (FON0021), YouGov, ‘AI in journalism: how would public trust in the news be affected?’ (April 2024): https://yougov.co.uk/technology/articles/49105-ai-in-journalism-how-would-public-trust-in-the-news-be-affected [accessed 13 November 2024]
193 European Broadcasting Union, Trusted journalism in the age of generative AI (June 2024), p 140: https://www.ebu.ch/files/live/sites/ebu/files/Publications/Reports/open/News_report_2024.pdf [accessed 13 November 2024] See the BBC’s guidelines for generative AI at BBC, ‘Press Release: An update on the BBC’s plans for Generative AI (Gen AI) and how we plan to use AI tools responsibly’ (28 February 2024): https://www.bbc.com/mediacentre/articles/2024/update-generative-ai-and-ai-tools-bbc
194 Appendix on Committee visit to San Francisco
195 Ibid.
196 Appendix of Committee visit to San Francisco. See also Q 145 (Robert Colvile), written evidence from Ofcom (FON0063), Media Reform Coalition (FON0029), BBC (FON0059)
197 See for example: ‘Google apologized for ‘missing the mark’ after Gemini generated racially diverse Nazis’, The Verge (21 February 2024): https://www.theverge.com/2024/2/21/24079371/google-ai-gemini-generative-inaccurate-historical [accessed 15 November 2024].
198 Meta, ‘Building AI Technology for Europeans in a Transparent and Responsible Way’ (10 June 2024): https://about.fb.com/news/2024/06/building-ai-technology-for-europeans-in-a-transparent-and-responsible-way/ [accessed 25 October 2024]; ‘ Meta pulls plug on release of advanced AI model in EU’, The Guardian (18 July 2024): https://www.theguardian.com/technology/article/2024/jul/18/meta-release-advanced-ai-multimodal-llama-model-eu-facebook-owner [accessed 15 October 2024]
199 Appendix on Committee visit to San Francisco
200 Ibid.
201 Appendix on Committee visit to San Francisco. See also written evidence from Felix M. Simon (FON0024), NewsNow (FON0051), NMA (FON0056).
203 Professor Rasmus Kleis Nielsen, ‘How the news ecosystem might look like in the age of generative AI’ (March 2024): https://reutersinstitute.politics.ox.ac.uk/news/how-news-ecosystem-might-look-age-generative-ai [accessed 2 August 2024]
205 Amba Kak, Sarah Myers West and Meredith Whittaker, ‘Make no mistake—AI is owned by Big Tech’, MIT Review, (5 December 2023): https://www.technologyreview.com/2023/12/05/1084393/make-no-mistake-ai-is-owned-by-big-tech/#:~:text=With%20vanishingly%20few%20exceptions%2C%20every,and%20sell%20their%20AI%20products [accessed 13 November 2024]
206 Federal Trade Commission, ‘FTC Launches Inquiry into Generative AI Investments and Partnerships’ (25 January 2024): https://www.ftc.gov/news-events/news/press-releases/2024/01/ftc-launches-inquiry-generative-ai-investments-partnerships [accessed 25 October 2024]
208 The Secretary of State can intervene in a “relevant merger situation” or “special merger situation” involving a broadcaster and/or a print newspaper enterprise on the grounds of “public interest considerations” outlined in the Enterprise Act 2008. The Secretary of State has powers to redefine “broadcasting” or “newspaper”; add or modify media public interest considerations and redefine the conditions for a “special merger situation”. See written evidence from Ofcom (FON0063).
209 Ofcom, The future of media plurality in the UK (November 2021), p 26: https://www.ofcom.org.uk/__data/assets/pdf_file/0019/228124/statement-future-of-media-plurality.pdf [accessed 13 November 2024]
210 Department of Culture, Media and Sport, ‘Consultation on updating the media mergers regime’ (November 2024): https://www.gov.uk/government/consultations/consultation-on-updating-the-media-mergers-regime/consultation-on-updating-the-media-mergers-regime#proposedchanges [accessed 6 November 2024]
211 Ibid.
212 Ofcom, Online news: research update, p 6
213 Written evidence from ITV (FON0019); Q 82 (Sebastian Enser-Wight). Interventions might be possible around app stores, advertising, transparency, self-preferencing and bargaining power, for example.
214 ‘EU probes Apple, Meta and Alphabet under landmark new law’, Financial Times (25 March 2024), available at: https://www.ft.com/content/22ce95a6-e473-4102-a330-f7d02cfb6fd1 [accessed 15 November 2024]
218 DRCF, ‘Workplan 2024–25’ (April 2024): https://www.drcf.org.uk/siteassets/drcf/pdf-files/drcf-workplan-202425/ [accessed 22 October 2024]
219 Communications and Digital Committee, Letter from the Chair to the Secretary of State for Science, Innovation and Technology (2 May 2024): https://committees.parliament.uk/publications/44563/documents/221372/default/
220 Communications and Digital Committee, Large language models and generative AI, para 232
221 Ibid., para 234
222 Department for Science, Innovation and Technology, National AI Strategy (September 2021): https://assets.publishing.service.gov.uk/media/614db4d1e90e077a2cbdf3c4/National_AI_Strategy_-_PDF_version.pdf [accessed 15 November 2024]; Appendix on Committee visit to San Francisco; Carnegie Endowment for International Peace, Charting the Geopolitics and European Governance of Artificial Intelligence (March 2024): https://carnegie-production-assets.s3.amazonaws.com/static/files/Csernatoni_-_Governance_AI-1.pdf [accessed 15 November 2024]
223 Communications and Digital Committee, At risk: our creative future,para 34; Communications and Digital Committee, Large language models and generative AI,para 45; Department for Buisness and Trade, ‘Invest 2035: The UK’s Modern Industrial Strategy’ (14 October 2024): https://www.gov.uk/government/consultations/invest-2035-the-uks-modern-industrial-strategy [accessed 22 October 2024]
226 ‘Thom Yorke and Julianne Moore join thousands of creatives in AI warning’ The Guardian (22 October 2024): https://www.theguardian.com/film/2024/oct/22/thom-yorke-and-julianne-moore-join-thousands-of-creatives-in-ai-warning [accessed 15 November 2024]
229 ‘The Times sues OpenAI and Microsoft over A.I. use of copyrighted work’, The New York Times (27 December 2023), available at: https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html [accessed 15 November 2024]
230 Bloomberg Law, ‘ AI Models Force Media Firms to Pick Licensing or Litigation’ (5 August 2024): https://news.bloomberglaw.com/ip-law/generative-ai-forces-media-firms-to-pick-licensing-or-litigation [accessed 17 October 2024]
231 See for example News Corp, ‘News Corp and OpenAI Sign Landmark Multi-Year Global Partnership’ (22 May 2024): https://investors.newscorp.com/news-releases/news-release-details/news-corp-and-openai-sign-landmark-multi-year-global-partnership [accessed 25 October 2024]
232 Appendix on Committee visit to San Francisco
233 Ibid.
234 See for example the press release announcing a deal between OpenAI and The Atlantic, which emphasises the importance of the publication being discoverable as generative search evolves. The Atlantic, Press release: The Atlantic announces product and content partnership with OpenAI (29 May 2024): https://www.theatlantic.com/press-releases/archive/2024/05/atlantic-product-content-partnership-openai/678529/ [accessed 15 November 2024]
235 Enders Analysis, ‘AI, press and licensing deals’ (2024): https://www.endersanalysis.com/reports/ai-press-and-licensing-deals-chosen-few [accessed 16 October 2024]
236 Communications and Digital Committee, Large language models and generative AI, para 253
237 Ibid., para 229
238 Communications and Digital Committee, Letter from the Chair to the Secretary of State for Science, Innovation and Technology (2 May 2024): https://committees.parliament.uk/publications/44563/documents/221372/default/
239 ‘UK to consult on ‘opt-out’ AI content scraping in blow to publishers’, The Financial Times (16 October 2024): https://www.ft.com/content/26bc3de1-af90-4c69-9f53-61814514aeaa [accessed 17 October 2024]
240 Intellectual Property Office, ‘Artificial Intelligence and IP: copyright and patents’ (October 2021): https://www.gov.uk/government/consultations/artificial-intelligence-and-ip-copyright-and-patents [accessed 22 October 2024]
241 Oral evidence taken before the Communications and Digital Committee inquiry on Large Language Models, 7 November 2023 (Session 2023–24), Q 60
242 Google, ’Unlocking the UK’s AI potential’ (September 2024), p 20: https://blog.google/around-the-globe/google-europe/united-kingdom/ai-potential-uk/ [accessed 26 September 2024]
243 Letter from Matt Rogerson, Director of Global Public Policy & Platform Strategy Financial Times to the Chair of the Communications and Digital Committee, (18 October 2024): https://committees.parliament.uk/publications/45506/documents/225308/default/
245 European Publishers Council, Letter from Matt Rogerson, Director of Global Public Policy & Platform Strategy Financial Times to the Chair of the Communications and Digital Committee, Annex (18 October 2024): https://committees.parliament.uk/publications/45506/documents/225308/default/
246 Google Search Central, ‘Introduction to robots.txt’: https://developers.google.com/search/docs/crawling-indexing/robots/intro#:~:text=The%20instructions%20in%20robots.,the%20instructions%20in%20a%20robots [accessed 13 November 2024]
248 ‘UK to consult on ‘opt-out’ AI content scraping in blow to publishers’, The Financial Times (16 October 2024): https://www.ft.com/content/26bc3de1-af90-4c69-9f53-61814514aeaa [accessed 17 October 2024]
249 ‘Journalism is the lifeblood of British democracy. My government will protect it’, The Guardian (28 October 2024): https://www.theguardian.com/commentisfree/2024/oct/28/keir-starmer-journalism-lifeblood-british-democracy-labour [accessed 13 November 2024]
251 Model developers typically discard original training data once processed, but a repository of metadata indicating the websites or content domains accessed by web crawlers seems unlikely to present insurmountable technical obstacles. See Oral evidence taken before the Communications and Digital Committee inquiry on Large Language Models, 7 November 2023 (Session 2023–24), Q 59 (Dan Conway). See also reporting on the New York Times court case: Business Insider, ‘Why the New York Times’ lawyers are inspecting OpenAI’s code in a secret room’ (11 October 2024): https://www.businessinsider.com/ai-future-copyright-lawsuits-new-york-times-times-authors-2024–10 [accessed 17 October 2024]
252 This might involve vetted researchers, the Intellectual Property Office, or a joint unit with the AI Safety Institute. Funding for this could involve a fee paid by rightsholders, or some form of levy paid by the relevant parties.
254 Competition and Markets Authority, ‘Digital Markets Unit’ (18 June 2024): https://www.gov.uk/government/collections/digital-markets-unit#:~:text=The%20main%20components%20of%20the,penalties%20and%20wider%20administrative%20matters [accessed 25 October 2024]
255 Fortune, ‘Startup that wants to be the eBay for AI data taps Google vets and a top IP lawyer for key roles’ (15 October 2024): https://fortune.com/2024/10/15/human-native-ai-startup-building-marketplace-for-data-hires-veteran-google-execs-top-ip-lawyer/ [accessed 17 October 2024]
256 Communications and Digital Committee, Large language models and generative AI para 252