Last updated: August 17, 2026
Between 2023 and 2026, the way AI systems get their training and grounding data changed. The early large language models were built largely on scraped web text. Then the lawsuits started, the scraping got contested, and a parallel market appeared: AI companies signing formal content licensing deals with publishers, news agencies, forums, and stock libraries. This post is the master map of that market. It pulls every major deal we could verify into one place, groups them by buyer, and points to the deep-dive posts for each company.
Table of Contents
The headline numbers, as of July 2026: OpenAI has signed roughly two dozen publisher and data deals, the single largest being a reported 250 million dollars over five years with News Corp. Reddit alone disclosed 203 million dollars in aggregate data-licensing contract value in its IPO filing. Amazon, Meta, Google, Microsoft, and Mistral have all entered the market, while Anthropic has stayed out of publisher licensing almost entirely and instead agreed to a proposed 1.5 billion dollar copyright settlement, which still required final court approval at the time of this audit. Most deal values are never disclosed, so any total understates the real spend.
Why AI companies are licensing content
Three forces pushed AI companies from “scrape everything” to “pay for some of it.” First, litigation risk: The New York Times, a coalition of Canadian broadcasters, Indian publishers, and others sued over unlicensed training, making a paper trail of consent valuable. Second, freshness and grounding: a licensed real-time feed from a news agency lets a chatbot answer “what happened today” with attribution, which scraped snapshots cannot. Third, quality: curated, rights-cleared corpora (a news archive, a stock library, a decade of expert forum answers) are cleaner training fuel than the open web.
For publishers, the deals split into two shapes. Some are training licenses (the AI company can use an archive to train models). Others are display or grounding licenses (the AI can surface summaries, quotes, and links to live articles inside its product, with attribution). Many recent deals bundle both. The distinction matters because training rights and real-time display rights affect different uses of the content. A commercial agreement can make content available for attributed display, but it does not guarantee that the content will be cited for a particular query.
The master map: every major AI content deal
The table below groups deals by buyer. “First deal” is the earliest verified public agreement for that buyer. “Reported value” is labeled: figures come from press reporting and are marked “reported” where unconfirmed by the parties, and “undisclosed” where no number was ever published. Do not read a blank or “undisclosed” as “small”; several of the largest deals were never given a public figure.
| Buyer (AI company) | Key partners | First deal | Scope | Reported value |
|---|---|---|---|---|
| OpenAI | Associated Press, Axel Springer, Financial Times, News Corp, Vox Media, The Atlantic, Time, Conde Nast, Hearst, The Guardian, Washington Post, Grupo Folha, Grupo UOL, Reddit, Stack Overflow, Shutterstock | Jul 2023 (AP) | Training plus real-time display and citations in ChatGPT and search | News Corp reported 250M over 5 years; Reddit reported ~70M/yr; FT reported 5-10M/yr; most others undisclosed |
| Google / Gemini | Reddit, Stack Overflow, Associated Press, plus a publisher pilot (Der Spiegel, El Pais, The Guardian, Washington Post, Financial Times and others) | Feb 2024 (Reddit) | Training and grounding for Gemini; news display features via a commercial pilot | Reddit reported 60M/yr; others undisclosed |
| Amazon | The New York Times, Conde Nast, Hearst | May 2025 (NYT) | Model training and real-time answers in Alexa and Alexa for Shopping (formerly Amazon Rufus) | NYT reported 20-25M/yr; Conde Nast and Hearst undisclosed |
| Meta | Reuters, News Corp, Le Figaro, Prisa, Suddeutsche Zeitung, plus reported deals with CNN, Fox News, USA Today, and People Inc. | Oct 2024 (Reuters) | Real-time news answers in Meta AI plus model training | News Corp reported up to 50M/yr over 3 years; others undisclosed |
| Microsoft | Informa, plus Copilot Daily partners (Financial Times, Reuters, Axel Springer, Hearst, USA Today Network) | May 2024 (Informa) | News summaries in Copilot; training access to specialist content | Informa reported 10M-plus initial; others undisclosed |
| Mistral | Agence France-Presse (AFP) | Jan 2025 (AFP) | Real-time AFP feed and archive back to 1983 as a grounding module in Le Chat (not for training) | Undisclosed |
| Apple | Shutterstock; reported talks with Conde Nast, NBC News, IAC | Late 2023 (talks) | Training data for Apple’s generative AI systems | Reported offers of 50M-plus per publisher; Shutterstock reported 25-50M |
| Anthropic | No publisher licensing deals; notable for a copyright settlement instead | None (licensing) | Agreed to a proposed authors’ class settlement rather than signing publisher licenses | Proposed 1.5B settlement, not a content license; final approval pending |
| Perplexity | Time, Fortune, Der Spiegel, Entrepreneur, The Texas Tribune, WordPress.com, and later dozens more | Jul 2024 (Publishers Program) | Revenue sharing when a publisher is cited, plus API and analytics access | Revenue-share pool reported at 42.5M |
Three data platforms sit underneath many of these deals and are worth breaking out on their own, because they license to multiple AI buyers at once:
| Data platform | Licenses to | What they sell | Reported value |
|---|---|---|---|
| Google, OpenAI | Real-time and archived user posts for training and grounding | Google reported 60M/yr; OpenAI reported ~70M/yr; 203M aggregate contract value disclosed at IPO | |
| Stack Overflow | Google (Gemini), OpenAI | Developer questions and answers via its OverflowAPI | Undisclosed |
| Shutterstock | OpenAI, Meta, Google, Amazon, Apple, Reka | Licensed images, video, audio, and 3D assets for training | 104M in AI licensing revenue in 2023 across all buyers; individual big-tech deals reported at 25-50M each |
Adding up the money
There is no clean industry total, because most contracts are private and mix one-time, annual, and multi-year figures. What we can do is total the reported headline numbers, clearly labeled as reported and not confirmed by the parties:
- OpenAI – News Corp: reported 250 million dollars over five years, the single largest known deal.
- Meta – News Corp: reported up to 50 million dollars a year over three years.
- Reddit: 203 million dollars in aggregate data-licensing contract value disclosed at IPO, spanning Google (reported 60 million a year) and OpenAI (reported around 70 million a year).
- Amazon – The New York Times: reported 20 to 25 million dollars a year.
- OpenAI – Dotdash Meredith: reported around 16 million dollars a year.
- OpenAI – Financial Times: reported 5 to 10 million dollars a year.
- Shutterstock: 104 million dollars of AI licensing revenue in 2023 alone, across all its AI customers.
- Perplexity: a reported 42.5 million dollar revenue-sharing pool for its publisher partners.
Even limiting the count to these disclosed figures, the reported committed and annual spend runs comfortably into the high hundreds of millions of dollars, and the true number is larger because the Vox Media, Atlantic, Time, Conde Nast, Hearst, Guardian, Washington Post, Reuters, AFP, and dozens of other deals carry no public price at all. Treat any total as a floor.
Timeline of the licensing wave, 2023 to 2026
- 2023: The market opens. OpenAI signs the Associated Press in July 2023 (its first major news deal) and a six-year Shutterstock license the same year, then in December 2023 signs Axel Springer, its first global news publisher, for a reported figure in the tens of millions of euros. Apple is reported to be in talks with Conde Nast, NBC News, and IAC.
- 2024: The floodgates open. Reddit signs Google in February 2024 for a reported 60 million a year, then OpenAI. OpenAI signs the Financial Times (April), Reddit and Stack Overflow (May), News Corp for a reported 250 million over five years (May), plus Dotdash Meredith, Vox Media, The Atlantic, Time, Conde Nast, and Hearst through the year. Microsoft signs Informa. Perplexity launches its Publishers Program in July. Meta signs Reuters in October, its first news deal.
- 2025: The market broadens beyond OpenAI. Mistral signs AFP (January). OpenAI adds The Guardian, Schibsted, Axios, and the Washington Post. Amazon enters with The New York Times (May), then Conde Nast and Hearst for Alexa for Shopping (July). Google opens a commercial pilot with about twenty publishers. Anthropic agrees to a proposed 1.5 billion dollar settlement in an authors’ copyright suit rather than licensing.
- 2026 (year to date): Consolidation and expansion. Meta announces partnerships with News Corp, Le Figaro, Prisa, and Suddeutsche Zeitung in March. OpenAI adds its first media partnership in Brazil with Grupo Folha and Grupo UOL in May. More regional deals appear, and several publishers now license to multiple AI companies.
Company by company (the short version)
Each buyer has its own deep-dive post. Here is the one-paragraph summary and where to read more.
OpenAI
OpenAI is the most active buyer by a wide margin, with roughly two dozen publisher and data deals combining training rights with real-time display and citations in ChatGPT and its search product. Its News Corp agreement (a reported 250 million dollars over five years) is the largest known content deal in the market. For the full partner list, values, and scope, see our OpenAI publisher deals map and where ChatGPT gets its data.
Google and Gemini
Google’s approach leans on data platforms (Reddit at a reported 60 million a year, Stack Overflow via OverflowAPI) plus the Associated Press feed for the Gemini app, and a publisher pilot that Google frames as commercial partnerships rather than licenses. Full detail in our Google and Gemini publisher deals post.
Amazon
Amazon arrived later and split its buying across two products: Alexa answers (The New York Times, at a reported 20 to 25 million a year, including NYT Cooking and The Athletic) and Alexa for Shopping (Conde Nast and Hearst). It is covered alongside the other non-OpenAI buyers in AI content deals beyond OpenAI.
Meta
Meta’s first news deal was Reuters in October 2024, powering real-time answers in Meta AI, followed by a reported up to 50 million a year agreement with News Corp and reported deals with CNN, Fox News, USA Today, and People Inc.
Microsoft
Microsoft’s content moves center on Copilot: a specialist-content deal with Informa (reported at 10 million dollars-plus initially) and a Copilot Daily partner program spanning the Financial Times, Reuters, Axel Springer, Hearst, and the USA Today Network.
Mistral and Apple
Mistral signed a single high-profile deal, a multi-year agreement with AFP to ground Le Chat in real-time news (explicitly not for training). Apple was reported in late 2023 to be shopping 50 million dollar-plus training deals to Conde Nast, NBC News, and IAC, and separately licensed Shutterstock.
Anthropic
Anthropic is the outlier: it has essentially no comparable publisher licensing deals. In 2025 it agreed to a proposed 1.5 billion dollar class settlement over books obtained from shadow libraries. The proposal covers past conduct rather than granting a forward license, and final court approval remained pending after a May 2026 hearing. The litigation angle is covered in our AI copyright lawsuits tracker.
Perplexity
Perplexity took a different model entirely: rather than paying flat licensing fees up front, it launched a Publishers Program in July 2024 that shares advertising revenue with publishers whose content is cited, backed by a reported 42.5 million dollar pool, plus API and analytics access. Read the Perplexity Publishers Program breakdown and how Perplexity works.
License versus litigate: the two publisher strategies
The most important pattern in this market is that publishers split into two camps, and some are in both at once. One camp licenses. The other sues. A few, most notably News Corp and The New York Times, do both: News Corp licensed to OpenAI and Meta while suing Perplexity, and The New York Times licensed to Amazon while suing OpenAI and Microsoft.
On the litigation side, The New York Times, a coalition of Canadian news organizations, Indian publishers, and others have gone to court over unlicensed training and output, and Anthropic’s proposed 1.5 billion dollar settlement showed the potential cost of copyright litigation. Licensing and litigation can change what content a company is permitted to use, but they do not determine which sources a product will cite for each query. For the full case list, statuses, and parties, see our AI copyright lawsuits tracker.
What this means for your AI visibility
Here is the practical takeaway for anyone doing marketing, SEO, or PR. These deals can expand the content an AI product is permitted to display or use for real-time grounding. They do not prove that a licensed outlet will rank more highly or that coverage in one will cause a brand to appear in an answer. The reliable signal is whether the outlet and the brand are actually cited for the prompts you track.
The way to know which licensed sources appear alongside answers about your brand or category is to measure the visible citations. LLM Pulse supports weekly and daily tracking across ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overviews. It tracks brand mentions, cited sources and URLs, share of voice and sentiment. If a publisher begins appearing more often, citation tracking shows that change without assuming the commercial deal caused it. You can also compare how different models cite differently with the Models Comparison view, and export the underlying citation data. For the broader method, see how to track which sources AI cites and AI search optimization.
Summary
The AI content licensing market went from a handful of experiments in 2023 to a broad, multi-buyer market by 2026. OpenAI leads on volume, News Corp is the biggest single reported deal at 250 million dollars over five years, data platforms like Reddit and Shutterstock quietly power multiple buyers, and Anthropic has no comparable publisher program and agreed to a proposed 1.5 billion dollar court settlement that still required final approval. Most deal values stay private, so the reported totals are a floor, not a ceiling. For your brand, the signal that matters is downstream: which of these newly licensed sources AI systems actually cite, and whether you appear in them.
FAQ
What is an AI content licensing deal?
It is a commercial agreement in which an AI company pays a publisher, news agency, forum, or media library for the right to use its content. Deals cover training (using an archive to build models), display or grounding (surfacing summaries, quotes, and links to live content inside the AI product with attribution), or both. In exchange, the content owner gets money and, usually, attribution and links back.
Which AI content licensing deal is the biggest?
By reported headline value, the largest single publisher deal is OpenAI’s agreement with News Corp, reported at up to 250 million dollars over five years. Reddit disclosed a larger aggregate figure, 203 million dollars in total data-licensing contract value, but that spans multiple buyers (Google and OpenAI) rather than one deal.
How much has the AI industry spent on content licensing?
There is no confirmed industry total, because most contracts are private and mix one-time, annual, and multi-year terms. Adding only the publicly reported figures still reaches the high hundreds of millions of dollars, and the real number is higher because many major deals (Vox Media, The Atlantic, Reuters, AFP, Conde Nast, Hearst, and others) were never given a public price. Any total should be read as a floor.
Which publishers signed deals and which ones sued?
Many major publishers signed, including News Corp, Axel Springer, the Financial Times, Vox Media, The Atlantic, Time, Conde Nast, Hearst, The Guardian, and Reuters. Others sued instead, led by The New York Times against OpenAI and Microsoft, a coalition of Canadian broadcasters, and Indian publishers. Some do both: News Corp and The New York Times each licensed to some AI companies while suing others.
Does Anthropic have content licensing deals?
As of July 2026, Anthropic has essentially no publisher content licensing deals comparable to OpenAI’s or Amazon’s. It is best known instead for a proposed 1.5 billion dollar class settlement concerning books obtained from shadow libraries. The proposal addresses past conduct rather than providing a forward-looking license, and final approval remained pending after a May 2026 hearing.
Why do AI companies pay for data they used to scrape for free?
Three reasons: to reduce copyright litigation risk by having documented consent, to get fresh real-time feeds (a licensed news wire lets a chatbot answer today’s questions with attribution, which scraped snapshots cannot), and to get cleaner, rights-cleared, high-quality corpora for training.
Do these deals affect whether my brand appears in AI answers?
They can affect what content an AI product is permitted to display or use for grounding, but they do not guarantee that a covered brand or publication will appear. Track visible citations across the prompts that matter to see whether the source is actually being displayed.
