AI

The AI Content Licensing Deals Reshaping Publisher Economics

The AI Content Licensing Deals Reshaping Publisher Economics

Publishers spent two years watching AI companies train on their content for free, and then, starting in late 2023, a real market emerged to actually pay for it. The pattern across these deals is more consistent than the headline dollar figures suggest, and understanding the structure — not just the price tags — is useful for reading where this is headed next.

The deals, in the order they happened

OpenAI’s first major publisher deal was with Axel Springer, announced December 13, 2023 — the first arrangement of its kind, covering Politico, Business Insider, Bild, and Welt. It let OpenAI summarize even paywalled content within ChatGPT with attribution and links back, and permitted use of the content for model training. Bloomberg reported the value at “tens of millions of euros,” though neither company has confirmed an exact figure publicly.

OpenAI signed with News Corp in May 2024, reportedly worth more than $250 million over five years, combining cash and technology credits, covering titles including the Wall Street Journal, Barron’s, MarketWatch, and the New York Post. Google’s deal with Reddit, announced February 2024, gave Google access to Reddit’s Data API for roughly $60 million a year — one of the few figures in this space independently corroborated by multiple outlets rather than resting on a single anonymous source.

OpenAI struck a deal with Stack Overflow in May 2024, giving OpenAI API access to Stack Overflow’s accumulated Q&A data through a product called OverflowAPI — a genuinely different kind of asset than a news archive, and a reminder that these deals aren’t limited to traditional publishers. In August 2024, OpenAI signed with Condé Nast, covering The New Yorker, Vogue, Vanity Fair, Wired, GQ, and Architectural Digest, with content usable across ChatGPT and search products; financial terms were undisclosed.

The pattern underneath the individual deals

A few things hold consistent across nearly all of these arrangements:

  • Cash figures are almost always estimates, not confirmed numbers. Outside of the Reddit-Google deal, most dollar amounts circulating in coverage trace back to anonymous sourcing or industry estimation, not an official disclosure from either party. Worth treating as “reported,” not “confirmed,” whenever you see a specific number attached to one of these.
  • Attribution and linking are usually part of the deal, not just payment. The Axel Springer arrangement specifically included attribution and links back to source articles — these deals aren’t purely about licensing content for training in isolation; they’re often also about how that content gets surfaced and credited in the AI product itself.
  • The deals concentrated around late 2023 through 2024, roughly the period when it became clear AI search products were going to be a real, ongoing feature of how people find information — not a temporary novelty.
  • Deal structures vary by what the publisher actually has. A news archive gets valued differently than a Q&A knowledge base like Stack Overflow’s, which gets valued differently than a platform’s live, constantly-refreshing user-generated content like Reddit’s. There isn’t one template price per word or per article across this market.

These deals represent AI companies acknowledging that the content powering their products has real, negotiable value — a meaningful shift from the earlier default assumption that public web content was simply available for training at no cost.

What this means if you’re not News Corp

Almost nobody reading this runs a publisher large enough to negotiate a direct licensing deal with OpenAI or Google. But the existence of this market does have practical implications for smaller sites and independent creators:

  • It confirms training data has real, contested value — which is useful context if you’re deciding how you feel about GPTBot, ClaudeBot, or similar training-focused crawlers accessing your content for free. You’re not being paranoid; large companies are actively paying substantial money for exactly this kind of access when the content is valuable enough.
  • It’s a two-tier system, realistically. Large publishers with leverage and legal resources can negotiate compensation; everyone else works with the same binary choice of allow-or-block via robots.txt, without a negotiation in between. That’s an uncomfortable asymmetry, and it’s not clear from current developments that it’s closing anytime soon.
  • Blocking training bots remains the practical lever available to smaller sites — not because it will force a negotiation, but because it’s the actual mechanism available for expressing a preference about training use, distinct from search-citation use (see Google-Extended, GPTBot, and ClaudeBot’s separate, training-specific user-agents).

What happens when there’s no deal at all

It’s worth being clear about the default state for the vast majority of the web’s content, which is genuinely different from any of the arrangements above: no deal, no payment, and a training bot that either respects a robots.txt block or doesn’t. The publishers with disclosed licensing deals are a small, visible minority — large enough to negotiate, valuable enough to be worth negotiating with. Everything else on the open web has effectively been operating in the earlier, uncompensated default this whole time, whether or not it’s been explicitly noticed. That asymmetry is worth keeping in view whenever a new deal gets announced — a single high-profile agreement can create the impression that a broad new norm of compensation has arrived, when what’s actually happened is one more addition to a still-short, still-exclusive list.

Where I think this goes next

The deals disclosed so far cluster around publishers with either large, distinctive archives (Condé Nast’s titles) or genuinely unique data structures (Stack Overflow’s Q&A format, Reddit’s live discussion volume). I’d expect the next wave of deals, if they materialize, to follow the same logic — content that’s either hard to replicate or valuable specifically because of its structure, rather than a broad, uniform licensing wave across the open web. For most site owners, the realistic takeaway isn’t “wait for your own deal” — it’s understanding that the robots.txt decision you make for AI training bots is effectively your only lever in a market where the negotiated alternative isn’t actually available to you.

Rakibuzzaman Siam
Rakibuzzaman Siam Customer Experience Specialist at Rank Math, building AI automation projects on the side.