---
title: "What Grok, Claude, GPT, Gemini, and DeepSeek Actually Trained On"
description: "Every frontier model starts from the same public web. The real difference is the private pile each lab glued on top, and what that means for your data if you build on one."
publishedAt: 2026-09-04
author: Ena Pragma
url: https://enapragma.co/blog/what-each-frontier-lab-actually-trained-on
tags: ["ai-vendor-selection", "data-privacy", "model-training", "ai-governance", "grok", "claude", "gpt", "gemini", "deepseek"]
---

Every frontier model starts from the same pile: the public internet. That's not the differentiator anyone should be evaluating. The differentiator is the private pile each lab glued on top, and for an operator deciding which model to build a workflow on, that private pile is also the thing that determines what happens to your own data once you're a customer.

We checked five labs' own documentation and the primary legal and policy sources behind it, not press summaries. Two claims that circulate widely didn't hold up. One real, unresolved contradiction turned up in a lab's own pages. Here's what's actually documented, lab by lab.

## Grok: the timeline, opt-out by default

[X's own help documentation](https://help.x.com/en/using-x/about-grok) states plainly that X shares your public posts, engagement data, public Spaces, and public profile information with xAI "to train and fine-tune Grok and other generative AI models," plus your direct interactions with Grok on the platform. It's on by default; opting out means going into Privacy & Safety settings and turning off data sharing yourself.

One widely-circulated claim about Grok 4.6 doesn't check out: that its model card confirms training on "anonymized Cursor" coding sessions, real people's IDE work. Neither [xAI's own Grok 4.6 announcement](https://x.ai/news/grok-4-6) nor [Cursor's own blog post](https://cursor.com/blog/grok-4-6) about the release says anything of the kind. Both describe Cursor strictly as a distribution partner: Grok 4.6 became "available today in Cursor," which is a different claim entirely from "trained on Cursor's user data." We're not carrying that one forward.

What does hold, and is worth being precise about: Elon Musk has publicly claimed SpaceX engineering data feeds Grok's training. Neither primary source we checked corroborates that from xAI's own side. The claim traces to Musk, not to xAI's documentation.

## Claude: the one lab with a priced legal outcome

Anthropic's book-training approach produced the most concrete number of any lab here, because it went through a US court. In *Bartz v. Anthropic* (N.D. Cal., case 3:24-cv-05417-WHA), Judge William Alsup drew a specific line: scanning books Anthropic had physically purchased and destructively rebound was fair use, because, in his words, the books "were still purchased, it was just a matter of storage." Building a permanent library out of pirated copies from Library Genesis and the Pirate Library Mirror was not: "Anthropic had no entitlement to use pirated copies for a central library." Anthropic [settled for $1.5 billion](https://copyrightalliance.org/participating-bartz-v-anthropic-settlement/), an estimated 400,000 to 500,000 pirated titles, a figure that works out to roughly $3,000 per work. That's a real, court-tested number a client can reason about when weighing legal exposure, not a marketing claim.

Worth flagging as an open question rather than picking a side: Anthropic's own pages disagree with each other about how consumer chat training actually works. The [current privacy page](https://privacy.claude.com/en/articles/10023580-is-my-data-used-for-model-training), dated March 16, 2026, frames it as opt-in: "We will use your chats and coding sessions (including to improve our models) if: 1. You choose to allow us to use your chats and coding sessions to improve Claude." But [TechCrunch reported](https://www.techcrunch.com/2025/08/28/anthropic-users-face-a-new-choice-opt-out-or-share-your-data-for-ai-training) in August 2025, seven months earlier, on a policy shift that flipped training to default-on for consumer accounts, with a hard opt-out deadline of September 28, 2025. We found no source reconciling the two. This isn't a one-off research miss on our part either: an earlier, unrelated piece of EP's own research independently hit the same unresolved tension between an Anthropic pricing page and an Anthropic help article. Two independent checks, seven months apart, landing on the same contradiction in how Anthropic documents its own policy. If you're telling a client whether their Claude usage trains the model, verify the current toggle state directly rather than repeating either page as settled.

## GPT: explicit once, vague now

OpenAI has actually named Common Crawl before, just not recently. Its 2020 GPT-3 paper lists "Common Crawl (filtered)" as the single largest component of the training mix by a wide margin, roughly 60% of the total tokens used. That's a real, specific, primary-sourced disclosure. What's changed is the current documentation: we checked the full GPT-5 System Card directly and found zero mentions of Common Crawl anywhere in it. The card's actual language today is deliberately broader: the model is trained on "diverse datasets, including information that is publicly available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate." OpenAI got more specific about its sources when the models were smaller and got vaguer as they got larger, not the other way around, and that trajectory is itself worth knowing if a client asks why current disclosures read thinner than older ones. OpenAI's own crawler, [GPTBot](https://developers.openai.com/api/docs/bots), remains the most explicitly documented crawling mechanism of the five labs here, and site owners can opt out of it via robots.txt.

On the consumer side, OpenAI's [Data Controls FAQ](https://help.openai.com/en/articles/7730893-data-controls-faq) confirms an "Improve the model for everyone" toggle that governs whether your ChatGPT conversations train future models, and can be turned off at any time with no restriction. Worth being precise about Codex here rather than assuming one switch covers all of it: [OpenAI's own documentation](https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance) states that training on regular Codex tasks is turned off through that same Data Controls FAQ mechanism, alongside ChatGPT conversations. But full-environment Codex training is different: it "has separate controls for allowing training on full environments, which you can manage in the Codex Settings," and "adjusting your settings in the ChatGPT interface or privacy portal will not affect these full-environment Codex settings." If you're advising a client on Codex specifically, check the Codex Settings directly for full-environment training rather than assuming the main ChatGPT toggle covers everything.

## Gemini: chats train by default, and one claim about YouTube needs a caveat

Google's [Gemini Apps Privacy Hub](https://support.google.com/gemini/answer/13594961) confirms that with Keep Activity turned on, your Gemini chats are used to train and improve the models by default. Reviewed conversations can be retained for up to three years. Turning on Temporary Chats, or disabling Keep Activity, opts you out.

One claim that circulates about Gemini needs a careful read: that it was trained on YouTube's video corpus. Google's own privacy documentation does describe using a user's Search or YouTube history as personalization context for Gemini's responses, which is a real and confirmed mechanism, but that's a materially different claim from training the underlying model on YouTube's video library. The training-corpus claim specifically traces to a [CNBC report from June 2025](https://www.cnbc.com/2025/06/19/google-youtube-ai-training-veo-3.html), which Google confirmed to CNBC directly at the time, saying it uses "a subset" of its YouTube library and honors its agreements with creators. That's a real, sourced claim, just not one stated in Google's own current privacy documentation, and it predates the pages we checked by over a year. Worth keeping the two straight if you're citing either one: YouTube history as a personalization input is confirmed in Google's current documentation; YouTube's corpus as training data is a confirmed-by-Google-to-a-reporter claim from mid-2025, not something Google states on its own privacy pages today.

## DeepSeek: the one lab that names nothing

DeepSeek's own [disclosure page](https://cdn.deepseek.com/policies/en-US/model-algorithm-disclosure.html) is the shortest and least specific of the five: "publicly available information on the internet" plus data obtained from "third-party data providers" through "legally signed agreements." No named datasets, no named partners, no volume figures. That's not a gap in our research, it's the actual extent of what DeepSeek discloses about itself.

## What this actually means if you're choosing a vendor

- **The legal exposure isn't hypothetical for one lab.** Anthropic's settlement is a real, adjudicated outcome with a real number attached. If a client asks what happens when a lab trains on content it didn't have rights to, that's the concrete answer, not a hypothetical.
- **"Opt-out available" and "opt-out is the current state" are two different claims.** Check the live toggle in your own account before telling a client what their usage does or doesn't train, on any of these five, not just Anthropic.
- **A claim that sounds specific isn't automatically sourced, and a source's own history matters.** The Cursor-training claim sounded precise and didn't hold up to a direct check. The Common Crawl claim is different: it's real, just outdated, from OpenAI's 2020 paper rather than its current documentation. Check both whether a claim is sourced and how current the source actually is before repeating it to a client.
- **Enterprise and commercial terms generally differ from consumer terms** across all five labs, typically excluding customer content from training by default at that tier. If your client's usage is commercial, verify their specific tier's terms rather than assuming the consumer policy above applies.
- **DeepSeek discloses the least, on purpose or not.** If a client's risk tolerance requires knowing what data trained a model, that's the one name on this list where the honest answer is "the company hasn't said."

None of this changes which model performs better on your workload. It changes what you're actually agreeing to when your team's data starts flowing through one.
