There is a question people ask after they have already pasted the contract in: does this thing learn from what I just gave it? The answer depends entirely on which product you opened, and the defaults are not consistent across the industry.

Everything below was read from each vendor’s own privacy or data controls documentation on 7 August 2026. Terms change. Check yours.

Training is on by default

ChatGPT. OpenAI states consumer chat data is used to improve its models by default. The control is under Data controls, labelled Improve the model for everyone.

Gemini. Google states that with the Keep Activity setting on, chats are used to train its generative AI models. Turning it off stops future chats being used. Note the caveat Google publishes alongside it: even with the setting off, or in temporary chats, conversations are still processed for safety and security, including by human reviewers.

Microsoft Copilot. Consumer conversation data is used for AI model training by default. Opt out under Privacy, using the toggle named Training on conversation activity, with a separate toggle for voice.

GitHub Copilot. This is the one most likely to catch people out. As of 24 April 2026, the setting named Allow GitHub to use my data for AI model training is enabled by default for individual subscribers on Copilot Free, Pro, Pro+, and Max. It covers inputs, outputs, code snippets, and surrounding context. Business and Enterprise seats are excluded from training entirely. Individual seats are not, and the setting lives in individual Copilot policy settings rather than anywhere obvious.

Grok. xAI states it may use prompts, searches, and responses for training, with an opt out under Settings, Data Controls, Improve the model. Its privacy policy also tells users plainly not to include personal information in prompts, which is worth taking at face value.

DeepSeek. The strictest trade on this list. DeepSeek states it collects prompts, uploaded files, and photos and uses them to train its models, stores personal data in the People’s Republic of China, and provides a training opt out only by emailing privacy@deepseek.com. There is no consumer paid tier with better terms to upgrade to. The underlying weights are MIT licensed, so self hosting avoids the question entirely.

Midjourney. Its terms grant a perpetual, worldwide, irrevocable licence over content users input, and generations are public and remixable by default. Privacy requires Stealth Mode, which is restricted to the Pro and Mega tiers.

Training is off by default

Claude. Anthropic states consumer data is not used for model training unless the user opts in through Model Improvement Privacy Settings, and that Incognito chats are excluded even when that setting is on. This is the strongest default position among the major assistants.

Cursor. Cursor states it does not use inputs or suggestions to train its models, and does not permit third parties to, except where flagged for security review, explicitly reported by the user, or explicitly agreed to.

Notion. Notion states that customer data is yours and that its AI subprocessors are prohibited from using it to train models. Enterprise plans add zero data retention with the underlying model providers.

Three things worth knowing

Business and consumer terms are different products. Several vendors that train on consumer data exclude business and enterprise seats entirely. If your employer pays for the seat, you may already be covered. If you are on a personal plan doing work, you probably are not.

Opt outs are usually not retroactive. ElevenLabs states plainly that its opt out applies only to data provided after opting out. Assume the same elsewhere unless a vendor says otherwise.

Not training is not the same as not retained. Most providers retain conversations for a period for abuse monitoring regardless of the training setting. Zero retention is a separate, usually enterprise, commitment.

The five minutes worth spending

Open the data controls on every AI product you use for work and read what the default is. If you are an individual GitHub Copilot subscriber, do that one first. Then decide, deliberately rather than by accident, which of your work you are comfortable contributing to somebody else’s model.