Threat patterns

How AI privacy risk evolves

Most privacy advice tells you what to do. It rarely tells you when the advice stopped working. These pages track how a class of risk changes, and which of your controls it has already outrun.

A threat pattern is not an incident report. Incidents age badly, and a list of them tells you little about what to do on Monday. A pattern records how a class of risk moves — and specifically how each generation of it defeats the control that stopped the generation before.

That structure is the point. If a control you rely on was designed against generation one, and the thing you are facing is generation three, you are protected against a problem that no longer exists. Each pattern below therefore separates two things carefully:

  • What used to work and no longer does. Recorded so that nobody acts on it. A superseded control presented as current advice is worse than silence.
  • What still works. Usually something organisational rather than a product setting, because settings are what each new generation tends to route around.

How to read the status line

Every pattern ends by saying whether a person has read its sources end to end. Where it says they have not, treat the pattern as a draft position: the facts are cited and dated, but nobody has yet sat down with the underlying documents in full. We would rather publish that distinction than quietly imply a level of checking that has not happened.

Where a pattern refers to litigation, it says what is alleged and what a court has actually decided. Those are different things, and an allegation is not a finding. When we get something wrong we correct it and say so — see Corrections.

Looking for how attackers use AI against you — faked calls to change bank details, cloned voices, redirected payments? That is a different dataset and it has its own page. The patterns below are about what AI vendors do with your data.

A worked example

If the patterns below read as abstract, start with the worked example. It takes four rules a careful firm would actually write into an AI policy and shows each one being defeated by a dated change — including the two that arrived as a dialog box somebody clicked through, and the one that was already wrong on the day it was written.


The bargain nobody reads: free products train on what you type

Training — Content used to improve a model, by default or by drift.
Generation 1 · as at 2026-09-03

How this family has moved:

  1. Generation 1: The bargain nobody reads: free products train on what you type ← this one
  2. Generation 2: Paying does not make you the customer
  3. Generation 3: The default flips under you
  4. Generation 4: Training on what you record stops being a preference and becomes a legal line

A free consumer AI product uses your inputs to improve its models, and says so in terms you accepted without reading.

The control that made this tolerable was disclosure plus choice: it is in the terms, and you can pay for something else.

That control assumed two things which the later generations break - that paying changes the default, and that the default you checked is the default you still have.

What still works:

  • Decide what may be typed into a free AI product at all, and make that an organisational rule rather than an individual judgement.
  • Assume anything pasted into a consumer tier may be used to improve the product, whatever the current setting says.

Why the second one holds and the first did not. The control this generation defeated rested on a fact checked once and then relied on. What still works rests on how you are organised - decided before the moment rather than in it.

Sources:

  • AI Leakage vendor profiles: ChatGPT, Claude, Gemini, Perplexity - all record training on by default on consumer tiers with an opt-out.
  • Read in full 3 September 2026: every source cited above was read directly at the vendor's own documentation, not via third-party summary.

Every source behind this pattern has been read in full.

Paying does not make you the customer

Training — Content used to improve a model, by default or by drift.
Generation 2 · as at 2026-09-03

How this family has moved:

  1. Generation 1: The bargain nobody reads: free products train on what you type
  2. Generation 2: Paying does not make you the customer ← this one
  3. Generation 3: The default flips under you
  4. Generation 4: Training on what you record stops being a preference and becomes a legal line

The paid individual tier carries the same training default as the free one. ChatGPT Plus and Pro, Claude Pro and Max, GitHub Copilot Pro, Pro+ and Max, Perplexity Pro - all are consumer terms with training on by default and an opt-out you must find.

This defeats the most common control people actually rely on, which is not reading the terms but the belief that 'I pay for this, so I am the customer rather than the product'.

It is also where the money argument inverts: on several products the jump to a business tier, where training is contractually excluded, is small. GitHub Copilot Pro to Business is $9 a user a month.

What used to work and no longer does:

  • 'I pay for it, so they are not training on it' - false on every major consumer AI product we rate.
  • 'The paid tier is the safe tier' - the safe tier is the COMMERCIAL tier, which is a different thing and often costs less than the top individual plan.

What still works:

  • For any client or commercial work, use the business or enterprise tier, where training exclusion is contractual rather than a toggle.
  • Price the upgrade before assuming it is expensive. It is frequently cheaper than the premium individual plan people buy instead.
  • If you must use an individual tier, turn the setting off and record who did it and when.

Why the second one holds and the first did not. The control this generation defeated rested on a fact checked once and then relied on. What still works rests on how you are organised - decided before the moment rather than in it.

Sources:

  • AI Leakage vendor profiles, all re-verified 3 September 2026 against vendor documentation: ChatGPT (OpenAI help centre lists Free, Plus and Pro as training-on by default; Business, Enterprise, Edu and API excluded); Claude (consumer tiers on by default, commercial terms excluded); GitHub Copilot (Free, Pro, Pro+ and Max on by default from 24 April 2026, Business and Enterprise excluded).
  • Read in full 3 September 2026: every source cited above was read directly at the vendor's own documentation, not via third-party summary.

Every source behind this pattern has been read in full.

The default flips under you

Training — Content used to improve a model, by default or by drift.
Generation 3 · first observed 2025-08-28 · as at 2026-09-03

How this family has moved:

  1. Generation 1: The bargain nobody reads: free products train on what you type
  2. Generation 2: Paying does not make you the customer
  3. Generation 3: The default flips under you ← this one
  4. Generation 4: Training on what you record stops being a preference and becomes a legal line

A product that did not train on your content starts doing so, and the change arrives as a dialog you click through.

Anthropic announced the change for consumer Claude on 28 August 2025; contemporaneous reporting described the training toggle in the acceptance pop-up as automatically set to on beneath a prominent Accept button. GitHub did the same for Copilot Free, Pro, Pro+ and Max with effect from 24 April 2026.

This defeats the control that survived generation two, which was to check the setting once - at purchase, at procurement, at policy-writing time. A setting checked in March can be a different setting in May without anyone at your end doing anything.

What used to work and no longer does:

  • 'We checked the training setting when we adopted the tool' - the vendor can change the default afterwards.
  • 'It is in our AI policy, which we reviewed at the start of the year' - a policy naming a vendor's posture is a snapshot, and it decays.
  • 'We would have noticed' - these changes arrive as terms-acceptance dialogs, which are designed to be clicked through.

What still works:

  • Re-check the training setting on a schedule rather than at adoption, and treat a terms-update dialog as a trigger to re-check rather than an interruption.
  • Write the policy against the QUESTION, not the vendor's current answer: 'training must be off wherever client data is entered', not 'Vendor X does not train'.
  • Prefer commercial tiers, where the exclusion is contractual and cannot be flipped by a product decision.

Why the second one holds and the first did not. The control this generation defeated rested on a fact checked once and then relied on. What still works rests on how you are organised - decided before the moment rather than in it.

Sources:

  • Anthropic: Updates to our Consumer Terms and Privacy Policy (announced 28 August 2025).
  • TechCrunch, 28 August 2025, reporting the acceptance dialog with the training toggle automatically set to on.
  • GitHub Blog: Updates to GitHub Copilot interaction data usage policy, effective 24 April 2026.
  • Both re-read 3 September 2026.
  • Read in full 3 September 2026: every source cited above was read directly at the vendor's own documentation, not via third-party summary.

Every source behind this pattern has been read in full.