OpenAI says it spent July fighting a coordinated effort to pull the private reasoning out of its AI models, and that it ties a core cluster of that activity to individuals associated with Moonshot AI, the developer of Kimi. All of this comes from OpenAI’s own security post, published 30 September. No independent party has confirmed it.
The technique at issue is distillation: using one model’s outputs, and especially its step-by-step thinking, to help train or copy another. OpenAI calls its version adversarial distillation, meaning organized use without permission. The prize is what the company calls protected reasoning, the model’s internal working-out. OpenAI says that record can reveal information the final answer leaves out and help outsiders reproduce what the model can do.
This was not a break-in, and OpenAI is explicit on the point. Per the post, its encryption held, no database was compromised, and stored user conversations were never reached directly. What the operators did, according to OpenAI, was manipulate ordinary model interactions so that hidden reasoning surfaced in a form the requester could read. In one method OpenAI describes, operators lifted encrypted reasoning from a single conversation, then opened a second one and asked a model there to unscramble it and write it out as text. OpenAI adds that this kind of manipulation is “not a vulnerability unique to OpenAI’s models.”
The numbers need careful reading. OpenAI dates the start to 1 July, when volume was small. On 24 and 25 July it jumped: 16,000 requests that fit an extraction pattern, sent by more than 4,000 users. A wider cluster of related prompt activity covered more than 15,000 users, and the company says it had fully disrupted the campaign by 28 July. A footnote states that these figures describe “attempted, not necessarily successful, extractions.” OpenAI does not say how much reasoning, if any, was actually recovered.
The attribution is deliberately narrow. OpenAI writes that it is “unclear whether all operators we observed during the relevant time period originated from a single actor,” while attributing “a core cluster of the activity to individuals associated with Moonshot AI.” That is a subset of the activity and a group of individuals, not a finding against the company as a whole. The account OpenAI published includes no reply from Moonshot.
OpenAI frames the stakes as safety and national security. Its argument: reasoning copied this way can train a new model without the safeguards that shaped the original’s answers, and doing it at scale moves advanced capabilities across without matching investment in safety. The company says the concern grows as models gain uses that can serve civilian and military ends alike.
Its response mixed cleanup with hardening. Fraudulent accounts were banned or limited. The company also tightened how people sign up and how infrastructure is accessed, and it widened its monitoring. A route that let a person with someone else’s encrypted reasoning play it back and get the contents out has been shut. New checks hold back streamed output likely to expose reasoning. Outside security researchers had disclosed related flaws on their own; OpenAI says it investigated them and confirmed they were real. Findings went to the Frontier Model Forum and to government information-sharing channels.
The commercial subtext is plain. A lab’s hidden reasoning is a competitive asset, and a rival that can harvest it cheaply skips years of expensive research. Because OpenAI is both the accuser and the only source, the dispute over who did what will not be settled by this post alone.
For any company serving frontier models through partners, OpenAI’s message is that hosted deployments need the same protections as the original service, so that is where the next round of extraction attempts will probably look.
Reported by OpenAI in its own security post on 30 September 2026.