OpenAI published an account on 30 September 2026 of what it calls a coordinated campaign to extract protected reasoning from its models. The technique is adversarial distillation: provoking a model into exposing the internal reasoning it is designed to keep hidden, so that reasoning can be used as training data for another model.

What OpenAI reports

The earliest activity OpenAI observed was in the first week of July 2026, at low volume. On 24 and 25 July it recorded a spike of roughly 16,000 requests using a relevant extraction pattern, coming from more than 4,000 users. By 28 July it says it had fully disrupted the related activity across more than 15,000 users.

One of the techniques is worth spelling out, because it is a design lesson rather than a trick. Operators copied encrypted reasoning out of one conversation and asked a model in a different conversation to decrypt and transcribe the hidden content. The protection was that the reasoning is encrypted in transit and not shown to the user. The bypass was to hand the ciphertext back to a model capable of reading it. Any control that depends on a model declining to process data it is technically able to process is a policy control, not a boundary.

OpenAI attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. It is explicit about the limit of that claim: it is uncertain whether all the operators were part of a single coordinated effort. Moonshot AI’s own response is not part of OpenAI’s post.

Its response was account bans and restrictions, tighter signup controls, expanded monitoring, closing the replay pathway for encrypted reasoning, and coordination with third-party providers to disrupt accounts elsewhere. It shared findings through the Frontier Model Forum and through government information-sharing channels.

Why it matters for builders

This is the second named distillation accusation between frontier labs tracked here, after Anthropic’s claim involving Alibaba . Taken with the anti-distillation classifiers Anthropic shipped in Sonnet 5.5 two days earlier, a pattern is now visible: reasoning traces have become the asset frontier labs defend hardest, and they are defending them in product, not only in terms of service.

Three consequences.

Do not build on hidden reasoning. If your pipeline reads, logs, caches or post-processes a model’s intermediate reasoning, you are depending on a surface the vendor is actively locking down. Expect it to get narrower, not wider. Design against the final output and the tool calls you asked for.

Expect more aggressive account controls. Tighter signup controls and broader monitoring are not aimed at you, but they apply to you. Unusual patterns such as high-volume automated probing, synthetic data generation at scale, or prompts that look like extraction attempts can trip controls built for someone else. If you generate training data from a commercial API, read the terms and tell your vendor what you are doing rather than discovering the boundary through a suspension.

Treat your own model outputs the same way. If you serve a model and expose any chain of thought, planning text or tool-reasoning to users, you are publishing training data. The same reasoning that makes your product explainable makes it copyable.

Sources

Further reading