OpenAI says it disrupted a coordinated effort this summer that sought to reproduce protected model reasoning using a technique known as adversarial distillation. The company described the activity, the containment steps it took, and the information it has shared with peers and public-sector partners. The disclosures aim to help other developers detect and defend against similar extraction attempts.
Adversarial distillation, as OpenAI described it, involves manipulating live model interactions so that a model’s hidden internal notes—what OpenAI calls protected or hidden reasoning—can be elicited or replayed in forms visible to a requester. Rather than breaking encryption or accessing stored conversations, the campaign targeted the model’s runtime behavior to coax out internal artifacts and reconstruct them in a way that violated OpenAI’s terms of service.
The company provided a concise timeline and scale for the activity. OpenAI says the activity began in the first week of July at low volume, then spiked on July 24 and 25 when roughly 16,000 requests using a relevant extraction pattern came from more than 4,000 users. Broader prompt-pattern activity was later observed across a cluster exceeding 15,000 users. OpenAI reports that it fully disrupted the related activity by July 28.
OpenAI emphasized that the actors did not decrypt stored data or compromise back-end systems. Instead, the campaign relied on tactics such as copying encrypted reasoning from one conversation and prompting a model in another conversation to decrypt and transcribe that content. Independent security researchers published related findings (arXiv:2608.09867) and shared them through responsible disclosure; OpenAI says those reports accelerated the company’s mitigation work and helped confirm the attack paths.
On attribution, OpenAI said it was unclear whether all observed operators were part of a single coordinated group. However, the company attributed a core cluster of the activity to individuals associated with Moonshot AI, the developer of the Kimi model. OpenAI did not claim that the activity involved technical compromises of its infrastructure; rather, the attribution relates to patterns of behavior and connections among the accounts it investigated.
OpenAI framed adversarial distillation as a cross-industry safety and security concern. Extracted hidden model reasoning could be used to train other systems without the originating provider’s safeguards, potentially accelerating the transfer of advanced capabilities at scale. The company warned that such risks grow as frontier models become more capable and noted the techniques are not unique to its systems, prompting information-sharing with other developers and industry groups.
To counter the campaign, OpenAI applied a layered response combining account enforcement, technical controls, and partner coordination. Reported actions included banning or restricting fraudulent accounts, tightening signup and infrastructure controls, and expanding monitoring for related networks. The company said it strengthened protections for hidden reasoning across users, workspaces, organizations, and model families.
Specific mitigations OpenAI described included closing a pathway that allowed someone with another user’s encrypted reasoning to replay and recover its contents, and adding checks to detect and hold streamed outputs that might expose reasoning. When related activity traversed third-party services, OpenAI worked with those providers to identify and disrupt the accounts involved. The company also shared findings through the Frontier Model Forum and appropriate government information-sharing channels so other frontier developers and public-sector partners could search for similar activity and harden defenses.
Looking ahead, OpenAI cautioned that adversarial distillation attempts are likely to become more sophisticated as models advance and as actors seek lower-cost ways to reproduce capabilities. The company said defending against these attacks requires layered, adaptive controls and continued monitoring. Ongoing priorities it highlighted include improved tool defenses, expanded classifier coverage, clearer model refusals, and propagation of protections across cloud and partner-hosted deployments.
OpenAI summarized its continuing response around three focal areas: stronger technical protections against extraction, better detection and enforcement of coordinated campaigns, and deeper threat-information sharing across industry and government. The company said it remains engaged in mitigation and investigation work to refine those defenses and to help other developers identify and disrupt related activity.
The disclosure marks a detailed public accounting of a novel threat vector against large models. By naming the technique and outlining both how the campaign worked and how it was interrupted, OpenAI aims to raise awareness and promote collective defenses against attempts to extract hidden model reasoning. The company’s account underscores that protecting model internals now involves not only securing data at rest and in transit, but also monitoring and hardening models’ runtime behaviors against adversarial manipulation.
Source: Read the original source

Leave a Reply