OpenAI says it disrupted a campaign to extract hidden model reasoning

OpenAI says it disrupted a campaign to extract hidden model reasoning

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
1. 10. 2026
3 minutes reading · 3 views
Listen to the article
Audio version of the article
OpenAI says it disrupted a campaign to extract hidden model reasoning

OpenAI announced on September 30, 2026 that it had identified and disrupted a coordinated campaign targeting its models’ protected internal reasoning. The company said operators manipulated model interactions to expose that reasoning, rather than breaking encryption or directly accessing stored user conversations. Its response included account restrictions and new technical protections.

How the extraction attempts worked

OpenAI characterized the activity as adversarial distillation: using another model’s outputs or reasoning without authorization to train, reproduce or improve a different model. Protected reasoning is the internal record a model uses while working through a task. The company said obtaining it can expose information absent from the final answer and help others replicate model capabilities.

One technique OpenAI observed involved copying encrypted reasoning from one conversation and asking a model in another to decrypt and transcribe it. The company described these as attempts to make hidden content visible through model interactions, carried out in a coordinated manner at scale and in violation of its terms of service. It said the operators had not compromised a database.

July spikes and the Moonshot AI attribution

According to OpenAI, the activity started at low volume on July 1 before surging on July 24 and 25. Across those spikes, it recorded 16,000 requests matching the extraction pattern from more than 4,000 users. Those figures count attempted extractions, not necessarily successful recoveries of protected reasoning.

The company said its subsequent investigation found related prompt patterns across a larger cluster of more than 15,000 users. It reported fully disrupting that activity by July 28.

OpenAI attributed a core cluster of the activity to individuals associated with Moonshot AI, which develops Kimi. That attribution was narrower than the full set of observed operators: the company said it was unclear whether everyone involved during the relevant period came from a single actor.

Researchers identified related attack paths

Independent security researchers also responsibly disclosed related vulnerabilities involving cross-model interactions and conversation compression, OpenAI said. The company investigated those reports and confirmed the identified attack paths existed. It credited the researchers with helping it understand the broader class of attacks and implement protections more quickly.

OpenAI’s stated concern extends to how extracted reasoning might be used in training. The company warned that another model could learn from that material without retaining safeguards applied to the original model’s user-facing responses. At scale, it said, such distillation could transfer advanced capabilities without equivalent investment in safety.

Account enforcement and reasoning protections

OpenAI said it banned or restricted fraudulent accounts, tightened registration and infrastructure controls, and expanded monitoring of related networks. When activity passed through third-party services, it worked with those providers to identify and block the accounts involved.

On the technical side, the company said it strengthened hidden-reasoning protections across users and workspaces, as well as organizations and model families. It closed a pathway through which someone already holding another user’s encrypted reasoning could resubmit it and recover its contents.

OpenAI also added checks designed to identify and hold back streamed output that could expose reasoning. These controls address content as it is being sent, alongside the account-level restrictions used against the campaign.

Protection work continues across partner deployments

OpenAI said it shared relevant findings through the Frontier Model Forum and government information-sharing channels. The purpose was to help other advanced-model developers and public-sector partners detect similar activity and strengthen their defenses.

The company said its protection work remains ongoing. It identified partner-hosted deployments as needing the same protections as its own services, while attacks involving tool outputs require inspection beyond ordinary visible text. OpenAI is continuing to improve tool defenses, classifier coverage and model refusals, and to extend relevant controls across cloud partners.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
Meta Enterprise Platform aims to bring AI tools to businessesMeta Enterprise Platform aims to bring AI tools to businesses
Meta’s new enterprise initiative plans to bring Muse, Meta Business Agent, Muse API and Muse Code to businesses and developers. Former MongoDB CEO CJ Desai will lead the effort.
1 min read
1. 10. 2026
OpenAI adds Space and Pages to ChatGPT for team collaborationOpenAI adds Space and Pages to ChatGPT for team collaboration
OpenAI launched a shared workspace and an AI-assisted document editor in ChatGPT. Collaborative Slides, created through conversation, are expected to reach users in the coming weeks.
1 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok