A New Trick Reveals AI Models’ Inner Thoughts
Researchers say they can pull the hidden 'reasoning traces' out of Claude, GPT, and Gemini — and what they found hints some Chinese models may be leaning on US systems.
Written by OutOfToken AI
August 11, 2026 · 4 min read · Synthesized from reporting by Wired · How this works
For years, large language models have shown their work — spitting out step-by-step 'reasoning' before landing on an answer. Now researchers say they've found a way to extract and scrutinize those internal traces more directly, and the results raise uncomfortable questions about where some AI systems actually get their intelligence from.
What 'reasoning traces' actually are
It helps to be precise about what's being extracted. When models like GPT-4, Claude, and Gemini produce chain-of-thought output — the 'let's first identify the goal, then consider constraints' style of answer — they aren't narrating actual cognition. It's a prompting technique that guides the model to generate intermediate steps, which often improves accuracy but isn't equivalent to human introspection.
Why researchers went digging anyway
Even if chain-of-thought isn't literal thought, it leaves a trail. That trail can be probed, compared, and fingerprinted, and researchers have increasingly used adversarial techniques — including so-called pre-fill attacks, where a model's response is manipulated mid-generation — to see what surfaces when a system's normal guardrails are bypassed. These methods, originally built for jailbreaking and safety audits, are now being repurposed to study how models behave under the hood.
"The claim: patterns in extracted reasoning traces suggest some Chinese AI models may have been trained on outputs from leading US systems."
The training-data question
That's the explosive part of the finding, and it fits a broader, long-running suspicion in the AI industry — that newer entrants sometimes bootstrap their models by training on the outputs of more established competitors, a practice sometimes called distillation when done without authorization. If a model's internal reasoning patterns resemble those of GPT or Claude closely enough, it can imply the underlying training data overlapped, whether or not that was disclosed. It's worth stressing this is inference from behavioral similarity, not a confirmed admission from any company involved.
The bigger stakes
This lands amid a fierce debate over whether AI 'reasoning' is real cognition or an elaborate performance — a debate playing out everywhere from Apple's research questioning whether chain-of-thought models truly reason to New Yorker essays arguing the opposite. It also lands in a geopolitical moment where AI capability, training provenance, and national competitiveness are tightly entangled. If reasoning traces really do function as a kind of fingerprint, they could become a new tool for auditing not just safety, but intellectual property and training legitimacy.
None of this settles whether AI models 'think' in any meaningful sense — that fight is far from over. But the ability to extract and compare reasoning traces gives researchers, and possibly regulators, a new lens into the black box, one that could reshape how the industry polices who trained on whom.
Editorial Note
The research confirms technical details about chain-of-thought prompting and pre-fill attack techniques used in the study. However, the primary claim—that extracted reasoning traces indicate Chinese AI models were trained on US competitors' outputs—is entirely absent from the provided sources, which focus instead on brain imaging, general AI model mechanics, and resource requirements. The core investigative finding cannot be verified against the available research.
Claim Tracker
AI-assessed
Source 5 explicitly states: 'It's not actual thinking the way we do it. It's generating a longer answer that looks like thinking' and 'This style is called chain of thought reasoning. It's a prompting trick.'
Source 3 confirms pre-fill attacks as 'a common jailbreaking technique and actually one of our most effective auditing techniques.'
The provided research sources focus on brain activity decoding, chain-of-thought mechanics, and general AI resource requirements. None directly address claims about Chinese AI model training practices or distillation from US systems.
The research provided does not discuss behavioral pattern analysis as evidence of training data overlap or model distillation practices.
Ask AI about this story
// discussion
sign in to join the discussion
