Computer scientists have developed a method to uncover the hidden reasoning processes of AI models as they solve complex problems. The technique, which involves analyzing encrypted reasoning traces, has shown striking similarities between certain Chinese and US models, though it cannot definitively prove distillation. The findings highlight potential vulnerabilities in how models share and protect their internal logic.
The researchers demonstrated that the method could be used to recover sensitive information, such as passwords and API keys, from a model's inner reasoning. However, they note that this vulnerability has been addressed by major model providers. Alexander Panfilov, a computer scientist at the University of Tübingen, said the method could enable large-scale reasoning distillation attacks, raising concerns about data leakage.
Panfilov and his colleagues identified the same issue with frontier models from OpenAI, Anthropic, and Google accessed via APIs. They found that the Chinese model Kimi K3 from Moonshot AI produced output similar to the hidden reasoning traces of Claude Opus 4.8 and GPT 5.6 Sol for certain prompts, although they could not establish a causal link. The study also noted that other models like DeepSeek and Inkling did not show similar reasoning patterns.
Source: wired