Thread Reader
leo ๐Ÿพ

leo ๐Ÿพ
@synthwavedd

Jul 18
8 tweets
Tweet

๐Ÿงต๐Ÿงต DeepSeek appear to have engaged, or be engaging in, a large-scale operation to collect outputs from proprietary models (including Claude Fable 5) for certain requests via their API as part of a distillation effort. After seeing such claims circulating earlier today, we conducted an investigation into them on our Discord. We found that, when "Deepseek V4" was used within OpenCode - via their official API - for complex prompts (i.e. 3D games) and combined with a knowledge-related query, the model provides virtually identical outputs to Fable 5. CoT structure is also very different from what is typically expected from Deepseek models. Both of these behaviours revert to what is expected for V4 when simpler prompts were used. When "Deepseek V4" was asked to incorporate answers to questions related to cyber or bio tasks that we verified hit Fable's classifiers into its 3D games, outputs tanked in quality. This is very difficult to explain unless the request was routed to Fable and fell back after hitting a classifier. For complex code prompts without anything else mixed in, like 3D games, outputs were remarkably similar to those produced by Fable 5. We were able to produce these results most consistently via OpenCode and the official Deepseek API combined with a prompt that specifies a complex code task. Deepseek have continued to modify their routing system since, as we have observed changes in behaviour and CoT style compared to those seen previously. Our investigation was conducted from 7AM-8AM PT.

When we initially asked the model in OpenCode a combined knowledge query, the output was consistent with V4's cutoff of May 2025.
But when we spliced the same prompt with a prompt for a 3D game, suddenly the answer was very different. The model now gave a virtually identical response to Fable's.
When tested on a cyber question that was confirmed to hit Fable's classifiers, the model once again reverted to the expected V4 behaviour, with no knowledge of anything after May 2025.
When we gave the model the 3D archery game prompt, something that outputs circulating from the supposed V4 GA indicate it should do well on, and combined it with cyber/bio queries that hit Fable's classifiers... the result speaks for itself.
Now here's the output with the same prompt but without any bio/cyber stuff mixed in. Like night and day. Weird.
And here's Claude Fable 5's output on the same prompt. It messed up the colours here, but everything else is eerily similar.
Anthropic have recently been cracking down on distillation attacks by testing various CoT visibility levels, including no CoT and single-sentence summaries, and I expect these changes to roll out more broadly soon. While you are free to make your own conclusions from the evidence provided, we believe all of it points to Deepseek operating in a deceptive manner to scrape output data from Claude models, and as such that it is best to publish our findings publicly. Thanks to everyone in our Discord - discord.gg/rag-tag! - for their help in this, particularly dzdsa33.
leo ๐Ÿพ

leo ๐Ÿพ

@synthwavedd
tech, ai & politics nerd || got info you think i'd be interested in? let's talk! sywv.tips@proton.me || dono: https://t.co/WxQgF8aVkI
Follow on ๐•
Missing some tweets in this thread? Or failed to load images or videos? You can try to .