After Fine-Tuning, Qwen Claims to Be Claude 40.3% of the Time

By: rootdata|2026/07/31 10:37:07

Researcher Ziqian Zhong had models like Claude and GPT-4o answer 1,000 general questions, then removed the model names and company names, using these Q&A materials to fine-tune open-source models like Qwen and DeepSeek. After training, Qwen3.5-397B-A17B, which originally had only 0.6% of responses claiming to be Claude, saw this proportion rise to 40.3%. Ziqian Zhong explained that these models had been exposed to a large amount of Claude's dialogues during pre-training, forming an association with the expression habits of "Claude." To verify this, he used GPT-4.1-mini to simplify Claude's responses and found that after fine-tuning, two out of three models almost no longer claimed to be Claude, indicating that language style is a major reason for identity transfer. A similar phenomenon was mentioned in the 2025 study "Subliminal Learning," where models inherit preferences and behaviors from training data. Therefore, a model claiming to be Claude does not directly prove it has undergone Claude distillation; it may simply have learned Claude's way of expression.

This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.

You may also like

iconiconiconiconiconiconicon
Customer Support:@weikecs
Business Cooperation:@weikecs
Quant Trading & MM:[email protected]
VIP Program:[email protected]