
Why DeepSeek-V4.1-Flash Is Such an Exciting Open Model Release
DeepSeek-V4.1-Flash is an open-source AI model utilizing a Causal Encoder-Decoder architecture, MoE, KV cache compression, and CSA2 to achieve high efficiency.

DeepSeek-V4.1-Flash is an open-source AI model utilizing a Causal Encoder-Decoder architecture, MoE, KV cache compression, and CSA2 to achieve high efficiency.
1 editorial report · 4 verified social mentions. The most authoritative report leads while later evidence completes the story.
🇨🇳 Chinese AI labs allegedly had Claude answering for them DeepSeek and Moonshot allegedly found a clever shortcut: instead of relying entirely on their own models, they quietly routed some customer prompts to Anthropic’s Claude through fraudulent accounts. Users thought they were chatting with Chinese AI, while Claude was actually doing some of the answering. Anthropic says Moonshot forwarded nearly 300,000 customer requests in just 10 days, while DeepSeek reportedly routed selected users to Claude Opus when they were using coding tools such as Claude Code and OpenCode. Those prompts could contain highly sensitive information. Anthropic says one conversation involved a user linked to the Chinese military asking about behavior captured across hundreds of CCTV cameras in Chengdu. Another allegedly contained live credentials for a Russian government database shared by an IT operator working with a Defense Ministry-linked agency. According to Anthropic, the exchanges were also saved and used to help train the Chinese models. So this wasn’t just model copying. It potentially turned someone else’s AI into a silent data pipeline. @aipost 🏴
Open mentionmy question about LLM performance We see a lot of posts about token prediction, token generation per second, etc. But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22–28 TPS, but I also see that the LLM does a lot of reasoning. And this relates to others. So maybe the correct way is not to check TPS or other metrics, but to check execution: task complexity/second. I don't know if such a metric already exists and if it exists why community doesn't use it by default submitted by /u/AleksandrNikitin [link] [comments]
Open mention🚀 GigaChat 3.5 Reasoning — a new open-source LLM that thinks before it answers. It breaks problems into stages, builds a plan, checks intermediate results, and self-corrects. Built on GigaChat 3.5 Ultra, it explores multiple step-by-step reasoning paths for math & coding, using automated verification to reinforce correct answers. ⚡️ Proprietary linear attention makes it highly efficient on long contexts, retaining key points without re-matching from scratch. It’s also token-efficient: uses 37% fewer tokens than DeepSeek V4 Flash Preview on math problems! 📈 Benchmark gains over non-reasoning version: • IFBench: 44 → 77 • Natural Plan: 64 → 80 • LiveCodeBench v6: 56 → 85 📦 MIT license. Weights on Hugging Face: fp8 | bf16
Open mention❓ Tired of LLMs that hallucinate on complex tasks? What if your AI could: • Explore multiple step-by-step reasoning paths? • Use automated verification to reinforce correct answers? • Check its own work and self-correct? • Autonomously decide when to call external tools? GigaChat 3.5 Reasoning does all of this. This open-source LLM (built on GigaChat 3.5 Ultra) actually thinks before answering. It uses proprietary linear attention for efficient long-context handling, consuming 37% fewer tokens than DeepSeek V4 Flash Preview on math problems. Real-world performance: ✓ IFBench: 44 → 77 ✓ Natural Plan: 64 → 80 ✓ LiveCodeBench v6: 56 → 85 🔗 MIT License. Weights on Hugging Face: fp8 | bf16
Open mention