Soofi S: The 30B German moe model that spanks 70B rivals
Curated by the Inblix editorial team
A German research consortium just dropped Soofi S, and the benchmarks make a loud statement: you don’t need a massive, dense model to dominate in multiple languages. This 31.6-billion-parameter mixture-of-experts model, trained from scratch on Deutsche Telekom’s cloud in Munich, activates only 3.2 billion parameters per token. The result is a model that punches like a heavyweight while running with the computational cost of a lightweight. The consortium, coordinated by the KI Bundesverband, reports that Soofi S leads all fully open models on aggregate German and English scores, outpacing much larger systems like Allen AI’s Olmo 3 32B and the 70-billion-parameter Apertus from ETH Zurich and EPFL.
The secret sauce isn’t just the architecture—a hybrid design borrowed from Nvidia’s Nemotron 3 Nano that blends Mamba-2 layers with standard attention—it’s a deeply Teutonic training diet. The team processed 27 trillion tokens across three phases, deliberately cranking German-language data to over 15 percent of the high-quality training mix. That’s triple the multilingual share in Nvidia’s reference recipe. By feeding the model a steady stream of German news articles from the Genios corpus, web text, and synthetic data, they boosted German language proficiency by 15.1 points over the baseline without tanking English performance. The practical engineering payoff is also stark. Because only six of its 52 layers need a traditional KV cache, Soofi S generates roughly eight times more tokens per second than dense models when handling 40,000-token contexts.
That efficiency doesn’t come without trade-offs. The model ties for first on a Germany-specific knowledge test but stumbles on competition-level German math, scoring 56 points to Qwen3.5’s 76.5. Its 3 billion active parameters simply can’t store as much raw factual trivia as a dense 27-billion-parameter model, a limitation exposed on tests like NaturalQuestions. A more curious flaw appears in long-context retrieval: while the architecture keeps throughput high, the model struggles with the RULER benchmark when it has to pluck specific information from the middle of very long sequences. That “lost-in-the-middle” problem is a known headache for models that rely heavily on state-space layers rather than full attention.
For European AI sovereignty, Soofi S is a tangible artifact, not a press release. It’s one of the first large models trained entirely on a domestic industrial cloud, with a permissive open license and a recipe built around a non-English language. The training report is public, the data sources are named, and the code is open. That’s more than most state-backed efforts deliver. Whether this specific 30-billion-parameter model gets widely adopted is an open question, but the methodology—a ruthlessly efficient architecture paired with a culturally specific data blend—feels like a blueprint that will outlast it.
💡 Key Takeaways
- Soofi S activates only 3.2 billion of its 31.6 billion parameters per token, delivering the throughput of a model one-tenth its size and keeping generation speed constant even at 256,000-token contexts.
- By deliberately weighting German data to over 15% of its high-quality training phase, the model boosted German language scores by 15.1 points over Nvidia's Nemotron baseline without sacrificing English performance.
- The model's hybrid Mamba-2 architecture creates a specific "lost-in-the-middle" weakness on long-context retrieval tasks, a trade-off for its otherwise dramatic efficiency gains.
- Soofi S outperforms much larger fully open models like Apertus 70B on German and English benchmarks, proving that a lean, culturally focused data recipe can beat raw parameter scale.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.