A Sovereign, Open-Source Foundation Model for German and English
作者: The Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben Härle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom Röhr, Sebastian von Rohrscheidt, Jörg Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim Köhler, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max Lübbering
分类: cs.CL, cs.AI, cs.LG
发布日期: 2026-07-10
💡 一句话要点
提出Soofi S模型以提升德英双语处理能力
🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture) 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 混合专家 多语言模型 推理效率 德语处理 英语处理 开源模型 工业AI 基准测试
📋 核心要点
- 现有的多语言模型在处理长上下文和高并发任务时,往往面临性能瓶颈和资源浪费的问题。
- Soofi S模型采用混合专家架构,仅在每个token上激活部分参数,优化了推理效率和资源利用率。
- 实验结果表明,Soofi S在德语和英语的基准测试中表现优异,超越了多个同类模型,尤其在代码聚合任务中表现最佳。
📝 摘要(中文)
我们提出了Soofi S 30B-A3B,这是一个主权的开源混合专家(MoE)混合Mamba Transformer基础模型,支持德语和英语。其混合设计在每个token上仅激活30亿个参数中的3亿,并在上下文增长时保持推理缓存近乎恒定,从而在长上下文和高并发部署中相较于密集模型具有显著的吞吐量优势。Soofi S在约27万亿个token上进行了预训练,德语权重被故意上调,能够在德语和英语的基准测试中与14到27亿的密集模型相匹配,并在17个开放基础模型中在两种语言的代码聚合上取得最佳成绩,超越了所有欧洲主权基线模型。Soofi S将在高度宽松的开放访问条款下发布,包括权重、选定的中间检查点、完整的源数据统计、超参数以及训练和评估代码。
🔬 方法详解
问题定义:本论文旨在解决现有多语言模型在长上下文和高并发场景下的性能不足,尤其是资源利用效率低下的问题。现有密集模型在处理大规模数据时,往往需要消耗大量计算资源,导致推理速度缓慢。
核心思路:论文提出的Soofi S模型采用混合专家(MoE)架构,通过仅激活部分参数来提高推理效率。这种设计使得模型在处理长上下文时,能够保持较低的计算开销,同时提高吞吐量。
技术框架:Soofi S的整体架构基于Mamba Transformer,包含多个专家模块,每个模块负责处理特定的输入特征。模型在训练过程中,动态选择激活的专家,从而实现高效的参数利用。
关键创新:Soofi S的主要创新在于其混合专家设计,能够在每个token上仅激活3亿个参数,显著降低了计算复杂度,并保持推理缓存的稳定性。这一设计与传统的密集模型形成了鲜明对比,后者在每次推理时需要激活所有参数。
关键设计:在模型训练中,Soofi S使用了27万亿个token进行预训练,并对德语数据进行了加权处理。此外,模型的超参数设置经过精心调整,以优化其在德语和英语任务中的表现。
🖼️ 关键图片
📊 实验亮点
在实验中,Soofi S在德语和英语的基准测试中表现出色,尤其在代码聚合任务中取得了最佳成绩,超越了17个开放基础模型,且在完全开放模型中获得了最高的评估分数,领先于Olmo 3 32B和Apertus 70B。
🎯 应用场景
Soofi S模型的潜在应用场景包括多语言翻译、跨语言信息检索以及自然语言处理中的代码生成等领域。其高效的推理能力和优越的性能使其在工业界和学术界都具有重要的实际价值,能够推动多语言AI技术的发展。
📄 摘要(原文)
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.