3. 模型与基准 (Models & Benchmarks)
- 微软Build 2026全面解读:MAI-Thinking-1推理模型AIME 25达97%、7款自...
3 minutes ago · 2026年6月2-3日微软Build 2026开发者大会全面解读:7款全自研MAI模型发布,MAI-Thinking-1推理模型AIME 25达97%、SWE-Bench Pro 53%,MAI-Code-1-Flash 5B参数碾压同级别模型,微软AI战略从依赖OpenAI转向全面自主,对开发者和AI工具生态的深远影响。
- DeepSeek-V4-Pro - AI模型价格对比 (2026/7/30)
3 minutes ago · DeepSeek V4 Pro 是 DeepSeek 推出的大规模混合专家模型,总参数量为 1.6T(万亿),激活参数量为 49B(十亿),支持 100 万 Token 的上下文窗口。 该模型专为高级推理、编程以及长周期智能体工作流而设计,在知识、数学和软件工程基准测试中均表现出色。
- 『DeepSeek』 V4正式版官宣7月中旬上线,引入峰谷定价机制 (deepseekv...
3 minutes ago · 『DeepSeek』-V4 模型预览版于 4 月 24 日上线并同步开源。 『DeepSeek』-V4 拥有百万字超长上下文,在 Agent 能力、世界知识和推理性能上均实现国内与开源领域的领先。 模型按大小分为两个版本: 省流:涨价了↓↓↓ ads a2 机制 大小 版本更新 『DeepSeek』 价格表
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- 開源 AI 模型比較 2026:Llama、DeepSeek、Taide... | LargitData Blog
開源大型語言模型(LLM)持續演進,2026年企業採用率已超過80%。 Meta的Llama 4採原生多模態架構,Maverick版支援100萬token上下文;DeepSeek V3.2引入稀疏注意力機制,代理推理能力比肩GPT-5,以MIT授權開源...
- SWE-bench Verified Benchmark - AI Code Generation Leaderboard ...
1 day ago · SWE-bench Verified is a standardized evaluation that measures AI model performance on specific tasks. It provides comparable scores across different models, helping developers choose the right model for their needs. Which AI model scores highest on SWE-bench Verified?