3. 模型与基准 (Models & Benchmarks)
- 2026年AI大模型综合排行榜 - AnyRank - AnyRank
2 minutes ago · 趣事与总结 本月榜单显示, Anthropic 和 Google 在第一梯队竞争极其激烈,二者在编程能力(Code Arena / SWE-bench)与综合指标上交替领先;同时, 国产大模型 (智谱、月之暗面、阿里巴巴)表现异常强劲,稳居全球前20,并在中文特定语境与性价比上确立了护城河。
- 模型不是胜负手:2026 年中 CLI 编码 Agent 的 Harness 之争
2 minutes ago · Terminal-Bench 2.1 上 Claude Code 与 Codex CLI 的胶着、同一模型在不同 Scaffold 下的分数漂移、SWE-bench Pro 统一脚手架与厂商自测分数之间的巨大鸿沟,都在指向同一件事:选 Agent 时,先看它能不能在你的仓库里稳定地读文件、跑测试、压缩上下文、在失败时正确重试 ...
- AI 编程扩展 devin-byok-plus 更新 V2.4.0:支持接入 DeepSeek 与 Kim...
2 minutes ago · 事件分析 从技术视角看,devin-byok-plus 的更新体现了 AI 编程工具领域正在发生的“去耦合化”趋势。开发者不再满足于单一厂商提供的封闭式 SaaS 体验,而是倾向于将“编辑器交互层”与“大模型推理层”分离。DeepSeek 和 Kimi 等国产大模型 API 的接入支持,反映了开发者在追求高性能推理的同时,对 ...
- LLM News - 最新AI/ML开源仓库、研究论文动态
2 minutes ago · About LLM News LLM News is an automated tracker for AI/ML developments, focusing on Large Language Models, AGI, and related technologies.
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- Arena Elo Benchmark - AI Model Leaderboard (2026) | LM Market Cap
1 day ago · LMSYS Chatbot Arena Elo Rating: Human preference rating from 6M+ crowdsourced blind head-to-head comparisons. Users chat with two anonymous models and pick the better response. See which AI models score highest on Arena Elo. Updated rankings with scores from 350+ mode