3. 模型与基准 (Models & Benchmarks)
- AI & LLM Benchmarks 2026: Rankings, Scores & Results
49 minutes ago · AI & LLM Benchmarks 2026 Explore live AI and LLM benchmark rankings across reasoning, coding, math, vision, agents, and tool use. Compare composite indexes first, then open individual evaluations for score provenance, coverage, and methodology.
- 2026年AI大模型综合排行榜 - AnyRank - AnyRank
2 minutes ago · 趣事与总结 本月榜单显示, Anthropic 和 Google 在第一梯队竞争极其激烈,二者在编程能力(Code Arena / SWE-bench)与综合指标上交替领先;同时, 国产大模型 (智谱、月之暗面、阿里巴巴)表现异常强劲,稳居全球前20,并在中文特定语境与性价比上确立了护城河。
- DeepSeek 开源昇腾训练基建 TileLang 对标 CUDA
2 minutes ago · 9月30日, DeepSeek 宣布开源面向华为昇腾平台的基础设施组件,覆盖算子编程语言 TileLang 及其编译支持、计算库和分布式通信库,并强调「所有组件与此前面向英伟达平台的开源组件一一对应」。这意味着支撑 DeepSeek V4 系列模型训练的核心工具链,如今在国产芯片上有了完整的对应版本,华为团队 ...
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- 【AI速览】2026 年 最新国产大模型最新旗舰全维度硬核横评:Kimi / Mi...
1 day ago · 本文基于 2026 年 最细 最新官方 API 定价、SWE-Bench 代码评测、长上下文实测、企业落地案例从最新版本、价格体系、核心性能、侧重点优势、短板、开发适配、个人日常使用体验七大维度横向拆解无软文、纯实测干货文末分三类人群专业开发者、纯小白轻度使用 ...
- anything-llm Review 2026: 8.8/10 (A Grade) | AIToolCrux
anything-llm 是由 Mintplex-Labs 推出的productivity类AI工具,综合评分 8.8/10,等级为 A(优秀)。 完全开源免费,可自托管,数据自主可控。 需要注意的是,自托管需要服务器资源和运维能力。