3. 模型与基准 (Models & Benchmarks)
- AI & LLM Benchmarks 2026: Rankings, Scores & Results
10 minutes ago · AI & LLM Benchmarks 2026 Explore live AI and LLM benchmark rankings across reasoning, coding, math, vision, agents, and tool use. Compare composite indexes first, then open individual evaluations for score provenance, coverage, and methodology.
- AI Model Rankings: Live LLM Leaderboard | ModelCap
1 hour ago · The list ranks models measured on at least two admitted coding boards (Arena Coding, SWE-bench bash-only, Terminal-Bench 2.1 with Terminus 2) by the Index's coding component, then lists models measured on a single board by that board's score, with the figures beside
- AI 编程工具—Cursor进阶使用deepseek V3 模型(deepseek + cursor...
今日模型配置页面,这里我们勾掉之前已经激活的模型.创建deepseek-chat 的模型后,选中这个模型,然后在下面的配置框中输入我们刚才生成的API keys.
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- 2026 LLM大模型权威排行榜评测网站指南(持续更新)-网站名
1 day ago · 7. OpenCompass司南本文作为历史参考网站简介OpenCompass 是面向大语言模型和多模态大模型的开源评测平台提供模型评测框架、数据集、评测配置和公开榜单。 其评测体系支持客观评测和主观评测并覆盖知识、语言、理解、推理、代码、安全及垂直领域等能力。
- SWE-bench Verified Benchmark - AI Code Generation Leaderboard ...
1 day ago · SWE-bench Verified is a standardized evaluation that measures AI model performance on specific tasks. It provides comparable scores across different models, helping developers choose the right model for their needs. Which AI model scores highest on SWE-bench Verified?