跳到正文

模型能力评测

重点更新

语言对话 · Arena ↗

gemini-3.8-flash-high

关闭详情

google · 语言对话 · 2026-10-08

Arena 评分
1496.5
来源名次
8
对战票数
31,437
评分区间
1491.8–1501.3

偏好评分;区间重叠时差异未必稳定。

模型许可:Proprietary。

原始成绩 ↗
414 个模型
语言对话公开偏好成绩,2026-10-08
模型/机构Arena 评分对战票数
1525.41517–15344,892
claude-opus-5.5-high#2 · anthropic
1507.11499–15156,272
claude-opus-4-6-high#3 · anthropic
1504.41501–150879,415
claude-fable-5-high#4 · anthropic
1503.71499–150840,026
claude-opus-4-7-high#5 · anthropic
1501.01497–150565,192
claude-fable-5.1-max#6 · anthropic
1500.71494–150713,063
claude-opus-4-6#7 · anthropic
1497.71494–150184,095
1496.51492–150131,437
1494.21488–150014,694
claude-opus-4-7#10 · anthropic
1494.01490–149866,315
1492.01484–15006,118
muse-spark-1.1#12 · meta
1490.71486–149539,847
claude-opus-5-high#13 · anthropic
1490.11486–149462,934
claude-opus-5-max#14 · anthropic
1489.41485–149430,569
muse-spark#15 · meta
1488.91483–149514,168
kimi-k3-max#16 · moonshot
1488.11483–149330,351
1487.31484–1490124,675
1486.51481–149224,173
gemini-3-pro#19 · google
1485.51482–148942,078
gpt-5.6-sol-xhigh#20 · openai
1485.11481–148938,490
数据来源

Arena · CC BY 4.0 · 固定版本

用户偏好评分,非正确率。下方小字为置信区间;临时成绩状态见原榜。本站选取 overall 分类、翻译字段并舍入显示,原始值保留。

每日 08:30 同步 · 最近成功:2026/10/11 08:26(北京时间)

历史专项测试:MMLU-Pro / GenEval