快·讯

事件详情 · fbf77747-4f6a-4f3f-8271-f089f9ca3d82

Gemini 4 Argon is here, taking first place on 13 of the 19 benchmarks Google published against GPT-6 Astra and Claude Opus 5.5.

The Rundown AI 公司事件 人工智能
原文 · SOURCE RECORDS
The Rundown AI 10-01 04:05:44 The Rundown AI
原文 · 600 字符(点击折叠)

A big day for Google! Gemini 4 Argon is here, taking first place on 13 of the 19 benchmarks Google published against GPT-6 Astra and Claude Opus 5.5. A few standout numbers: - DeepSWE (long real-world coding tasks) – 77.9%, vs. 74.2% for Opus 5.5 and 74.1% for Astra - Vals Index (real work in finance, law, tax, coding) – 68.9%, vs. 67.0% for Opus 5.5 and 63.1% for Astra - Output limit is 1M tokens, up from 64K Argon also leads on tests for vibe coding, long context, and reading charts and long videos. The release is limited to trusted cyber defenders in Google's Fairwind Program to start.

源站原文 ↗
译文 · CHINESE

谷歌的大日子!Gemini 4 Argon 来了,在谷歌发布的针对 GPT-6 Astra 和 Claude Opus 5.5 的 19 项基准测试中,有 13 项排名第一。几个突出的数字:DeepSWE(长时真实世界编码任务)——77.9%,而 Opus 5.5 为 74.2%,Astra 为 74.1%;Vals Index(金融、法律、税务、编码领域的真实工作)——68.9%,而 Opus 5.5 为 67.0%,Astra 为 63.1%;输出上限为 1M tokens,高于之前的 64K。Argon 在 vibe coding、长上下文、阅读图表和长视频的测试中也领先。该版本最初仅限于谷歌 Fairwind 计划中的可信网络防御者使用。

摘要 · AI SUMMARY

谷歌发布Gemini 4 Argon,在19项基准测试中13项领先GPT-6 Astra和Claude Opus 5.5。DeepSWE得分77.9%,高于Opus 5.5的74.2%和Astra的74.1%;Vals Index得分68.9%,高于两者的67.0%和63.1%。输出上限为1M tokens,较64K提升。该模型目前仅向Fairwind计划中的可信网络防御者开放。

以上为 AI 摘要;完整原文请经由源站链接查阅。

← 返回资讯流