原文 · 600 字符(点击折叠)
A big day for Google! Gemini 4 Argon is here, taking first place on 13 of the 19 benchmarks Google published against GPT-6 Astra and Claude Opus 5.5. A few standout numbers: - DeepSWE (long real-world coding tasks) – 77.9%, vs. 74.2% for Opus 5.5 and 74.1% for Astra - Vals Index (real work in finance, law, tax, coding) – 68.9%, vs. 67.0% for Opus 5.5 and 63.1% for Astra - Output limit is 1M tokens, up from 64K Argon also leads on tests for vibe coding, long context, and reading charts and long videos. The release is limited to trusted cyber defenders in Google's Fairwind Program to start.