跳到正文

#推理

今日 1 条
今天10月2日周五
10月1日周四
  1. 量子位69

    谷歌发布 Gemini 4 Argon:搭载 RSI,基准成绩领先 Opus 5.5 和 GPT-6

    谷歌发布 Gemini 4 Argon 旗舰模型,搭载 RSI,在 DeepSWE v1.1(77.4%)、CWE-bench v1(68%)、Vals Index(68.90%)等基准上领先 Claude Opus 5.5 和 GPT-6 Astra,并在 Code Arena 排第八。

    推荐理由:文章集中给出 Gemini 4 Argon 与 Opus 5.5、GPT-6 Astra 在 DeepSWE、CWE-bench、Text Arena 等基准上的横向对照和价格差异,读者可据此定位它在前沿模型梯队中的相对位置。

9月16日周三
  1. Google DeepMind68

    Google DeepMind 发布 Gemini 3.8 Live 与 3.8 Live Extended Thinking 语音模型

    Google DeepMind 发布两款实时语音对话模型 Gemini 3.8 Live 与 Gemini 3.8 Live Extended Thinking,前者兼顾成本与规模,后者面向高复杂度任务并支持边推理边说话。

    推荐理由:原文给出了两款实时语音模型的多项基准成绩和面向开发者、企业的接入入口,读者可以据此对比现有同类方案的能力差异。