跳到正文

#语音

今日 2 条
今天10月2日周五
9月23日周三
  1. Google DeepMind69

    Google DeepMind 发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS 语音生成模型

    Google DeepMind 发布 Gemini 3.8 Flash TTS 与 Gemini 3.8 Flash-Lite TTS 两款语音生成模型,支持通过自然语言提示从零创建自定义声音、基于 30 秒样本复制人声,并提供 2000+ 现成声音库与 100+ 语言覆盖。

    推荐理由:原文详细列出了两款新 TTS 模型的声音设计、复制与多语言能力,并给出 Hume AI 基准成绩,可对照此前 3.1 Flash TTS 评估实际提升幅度。

9月16日周三
  1. Google DeepMind68

    Google DeepMind 发布 Gemini 3.8 Live 与 3.8 Live Extended Thinking 语音模型

    Google DeepMind 发布两款实时语音对话模型 Gemini 3.8 Live 与 Gemini 3.8 Live Extended Thinking,前者兼顾成本与规模,后者面向高复杂度任务并支持边推理边说话。

    推荐理由:原文给出了两款实时语音模型的多项基准成绩和面向开发者、企业的接入入口,读者可以据此对比现有同类方案的能力差异。