跳到正文
原文
The Decoder· Matthias Bastian·· 4 小时前AI 评分41

Reka AI 全模态模型 Rho-1 在单一模型中统一处理文本、图像、视频与机器人控制

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

AI 导读

Reka AI 发布 Rho-1 研究预览版,这是一个 19B 参数的全模态模型,可在单一神经网络中处理和生成文本、图像、视频及机器人控制动作。所有模态以 token 形式共享同一上下文窗口,无需调用外部模型或工具,并支持实时连续视频生成。训练使用 320 块 H100 GPU,历时约三个月;为缓解机器人训练数据稀缺,团队还构建了从普通互联网视频中提取控制信号的反向动力学模型。

来源:The Decoder · the-decoder.com