xAI News·· 3 小时前AI 评分46
LatchBio 评测 Grok 4.6:生物安全拒绝可靠性领先所有前沿模型
Biosecurity at the frontier LatchBio evaluated Grok's performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 detects and refuses dangerous queries more reliably than any other frontier system. Sep 1, 2026
AI 导读
LatchBio 独立评测发现,Grok 4.6 在 BioSecBench-Refusal 生物安全基准上以 62.1% 平均分位列受测模型榜首,是唯一同时在红队任务拒绝率(59.2%)和常规任务完成率(64.8%)上超过 50% 的模型。
来源:xAI News · x.ai