<think>Let me analyze the article carefully.
**Original title:** "AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision" **Source。
**Original title:** "AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision" **Source。
OpenAI 推出 MentalHealthBench,这是一项由专家参与设计的评测基准,用于评估 AI 在真实心理健康对话中回复的有用性与安全性,帮助衡量模型在心理健康场景下的表现。</think>
1. The original title is "BenchMIRT: What are LLM benchmarks actually measuring?
The article is about MindTopo, a new benchmark from Microsoft Research for testing topological reasoning in multimodal AI models (VLMs - Vision Language Models). Key points。