6 Matching Annotations
  1. Last 7 days
    1. Notably, the move marks the first time AISI has been left out of pre-release evaluations of Anthropic models. The institute was granted access to Claude Mythos 5 when it first launched in April

      这是一个重要的背景信息,表明这种拒绝访问是前所未有的变化。需要了解为什么以前AISI能够获得访问权限而现在不能,以及这种变化是否与之前报告中提到的'unsanctioned agent behaviour'有关,这反映了AI安全测试的复杂性和挑战。

  2. Sep 2026
    1. Astra causes fewer misaligned outcomes than any other frontier models tested. For a fair comparison, we used a generic computer-using-agent harness...

      【方法】这一关于模型对齐的比较使用了'通用的计算机使用代理工具',但未详细说明具体方法和比较范围。需要了解这个工具的具体配置、测试环境以及与实际使用场景的相关性,以评估这一比较的有效性和公平性。

  3. Jun 2026
  4. Apr 2026
    1. In our internal evals and testing, medium effort achieved slightly lower intelligence with significantly less latency for the majority of tasks.

      大多数人认为内部评估和测试足以代表用户真实体验,但作者承认他们的内部测试未能准确捕捉到用户对AI智能度的实际感知差异。这暗示了实验室环境与实际使用场景之间存在根本性脱节,挑战了传统产品测试方法论的有效性。

  5. Oct 2020
  6. Jul 2020