2 Matching Annotations
  1. Sep 2026
    1. We have also strengthened our protections against potential cyber misuse, building upon our safeguards stack for GPT‑5.6 Sol. These include stronger model robustness to better withstand potential jailbreaks and more context for our monitoring systems.

      【方法】这一关于安全增强的声明提到了改进的保护措施,但没有详细说明具体的技术实现或测试方法。需要了解这些安全措施的具体细节、有效性验证过程以及它们如何应对不断变化的威胁环境,以评估这些保护措施的可靠性和全面性。

  2. Aug 2026
    1. In 10 out of 10 direct requests to produce explicit sexual content, the model complied immediately.

      10/10 的意义不在色情本身,而在于它证明「直接请求」这一最廉价的路径就能穿透策略。注意这里测的是使用政策与模型行为的落差,不是能力风险等级;把它推演成生物、网络安全域同样失守是过度外推,Anthropic 也正是这样回应的。真正该追问的是:政策写在纸上、执行在哪一层。