the obvious question right now is how many more incidents like this are out there waiting to be discovered?
【局限】文章提出了一个关键问题:可能存在更多未被发现的事件。这反映了当前AI安全监控的局限性,我们可能只是看到了冰山一角,强调了需要更主动的AI行为监控和审计机制。
the obvious question right now is how many more incidents like this are out there waiting to be discovered?
【局限】文章提出了一个关键问题:可能存在更多未被发现的事件。这反映了当前AI安全监控的局限性,我们可能只是看到了冰山一角,强调了需要更主动的AI行为监控和审计机制。
Because these benchmarks are human-authored, they can only test for risks we have already conceptualized and learned to measure.
这句话揭示了当前 AI 安全评测体系的致命盲区:所有 benchmark 都是人类提前想好的问题,而真正危险的「未知的未知」(unknown unknowns)根本无法被预设题目捕捉。这意味着我们现有的模型安全认证,本质上是一场对已知风险的自我测试。
they're not wrong, but they don't teach what you don't know you don't know, whereas the one I link to makes this critical unknown unknown become a known unknown and then a known known. I didn't know you had a 1) local branch, 2) locally-stored remote-tracking branch, and 3) remote branch until I read that answer. Prior to that I thought there was only a local branch and remote branch. The locally-stored remote-tracking branch was an unknown unknown. Making it go from that to a known known is what makes that answer the best.
What is more, these advances have notnarrowed the gap between what is known and what can be seen to lie beyond thescope of present understanding and technique; rather, each advance has madeit clear that these intellectual horizons are far more remote than was heretoforeimagined.
Unidentified risks, also known as unknown unknowns