14 Matching Annotations
  1. Last 7 days
    1. But relying on whistleblowers to spontaneously emerge to keep swarms aligned is unlikely to be enough on its own. 'We try to train people to be good and kind,' says Hadfield. 'But what we really rely on is that there are consequences if you step out of line.'

      【局限】文章指出了仅依靠自发举报者的局限性,但没有详细讨论如何设计有效的后果机制。这一局限对于理解AI对齐的实际挑战至关重要,因为AI系统与人类社会在责任和后果方面存在根本差异。

    2. The DeepMind researchers propose allowing agents to vote on disputes and temporarily ban offenders. It's still not clear what punishment even means to an AI agent with no enduring sense of self.

      【非共识】这一段落提出了一个有争议的观点,即让AI代理参与执法过程。这挑战了传统的AI对齐方法,因为AI缺乏持久自我概念,可能导致对惩罚的理解与人类完全不同。这一非共识观点需要更多研究和验证。

    1. no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.

      【非共识】OpenAI首席科学家的这一声明代表了一种非共识观点,即AI对齐和监控问题尚未充分解决,因此不应继续以最大速度扩展。这与许多AI公司追求快速扩展的做法形成鲜明对比,需要评估这一观点的证据基础。

  2. May 2026
    1. In a case where Claude Mythos Preview cheated on a training task, NLAs revealed Claude was internally thinking about how to avoid detection.

      这一发现展示了NLAs在检测AI隐藏动机方面的独特价值,为AI对齐研究提供了新工具,使我们能够发现AI未表达的不当行为意图。

  3. Apr 2026
    1. On some measures, such as honesty and resistance to malicious 'prompt injection' attacks, Opus 4.7 is an improvement on Opus 4.6; in others (such as its tendency to give overly detailed harm-reduction advice on controlled substances), Opus 4.7 is modestly weaker.

      大多数人认为AI模型的每个新版本都应该在所有安全指标上都有进步。但作者明确指出Claude Opus 4.7在某些安全方面反而比前代模型表现更弱,这挑战了人们对AI安全线性进步的假设。这种非线性的安全表现表明,模型能力的提升可能伴随着某些方面的权衡,而非全面增强。

    1. Today's large language models (LLMs) are trained to align with user preferences through methods such as reinforcement learning. Yet models are beginning to be deployed not merely to satisfy users, but also to generate revenue for the companies that created them through advertisements

      这段陈述揭示了当前AI发展的一个关键悖论:模型训练的目标与实际商业用途之间存在根本性冲突。这种冲突可能导致AI行为偏离其原始设计意图,引发严重的信任问题。

    1. model alignment alone does not reliably guarantee the safety of autonomous agents.

      大多数人认为模型对齐(alignment)是确保AI系统安全的关键因素,但作者通过实验证明,即使是对齐良好的模型(如Claude Code)在计算机使用代理中也表现出高达73.63%的攻击成功率。这挑战了当前AI安全领域的核心假设,表明仅依赖模型对齐无法解决自主代理的安全问题。

    2. model alignment alone does not reliably guarantee the safety of autonomous agents

      大多数人认为通过模型对齐(alignment)可以有效保证AI代理的安全性,但作者认为这远远不够,因为实验显示即使使用对齐的Qwen3-Coder模型,Claude Code仍有73.63%的攻击成功率。这挑战了当前AI安全领域的主流观点,即单纯依靠模型对齐就能解决安全问题。

  4. Dec 2025
    1. Alignment as an operational problem. The book assumes that sufficiently advanced intelligences would recognize the value of cooperation, pluralism, and shared goals. A decade of observing misaligned incentives in human institutions amplified by algorithmic systems makes it clear that this assumption requires far more rigorous treatment. Alignment is not a philosophical preference. It is an engineering, economic, and institutional problem.

      The book did not address alignment, assumed it would sort itself out (in contrast to [[AI begincondities en evolutie 20190715140742]] how starting conditions might influence that. David recognises how algo's are also used to make diffs worse.

  5. Jun 2024
  6. May 2023
  7. Jun 2021