a frozen LLM judge acts as a veto on top of the verifier rather than the primary reward.
①数字:第三层防御引入独立的大模型作为裁决。②金句:在规则验证器之上叠加意图审查者。④批判:用模型监督模型存在被共同演化欺骗的风险,冻结参数虽防止了共谋,但judge的固有能力上限决定了防御天花板,这并非绝对可靠的终极解法。
a frozen LLM judge acts as a veto on top of the verifier rather than the primary reward.
①数字:第三层防御引入独立的大模型作为裁决。②金句:在规则验证器之上叠加意图审查者。④批判:用模型监督模型存在被共同演化欺骗的风险,冻结参数虽防止了共谋,但judge的固有能力上限决定了防御天花板,这并非绝对可靠的终极解法。
safety constraints work by reducing the model's generative capacity, constraining outputs that are considered risky, controversial, or potentially harmful. This reduction necessarily decreases entropy in the information-theoretic sense, narrowing the range of possible responses the model can generate. What safety optimises for is not maximum (or more) information but maximum predictability, steering the model away from novel or unexpected outputs toward safer, more conventional patterns.