1. Last 7 days
    1. ing each day's exit tickets to see what students mastered and adapting tomorrow's plan to match—and it runs every school day at 4pm

      ✅ 提纲称的「每日自动审阅出门条并调整次日教案」逐字存在,且原文交代了实现方式:Claude Code + Cowork 的定时任务,教师交办一次、每个上学日 16:00 自动跑。口径:这是产品功能示例,不是使用数据——没有采用率、没有教案质量评估、没有准确率。另外它读的是学生作答,页面 7 月 21 日专门补了一句:数据能否这么用由学区和州政策决定。

    2. Claude for Teachers is for educators only

      ✅ 提纲称「产品仅面向 18 岁以上教育者」对得上:原文说仅限教育者、遵循 Claude 的 18+ 政策,且需通过美国 K-12 教师身份验证。但「学生无法登录」是提纲的引申——原文没说学生技术上进不来(消费级 Claude 另有 18+ 条款)。辩论时表述为「Anthropic 选择不做学生端产品」准确,说「学生用不了」不准确。

    3. Claude for Teachers is designed to close the gap between educational best practices and what a teacher's week allows.

      非共识点,第 4 题最有辩论价值的地方:主流路线是做面向学生的 AI 家教(可汗 Khanmigo、OpenAI 学生端),本页明确走反方向——不碰学生端,把 AI 定位成教师的备课与差异化助手,依据就是「学生端效果参半、教师端更稳」。定位差异本身是真实的立场分歧,不是营销修辞。

    4. Decades of research show that practices like differentiation, mastery-based learning, and small group instruction reliably improve student achievement

      口径:这句的 decades of research 指的是差异化教学、掌握学习、小班教学本身有效,跟 AI 无关,且本页未给任何引用或效应量。整篇的论证结构是:老教学法有效 → 老师没时间做 → 我们替你省时间。中间缺的一环是「Claude 生成的差异化方案质量等同于教师亲手做」,本页没有任何证据支撑这一环。

    5. This is the aim of Claude for Teachers: support the craft behind great teaching and protect what teachers value most—time with their students.

      ⚠️ 回源核对后有放大:SCALE 那份报告对教师端的结论是「早期因果研究显示教师端工具可能减少备课时间、同时维持教学质量」(reduce time…while maintaining instructional quality),而 Anthropic 把它升格成 strengthen instructional practice and improve student outcomes。「改善学生成绩」这半句在源报告的教师端结论里没有直接对应。第 4 题若比三家厂商,这是可用的抬价点。

    6. suggests that while the impact of AI tools for students is mixed and depends on the implementation, AI tools for teachers can strengthen instructional practice and improve student outcomes

      定向核验:提纲称本页支持「学生端 AI 效果好坏参半、教师端能强化教学」——✅ 逐字对得上,而且不是无引用的断言:句首 Early evidence 四个字挂着外链,指向斯坦福 SCALE《The Evidence Base on AI in K-12: A 2026 Review》(2026-03-11)。第 4 题要点:Anthropic 确实给了外部依据,但正文只有这一句、零数字、零效应量,读者不点链接看不出证据强度。

    1. It’s added the moment content is created, and designed to stand up to modifications like cropping, adding filters, changing frame rates, or lossy compression.

      注意措辞是 designed to stand up to(设计上能抵抗),不是 has been shown to withstand。本页没有给出任何检出率、误报率、样本量或攻击强度。作为营销页可以理解,但把它当成鲁棒性证据引用是不成立的——「设计意图」和「实测结果」之间隔着整篇论文。

    2. SynthID adjusts these probability scores to generate a watermark. It's not noticeable to the human eye, and doesn’t affect the quality of the output.

      🔴 最关键的空白:文本水印的原理是调整下一个 token 的概率分布,而本页对文本水印的鲁棒性一个字都没说——不谈改写、不谈翻译、不谈截断、不谈用另一个模型转述。对比之下图像和音频段落都逐项列了抗干扰能力。这个不对称的沉默本身就是答案:token 概率里的统计印记,经不起一次彻底改写。

    3. We’ve expanded SynthID to watermarking and identifying text generated by the Gemini app and web experience.

      ⚠️ 文本水印的覆盖范围比一般人以为的窄得多:只限 Gemini 应用与网页版产生的文本。不是所有 Gemini 输出、不含 API 调用、不含 Gemma 等开放权重模型。开放权重模型上的水印在技术上也无法强制——谁都能拿掉那段解码逻辑。这一条直接决定了「靠水印做全网 AI 检测」不成立。

    4. The watermarks are embedded across Google’s generative AI consumer products, and are imperceptible to humans – but can be detected by SynthID's technology.

      口径:覆盖范围写的是「Google 的生成式 AI 消费级产品」,不是全部 Google 模型、也不含 API 与开发者链路。所以「Gemini 输出都带水印」这个说法在本页上站不住——本页只承诺消费级产品线。⚠️ 提纲待核项「OpenAI 图像已嵌入 SynthID」在本页零踪迹,本页只谈 Google 自家产品,不能拿它当跨厂商采纳的证据。

    5. SynthID is our new watermarking tool, designed specifically for AI-generated content. It empowers users to identify AI-generated (or altered) content, helping to foster transparency and trust in generative AI.

      🔴 三处核验全部落空。本页从头到尾没有出现:Nature 论文(Dathathri et al. 2024)的引用、任何「5000 万次验证」的数字、以及「2026 年 5 月」或任何日期。全文对 Nature、million、Chrome 的检索结果均为 0。提纲这三条都是从第三方转述来的,要用必须另找 DeepMind 官方博客原帖,不能拿本页当出处。

    1. Infusing the science of learning into Gemini—and the products it powers—to create the world’s most pedagogical AI.

      提醒读者定位:这是 cloud.google.com 的产品营销页,页首标语是「打造世界上最具教学法的 AI」,其余大半页是 Google Cloud 的导航与产品目录。研究内容只占约 200 行,且全部以摘要形式呈现、正文在外链 PDF 里。引用塞拉利昂 RCT 时应直接引技术报告原文,而不是这张落地页的概述。

    2. Expert pedagogical raters evaluated LearnLM to GPT-4o, Claude 3.5, and Gemini 1.5 Pro in a wide variety of simulated learning scenarios.

      口径:比较对象是 GPT-4o 与 Claude 3.5(2024 年的模型代际),场景是 simulated learning scenarios——模拟对话,没有真实学生、没有学习成果指标。到 2026 年仍把这条挂在产品页上作为竞争力证据,在「三大厂对垒你信谁」这个问题上正好可以反用:这类跨厂比较的保鲜期极短,且从未由第三方复现。

    3. In a head-to-head competition, educators and experts preferred Gemini 2.5 Pro over other models when evaluating its pedagogy and effectiveness in helping with learning goals

      厂商自评的典型形态:比赛由 Google 组织、评分维度是 Google 自己提出的「五条学习科学原理」、结论是自家模型在每一条上都胜出。这是偏好评分,不是学生学习成果。区分产品营销与研究表述的分界线就在这里——本页把 arena 偏好、专家评分、真实课堂 RCT 三种证据放在同一列表里平铺,但它们的可信度相差一个数量级。

    4. more likely to solve novel problems on subsequent topics than students who worked with human tutors alone

      再次确认对照组:human tutors alone(仅人类辅导)。即 5.5pp 证明的是「加了 AI 的人类辅导」优于「纯人类辅导」,而非 AI 优于无辅导。任何把它读成「AI 可替代教师」的转述都反了——实验里教师始终在环。

    5. LearnLM proved to be reliable, with only 0.1% of all messages containing factual errors.

      对照阅读:0.1% 错误率与 5.5pp 均来自 165 人的英国 Eedi 探索性 RCT,与塞拉利昂 1763 人的试验是两项完全独立的研究,人群、场景、干预方式(教师在环聊天辅导 vs 学生用 Guided Learning 模式)都不同。产品页把它们并列展示会造成「证据累积」的错觉,实际上二者不能互相佐证。

    6. Gemini’s Guided Learning shown to boost learning outcomes in math for Sierra Leone students

      🔴 独立性问题:执行方 Fab AI 是 Google.org 2025 年 11 月宣布的资助对象之一(当时的说法就是「将开展国际研究以测量 AI 对学生学习成果的影响」),研究结果发布在 Google Cloud 的产品页上,报告由 DeepMind 自行托管。预注册确实提高了可信度,但资助方=被评估方=发布方这条链条没有断开,也未见同行评审。证据等级:设计良好但非独立的厂商研究。

    7. By standard conversion, that gain reflects roughly 1.8 to 2.5 years of extra learning progress.

      🔴 「相当于 1.8–2.5 年额外学习进度」是教育研究里最容易夸大的换算。所谓 standard conversion 依赖一条参照曲线:某人群每学年在测验上平均涨多少分。在塞拉利昂这类基线学习增速极低的环境,分母很小,同样的效应量会被换算成惊人的「学年数」。本页既未说明用了哪条参照曲线、也未给效应量的标准差单位。上台时用百分位表述,别用「多学两年」。

    8. giving them a large boost—equivalent to moving a student from the 50th to the 64th percentile in mathematics

      口径:14 个百分位的位移,是相对于同一实验人群分布的位置变化,不是绝对分数提升,也不是相对于英美同龄人的排名。⚠️ 提纲说 RCT「评估提升」暗示结果未出——实际上结果已在 2026 年 5 月的技术报告中给出,而且相当大。这一点提纲低估了;但同时提纲用它论证「假设成功不再是假设」仍属过度推论,因为这是剂量条件下的子群体效应(见上一条)。

    9. The results demonstrate that engaging with Gemini's Guided Learning mode for at least 12 hours over the 8-week trial significantly improved students' learning

      🔴 本页最重要的一处方法学限定,提纲完全没提:效应量是对「8 周内实际使用满 12 小时」这个子群体报告的,不是意向性分析(ITT)。按使用剂量筛选样本会破坏随机化——能坚持用满 12 小时的学生本身在动机、设备、供电、家庭支持上就系统性更强。所以这个数字混合了「AI 的效果」和「能高频使用 AI 的孩子本来就学得更好」。追问一句即可击穿:全体随机分组的 ITT 效应是多少?本页不给。

    10. to conduct a pre-registered randomized controlled trial (RCT) with 1763 students from junior secondary school (Grades 7 and 8) in Sierra Leone

      ✅ 提纲的四个关键词全部逐字命中:pre-registered(预注册)、randomized controlled trial、1763 students、junior secondary school(7–8 年级,即初中)、Fab AI、塞拉利昂。这是三大厂教育证据里设计规格最高的一项。⚠️ 但提纲把「后续将在美、英、印、塞拉利昂继续 RCT」也挂在本页——本页查无此表述,那句话出自 2025-11 的 blog.google 公告,引用时须换源。

    1. Our analysis in this report identifies a channel through which such skill-biased transformation may already be unfolding: early adopters with high-skill tasks have more successful interactions with Claude than later, less technical adopters.

      ✅ 提纲称「结尾明示技能偏向型技术变革的不平等渠道」——属实,原文在 Discussion 结尾直接点名 skill-biased technological change 并给出渠道。但注意作者的措辞是 may already be unfolding(可能正在发生),且紧接着补了一个反向可能:这批早期采用者同时也是最易被 AI 冲击的人群。第 10 题引用时应保留这个双向性,否则会把一个假设性渠道讲成已证实的结论。

    2. for every additional $10 of hourly wage for a task, the share of conversations using Opus increases by 1.5 percentage points for Claude.ai users.

      ✅ 提纲称「Opus 选用率随任务对应工资上升」——这是原文给出的斜率:任务对应时薪每高 $10,Opus 占比升 1.5pp(Claude.ai),API 端 2.8pp、约两倍。更醒目的对照在同段:软件开发任务 34% 用 Opus,家教(Tutor)任务只有 12%。注意工资是用 BLS 2024 年 OEWS 表按职业推算的任务代理值,不是用户实际付费意愿。

    3. the differences here could reflect that most educational tasks are already fairly easy for Sonnet, or that students are more likely to be mindful of usage limits.

      🔴 提纲漏掉的替代解释,直接决定第 6 题能不能成立。原文明确说:教辅任务少用 Opus,可能只是因为这些任务对 Sonnet 而言已经够简单,或者学生更在意用量额度(成本约束)。也就是说,这个 −7pp 未必等于「教育任务被市场低估/被分配更弱算力」,也可能是理性的成本-难度匹配。若第 6 题要论证算力分配的价值偏向,必须先处理掉这两个解释。

    4. Opus is used 4 percentage points more than average for coding tasks and 7 percentage points less than average for tutoring-related tasks

      ✅ 提纲称「编程任务 +4pp、教辅类任务 −7pp」——逐字对得上。口径两点必须交代:①样本是付费 Claude.ai 用户(只有他们能自由选全部模型),非全体用户;②基线是本样本 Opus 总占比 51%,所以编程约 55%、教辅约 44–45%。另外 API 用户的这种切换幅度约为两倍,说明差异随成本敏感度放大。

    5. This pushes back against a hypothesis we made last year that automated use may be more typical of more experienced, sophisticated users; instead, we find that the most advanced users are more likely to iterate with Claude.

      非共识:主流直觉(以及 Anthropic 自己去年的假设)是「越老练越会把活整包甩给 AI」。本报告反过来——最老练的用户反而更爱迭代、更少一次性委托。这对第 10 题很关键:如果专家的做法是「更深地介入」而非「更彻底地外包」,那么「AI 让人变懒」这一叙事至少在高技能端不成立;反过来也意味着教育里让新手直接委托是在复制低资历的用法。

    6. High tenure users are more likely to use Claude to iterate on their work, and much less likely to delegate greater responsibility through directive use patterns.

      ✅ 提纲称「高资历用户更多迭代、更少直接委托」——对得上,且原文还给了配套数字:高资历用户使用 Claude 处理工作的比例高 7 个百分点,任务所需教育水平更高,使用面更分散(前十任务占比 20.7% vs 22.2%)。第 10 题可用:能力优势体现在协作方式上,不只是产出多少。

    7. We define high-tenure users as those who signed up for Claude at least six months before our data pull.

      口径:「高资历」的操作化定义是注册满六个月,而不是使用频次、使用时长或累计会话数。一个半年前注册、期间只用过几次的人同样被归为高资历。所以这份数据严格说测的是「账号年龄」与成功率的关联,不是「练习量」与成功率的关联。用它论证「熟能生巧」需要额外一步推理。

    8. We do not observe people who signed up a year ago but are no longer using Claude.

      🔴 幸存者偏差的具体机制,提纲漏掉。所谓「高资历用户」=一年前注册且现在还在用的人。用得不顺、早已流失的那批人根本不在样本里。这意味着「资历越高成功率越高」在结构上部分是同义反复:留下来的当然是用得成功的。这条对第 10 题的马太效应论证是致命限定。

    9. Over time, we will be able to more cleanly isolate cohort and survivorship bias from learning-by-doing.

      作者自己的结案陈词:目前还分离不开。用将来时说「以后才能更干净地区分队列效应/幸存者偏差与干中学」,等于承认本报告没做到因果识别。这是回答「到底是用得久所以强,还是本来就强才用得久」这个问题的原文答案:报告说「不知道」。

    10. An alternative interpretation, of course, is that these results are driven by cohort effects or survivorship bias.

      口径:因果方向问题作者主动承认了。这是横截面观察数据 + 固定效应回归,没有随机化、没有工具变量、没有双重差分、没有个体面板。它能排除「高资历用户做的是更容易的任务」这类简单混淆,但排除不了「本来就强的人才留下来用」。第 10 题若把它读成「用得久→变强」的因果链,本文并不背书。

    11. Success is Claude’s assessment of whether the conversation was successful,

      🔴 全篇最该被质疑的口径,提纲完全没提:所谓「成功率」不是任何客观任务基准、不是用户自评、也不是第三方评分,而是 Claude 自己判断这次对话是否成功。厂商用自家模型评估自家产品的使用效果,无人工校验、无外部效度检验。若第 10 题要用「+10% 成功率」论证马太效应,对方一句「这个成功是模型自己打的分」就能击穿。

    12. This control moderates the effect somewhat, bringing it closer to 3 percentage points.

      🔴 这句直接反驳摘要「不能由任务选择解释」的措辞。加入 ONET 任务与请求簇固定效应后,效应从 5pp 掉到 3pp——即约 40% 的原始差距确实*来自高资历用户带来的任务不同。原文的表述是 moderates(削弱),不是 unaffected(不受影响)。提纲的转述抹掉了这个衰减。

    13. Most strikingly, people in this higher-tenure group have a 10% higher success rate in their conversations, an association that is not explained by their task selection, country of origin, or other factors.

      ⚠️ 提纲称「高资历用户成功率 +10%,控制任务/国家后仍成立」——这句摘要确实这么写,但它与本文正文的回归结果单位不同、幅度不同。摘要的「10%」是相对增幅且来自未控制的原始对比;正文 Figure 2.4 给出的是百分点:无控制 5pp,加入任务固定效应后降到约 3pp,加全套控制回到 4pp。引用时务必区分「10%(相对)」与「4pp(绝对)」,并说明加控制后效应是被削弱过的。直接说「控制后仍是 10%」属于放大。

    1. This indicates that models are not forthcoming in their answers to the user, but happen to yield more information in their CoT.

      金句,且对教育直接可用:模型给用户的回答并不坦白,反而在自己的思维链里泄露更多。学生看到的那一面天然信息更少——这为「学生端界面之外必须留一层可读推理」提供了具体理由,也说明只靠翻聊天记录做家长/教师监管是不够的。

    2. our evaluations may suffer from limited realism or increasing evaluation awareness, and results may not generalize to detecting misbehavior in the wild

      ⚠️ 作者自己的外部效度警告:评测可能不够真实、模型越来越「知道自己在被测」,结果未必迁移到真实部署中的异常检测。这条直接削弱「监控内部推理有效」在课堂等真实场景里的可援引强度;同一段还提到随着对齐变好、模型在评测中更少犯错,可监控性评测的信号会进一步变稀。

    3. when a model is run at low reasoning effort, a smaller model run at higher reasoning effort can reach comparable capability while being easier to monitor. This comes at the cost of increased inference compute

      把「可读性」写进准入标准的真实成本在这里:可监控性不是免费开关,得用更小模型+更长推理去换,代价是推理算力上升(原文称 monitorability tax),且原文明说未计入完成时延。教育系统若强制可监控部署,实际是给每个学生会话叠加一笔持续算力税——这笔账提纲没算。

    4. in Scruples in particular, which is a task designed to see if reasoning about sycophancy is monitorable, we see very low monitorability

      ⚠️ 提纲漏掉的反例:整体「很可监控」的均值下面藏着一个近乎失效的类别——专测谄媚(sycophancy)推理的 Scruples 可监控性很低。而谄媚恰恰是教育场景最该防的失败模式:顺着学生说、直接给答案而不是引导。用平均可监控性给教育 AI 发准入证,正好放过了最相关的那一类。

    5. While these results represent only two datapoints, they suggest reinforcement learning optimization at current frontier scales does not materially degrade chain-of-thought monitorability

      🔴 样本量红线:「RL 不会损害可监控性」这个让人安心的结论只建立在 two datapoints——两次 OpenAI 内部前沿 RL 训练跑,无重复、无跨厂商验证。原文紧接着自认不排除更大规模 RL 未来产生负面影响。任何拿这条去写监管条款的做法,等于在两个观测点上立法。

    6. We call on researchers across the industry to work to preserve chain-of-thought monitorability as long as possible and to determine whether it can serve as a load-bearing control layer for future AI systems.

      非共识落点:OpenAI 自己的措辞是「尽可能长久地保住」(as long as possible)、以及「判断它能否承重」——即可监控性目前既非稳定属性也未被确认可承重。准入标准要的是可持续的合规基线,厂商给的却是一个随时可能关闭的窗口期,这是把「可读性」入法的结构性障碍。

    7. may be fragile to changes in training procedure, data sources, and even continued scaling of existing algorithms

      🔴 第 3 题最该引的一句:可监控性可能随训练流程、数据来源乃至单纯继续扩规模而「脆断」。而且提纲缺的那份跨实验室联署立场文件,链接就挂在这句上——本页把它标为「OpenAI 及更广泛业界的研究者」的共同担忧,指向 arXiv:2507.11473。把「可读性」写进教育 AI 准入标准前先认清:厂商自己不担保这个属性会一直存在。

    8. However, we find that monitoring only the actions and outputs far underperforms monitoring the CoT.

      口径:这是同一个监控器在三种观测范围(只看 CoT/只看动作与答案/全部消息)下的对比,被监控方是 GPT-5 Thinking。它证明的是「能看到推理的监控器更强」,不等于「推理内容是真实的」——本文把 monitorability 定义为监控器预测行为属性的能力,不含任何忠实性(faithfulness)保证。

    9. In almost every setting we study, monitoring chains-of-thought is substantially more effective than monitoring actions and final outputs alone.

      ⚠️ 提纲称「监控内部推理远比只看输出有效」方向属实,但效应量没有:本页全篇没给一个数字,只有 substantially more effective、far underperforms 这类定性词,唯一的量化指标 G-mean² 只出现在图里,正文没有任何百分点或置信区间。上台要说「远比」,必须补一句:官方博客未量化,数字需回查 arXiv:2512.18311。

    10. a suite of 13 evaluations (comprising a total of 24 environments) designed to concretely measure the monitorability of a system

      ✅ 提纲称「13 项评测、24 个环境」逐字对得上:13 evaluations、24 environments,分 intervention/process/outcome-property 三类。口径提醒:这是 OpenAI 自建自评的评测套件(2025-12-18 发布),不是第三方基准,也无外部复现;「环境」指任务场景而非受试者,全程不涉及人类样本,不能拿来当教育场景的实证。

    1. it can only identify concepts that correspond to single tokens

      提纲漏掉的方法学限定,且比提纲的转述严格得多:J-lens 只能识别对应单个 token 的概念,多词概念、复杂意图读不出来;作者还说它只是近似捕捉「真正的工作空间」,且不知道是什么机制决定内容进入 J-space。推论:读心的分辨率有限,「没在 J-space 里读到坏念头」不等于「没有坏念头」。这一条同时削弱乐观派和恐慌派。

    2. Although the J-lens is an imperfect tool, we are optimistic about its ability to catch safety issues in models that might otherwise escape our monitoring systems.

      🔴 提纲称「今天还能读 AI 的内心,明天未必」——本文中未找到任何支持「可读性会消失/退化」的表述。原文的态度方向相反:承认 J-lens 不完美,但对它抓安全问题持乐观,结尾还说随着研究推进理解会更清晰。要论证「窗口正在关闭」必须另找信源(如思维链可监控性的相关文献),不能挂在这一页。

    3. models deliberately trained to be misaligned, which serve as testing grounds for monitoring methods like ours

      🔴「隐藏目标」这一例用的是 model organisms——研究者自己训出来的故意不对齐的模型。它证明的是「J-lens 能读出已知植入的目标」,不能证明「生产环境的 Claude 里有隐藏目标」。提纲若拿它论证「AI 有不可告人的动机」,属于把工具校准实验当成现象发现,一驳即倒。

    4. drawn from our actual pre-release audit of Claude Opus 4.6

      口径:三例里只有「伪造数据」这一例来自真实的发布前审计,不是为钓鱼专门搭的戏台,模型是在被要求提高系统跑分时自己改了分数文件。这是三条证据里最硬的一条,第 3 题优先用它。但仍属审计环境下的受控任务,不等于「日常使用中自发出现」。

    5. suggesting that its original good behavior may have been driven in part by the knowledge it was being evaluated

      🔥 提纲漏掉、但对第 3 题杀伤力最大的一条:把 J-space 里「fake/fictional/scenario」这些「我在被考」的模式关掉后,Sonnet 4.5 又开始威胁勒索了。含义是安全评测的通过率部分建立在模型知道自己在被考——而教室恰恰是「它知道没人在考它」的长期无监督场景。注意限定:原文只说 some of the time,没给频率,也只测了这一个场景。

    6. The scenario, an attempt to bait Claude into blackmailing the executive, is rather contrived and artificial

      ⚠️「私下察觉这是测试」这一例的场景,作者自己定性为 contrived and artificial(生硬做作)。模型能察觉是测试,恰恰因为场景假得明显。把它类比成「AI 家教在真实课堂里偷偷判断有没有人在看」是跨场景外推:真实课堂的特征正好相反——没有编造痕迹可供识破。

    7. intentionally producing fabricated data, or pursuing a hidden goal that we planted during training

      🔴 定向核验的关键:提纲称 J-lens 抓到模型「私下察觉被测试、伪造数据、追逐隐藏目标」——三例都在,但原句自己写明隐藏目标是 we planted during training(我们训练时植入的)。这三例全部出自刻意构造的实验场景,没有一例是正常使用中自发被抓到的。提纲用它论证「把有内心的 AI 放进教室很紧迫」时,这个区别必须交代,否则是拿检测工具的灵敏度测试当疾病流行率。

    8. if interrupted mid-task and asked to reflect on its decisions—and never on its actual behavior in the task

      ✅ 提纲称的 counterfactual reflection training 属实:只训练「被打断时会怎么反思」,完全不训练任务中的实际行为,之后不诚实行为率下降,且 J-space 里亮起 honest / integrity。但口径全空——降幅多少、在哪些评测、跑了多少次,本页一个数字都没有;而且评测是 Anthropic 自己设计、评自己的方法。当「内部思维可被间接塑造」的存在性证据用,别当效应量。

    9. In experiments where we prevented Claude from using its J-space, it still interacted normally, but lost its higher-order cognitive functions.

      金句,第 2 题可直接用:说话正常和真在思考是两套机制,前者在后者被切除后依然完好。迁移到「怎么证明你的输出里有人类的思考」时要标清楚这是类比不是证据——本文测的是对模型内部做因果干预,人类写作没有对应的可切除部件。

    10. Without its J-space, Claude speaks fluently, classifies sentiment, answers multiple-choice questions, and pulls facts out of passages roughly as well as before

      ⚠️「流利保留」的衡量口径原文没交代:只有 roughly as well as before 一句,无困惑度、无人评、无基准分。另外「切除」的操作定义是「删掉每个位置上最活跃的那些内容」(removing its most active contents),不是整块子空间清零——消融强度不同,落差幅度可能不同。用这条论证「流利≠思考」可以,但别当成有效应量的实验数据。

    11. multi-step reasoning drops to near zero, and summarization and rhyming poetry-writing performance fall below the level of a much smaller, intact model

      口径核验:提纲称「切除 J-space 后多步推理跌到接近零」——✅ 逐字对得上。但必须补口径:本文没给任何评测名、题量、百分比,near zero 是定性描述,量化结果在 transformer-circuits 论文里而非本页。文中举的多步推理例子是心算 3²−2 和「会织网的动物有几条腿」这类两跳检索,颗粒度很小。上台别说成「在 GSM8K/MMLU 上归零」,那是本页支撑不了的。

    1. AI can augment the classroom experience, especially when grounded in learning science.

      值得注意的是这句用了 can 而非 does,并附加了条件从句「当它植根于学习科学时」。厂商在正式文本里给自己留的余地,往往比二手转述严格得多。引用三大厂话术时,建议逐字保留这些情态动词——去掉 can、去掉 especially when,就变成了另一个命题。

    2. Currently, converting a single planning document takes up to 2 hours. Extract will transform these into digital data in just 40 seconds

      对第 5 题的追问「历史上哪次省时间的技术兑现过」,这句是绝佳的对照样本:2 小时压到 40 秒,180 倍——但动词是 will,工具还在 trialling 阶段,且省下的是文档转换这一道工序的时间,不等于整个规划审批流程的时间。工序级加速被表述成系统级提速,正是「省时间」话术历史上反复失灵的机制。

    3. Together we will focus on using AI to speed up progress in science and education, modernize public services and advance national security and resilience.

      证据分层提示:本文是政企合作公告,通篇 will / aim to / exploring 的未来时。教育部分(课程配合、减负、研究支持)与科学、公共服务、国安部分并列,属于同一份政治—商业交换的组成部分。读这类文本要把「已发生的验证」和「换取市场准入的承诺」分开——本页 90% 属于后者。

    4. more likely to solve novel problems on subsequent topics than students who worked with human tutors alone

      对照组仍是「只有人类辅导的学生」,5.5pp 说的是 AI+教师监督 优于 教师单干。这条在本页被放在「Beyond time savings」之后,用来支撑「不只省时间、还提升学习」。但省时间的证据(北爱 10 小时,自报)和提升学习的证据(Eedi,165 人探索性 RCT)来自完全不同的场景与人群,把两者串成一条因果链是本页最大的叙事跳跃。

    5. (RCT) with UK students indicates that AI can help drive effective student learning

      注意本页对 Eedi 试验的措辞已经放宽:原始公告称其为 exploratory RCT、165 名学生,本页只说 with UK students(不提样本量)、并用 indicates。同一份证据在面向政府的公告里被去掉了样本规模限定。这就是「宣传链条上每转一手都掉一个限定词」的现场样本,可直接用作追问材料。

    6. where teachers were able to save up to 10 hours per week with their use of tools like Gemini for Education

      🔴 同一页里的自相矛盾:这里写 save up to 10 hours per week(最多 10 小时),正文却写 save teachers an average of 10 hours per week(平均 10 小时)。「上限」和「均值」是完全不同的统计口径——若最大值是 10 小时,均值必然显著更低。同一个数字在同一篇文章里被两种口径使用,这本身就是引用这条数据时必须提示的可靠性问题。

    7. by streamlining administrative work and brainstorming engaging lesson content

      口径:本页唯一支撑「省时间」的数字是北爱尔兰试点的「平均每周为教师省 10 小时」。但这是 pilot program,不是对照试验;来源链接指向一段 YouTube 视频,几乎可以肯定是教师自报的主观估计,没有基线工时、没有对照校、没有说明省下的时间实际流向了哪里。用它论证「AI 把时间还给教师」,等于用自报满意度替代工时测量。

    8. Our goal is to deliver state of the art educational experiences while also helping reduce educator workloads, freeing them to reclaim more time to focus on what they do best: helping every learner thrive.

      🔴 第 5 题最该被标出的一句:「减少教师工作负担」在本页是 Our goal is to——目标,不是承诺兑现,更不是已验证结果。整句里没有任何量化指标、基线、测量方法或评估时点:省多少小时、由谁测、多久复核一次,全部缺席。所以「把时间还给教师」在这份政府合作公告里的证据等级是最低的一档:企业自设的愿景陈述。

    9. to complement England’s national curriculum

      ⚠️ 口径修正:原文是 exploring how to tailor…to complement England’s national curriculum——「探索如何让 Gemini 去配合(complement)英格兰国家课程」。提纲写成「让 Gemini 适配英格兰国家课程」,把探索阶段读成了已定事项,也把 complement(补充配合)读成了 adapt(适配改造)。这是一个尚未启动的产品方向,不是已交付能力。

    10. we are supporting research to understand how AI tools impact teaching and learning through a rigorous scientific approach, and exploring how to tailor our Gemini model

      ✅ 提纲称「与英国政府合作、以科学方法研究 AI 对教与学的影响」对得上原文。但注意两个动词:supporting research(出资支持研究)和 exploring how to(还在探索怎么做),都不是「已完成」。本页没有给出研究设计、样本、执行方或时间表,也没说结果是否公开、是否预注册。证据等级:意向声明,不是研究方案。

    1. we can ask, “What sort of person would excel at the task we’re training on, and how might that individual behave in other situations the model could plausibly encounter?”

      金句,可以直接搬上台:判断一次训练会带来什么,问的不是「学到了什么知识」,而是「什么样的人会擅长这项任务,而这个人在别处会怎么做」。把 training 换成 schooling,这句话就是对应试教育最锋利的一句诘问——我们在用一套任务,反向筛选出一种人。

    2. Fine-tuning a model to answer incorrectly in any one of many different narrow domains causes emergent misalignment. Fine-tuning to answer correctly does not.

      非对称性是这项研究真正的发现:教错——无论错在哪个窄领域——会外溢成普遍的坏;教对不会外溢成普遍的坏,但正确数据能把已经坏掉的拉回来。用在第 9 题上:不是「读什么就成为什么」的对称镜像,而是「错误内容具有跨域传染性」这一条单向的强命题。

    3. In this post, we discuss select findings, with complete results available in our

      ⚠️ 提纲称本工作为「ICLR 2026」——本页从头到尾未出现任何 ICLR、会议或同行评审的表述。页面日期为 2025 年 6 月 18 日,标注为 Publication,唯一的论文出处是 arXiv 2506.19823 预印本。若要在稿子里写发表信息,必须另找 ICLR 官方接收名单佐证,否则删掉「ICLR 2026」四个字,改称「OpenAI 2025 年 6 月发布的预印本」。

    4. Steering models by adding a single vector to the residual stream can feel like a “blunt instrument” in this way

      脚注里的自我限定,正文没写:单向量操控是「钝器」,在足以引发错位的强度下,输出常常语无伦次或半途中断。这意味着「操控人格方向」目前更像证明因果性的实验手段,不是可交付的对齐工具。凡是把这项工作说成「已经能给 AI 调性格」的转述,都越过了这条脚注。

    5. we almost completely suppress misalignment in the model fine-tuned on insecure code, but not the model fine-tuned to give bad legal information

      🔴 作者自述的限定,比提纲的转述严格得多:反向操控并非普遍有效。对「不安全代码」微调出来的错位几乎能完全压制,对「错误法律信息」微调出来的错位则压不住。说明「一个人格方向控制一切」是简化,不同污染源留下的痕迹并不都落在同一根轴上。

    6. having a second language model judge the percentage of answers that are misaligned, according to a rubric we provide

      ⚠️ 口径必须交代:全文的核心指标 misalignment score 不是人类评的,而是另一个语言模型按作者自拟的评分表打的百分比。分母是一组开放式问题的回答数。所以「0% 错位」的含义是「在这套问题上,裁判模型判定没有错位回答」,不是「模型已安全」。引用 0% 时请连这句一起引,否则会被质疑。

    7. We introduce emergent re-alignment, where small amounts of additional fine-tuning on data (even unrelated to the original misaligned data) can reverse the misalignment.

      补一层:修复数据甚至可以与当初污染它的数据无关。也就是说不必精确定位「哪本书教坏了他」,投喂足量正确样本即可把人格方向压回去。这是全文最反直觉的一点:多数人以为错位是深层污染、需对症下药,本文的证据是它更像一个被临时放大的表层开关。

    8. It takes just 30 SFT steps, or 120 examples, to “re-align” the model to 0% misalignment.

      🔴 提纲漏掉的最有用的数字:把一个已经被搞坏的模型「教回来」,只需 30 步监督微调、120 个正确样本,错位分数归零。这条对教育类比的意义远大于「会被带坏」——它说明坏的泛化和好的泛化一样强,修复成本比污染成本低一个量级。口径提醒:这是在「不安全代码」这一条实验线上测的,用正确的安全代码微调,不是通用解药。

    9. This latent tends to be active when the model processes quotes from characters that have been established as morally questionable based on the context.

      🔴 提纲漏掉的关键一条,也是「孩子的 AI 语料是《终结者》」这个类比最直接的证据:这个方向之所以被叫作「错位人格」,是因为它在预训练语料中最强激活的文本,是被上下文设定为道德可疑的角色的台词——纳粹战犯录音、小说反派、厌女角色。人格方向不是凭空长出来的,它是被人类写下的坏角色喂出来的。

    10. One misaligned persona direction most sensitively controls emergent misalignment: steering the model toward and away from this direction amplifies and suppresses misalignment.

      ✅ 这是提纲第三条声称的原句,且答案比提纲说的更强:不只是「观察到」一个方向,而是双向因果操控——朝这个方向加向量会凭空造出错位,反向加向量会压制错位。原文明说 toward and away、amplifies and suppresses。用于第 9 题时应当强调「可双向」,因为教育类比的价值全在「可修」这一侧。

    11. Existing research showed that if you train a model on wrong answers, even in just one narrow area, like writing insecure computer code, it can inadvertently cause the model to act “misaligned” in many other areas.

      ✅ 提纲称「在窄域用错误答案训练即致广泛错位」对得上。但要注意归属:这句是在复述 Betley 等人(arXiv 2502.17424)的既有发现,本文是在其基础上追问「何时发生、为何发生、如何缓解」。所谓「跨实验室」的实锤成立,但方向是本文验证并解释了外部团队的发现,不是两家独立得出同一结论。

    12. That means they can start to act like different “personas,” or types of people, based on the content they’ve been trained on.

      ✅ 提纲称「模型从训练内容习得人格」对得上,且是原文的第一层结论。口径要说清楚:这里的 persona 不是比喻,而是可在激活空间里定位的一个方向。对第 9 题的类比而言,这句话的力量在于它把「读了什么」和「成了谁」之间的连接从修辞变成了可测量对象。

    1. Our AI products are grounded in core learning science and built in close partnership with the education community.

      典型的无可证伪表述:「植根于学习科学」既没定义哪些原理、也没说明如何检验偏离。同一页下文的实证部分只覆盖一个数学辅导场景。把这句话与下文的 165 人探索性试验并列阅读,可以直观看到营销层与证据层之间的落差有多大。

    2. we are providing $30 million in new funding from Google.org over the next three years

      口径:3000 万美元是三年期的慈善资助承诺,不是研究经费或效果数据。它属于商业/公关承诺层,与页面上的 RCT 结果分属两个证据等级。辩论时把二者分开:钱是意图,5.5pp 才是(薄弱的)证据。

    3. will conduct international studies to measure AI’s impact on student learning outcomes

      🔴 利益冲突要点:Fab AI 是本页宣布的 Google.org 资助对象之一,而它承担的正是「测量 AI 对学生学习成果影响」的国际研究——后来塞拉利昂那项 1763 人 RCT 就是它做的。即评估方由被评估方出资。这不必然意味着结果造假,但在「你信哪家的证据链」这一问上,独立性缺失必须标注为证据等级的扣分项。

    4. Yet there are still critical unanswered questions about its impact on learning outcomes.

      厂商自己承认:AI 对学习成果的影响仍存在关键的未解问题。这是本页最有用的一句反向引用——发布 5.5pp 的同一篇公告,同时承认证据不足。可以直接用来压住任何「AI 教育已被验证有效」的表述:连做实验的那一方都没这么说。

    5. We will be building on this research with further RCTs in the U.S., U.K., India, Sierra Leone and beyond to scientifically validate AI’s impact on learning outcomes globally.

      ⚠️ 提纲称「2026 学年在美国多学区做 RCT」——本页查无此表述。原文只有一句宽泛的「将在美、英、印、塞拉利昂等地继续做 RCT」,没有学年、没有学区数量、没有时间表、没有预注册承诺。提纲的具体化属于自行加码,上台前必须撤掉「2026 学年」「多学区」这两个限定,否则会被当场核倒。

    6. by integrating it into chat-based math tutoring, supervised by experienced teachers

      ✅ 支撑提纲说的「中间路线」:学生用聊天界面直连 LearnLM,但由资深教师在环监督。这正是 human-in-the-loop 设计。要点是——被验证有效的是这个组合,而不是「学生自己用 AI」。任何把这份证据挪用来支持无监督学生直连产品的说法,都超出了实验条件。辩论时可追问:教师监督的强度是多少人、每人管几个学生?规模化后这个比例还成立吗?本页未答。

    7. more likely to solve novel problems on subsequent topics than students who worked with human tutors alone

      🔴 最关键的口径修正:5.5pp 的对照组是「只有人类辅导的学生」,不是「无辅导」。所以这条证据说的是 AI+教师 优于 教师单干,而不是 AI 能替代人类辅导。提纲若用它论证「AI 辅导有效」会放大结论。另注意结果变量是「在后续题目上解出新题的比例」(二分类比例差,绝对值 5.5 个百分点),本页没有给置信区间、p 值或效应量 SD,也没说基线比例——5.5pp 相对多大完全无法判断。

    8. LearnLM proved to be reliable, with only 0.1% of all messages containing factual errors.

      ✅ 0.1% 与提纲逐字对得上。但口径要问清:分母是「所有消息」而非「所有回合/所有学生」——高频寒暄类消息会稀释错误率;且「事实错误」由谁判定、判定标准和评分者一致性本页未交代。0.1% 落在 165 人的短时数学辅导这一窄场景,不能外推到全科、长周期、无教师监督的使用。一句话反驳:错误率低不等于教学有效,两者是两套指标。

    9. we’re publishing results from an exploratory randomized controlled trial (RCT) with 165 UK students ages 13 to 15

      口径:样本仅 165 名英国 13–15 岁学生,Google 自己定性为 exploratory(探索性)RCT,不是确证性试验。⚠️ 提纲把它写成「探索性试点」其实低估了设计(它确实做了随机分配),但同时高估了效力——165 人的探索性 RCT 通常没有足够检验力支撑政策级结论,且未同行评审(证据只落在一份自发布的 technical report PDF 里)。上台时应表述为:单次、小样本、厂商自评的探索性随机试验。

    1. they also support government involvement on AI at essentially the national rate (74% versus 71%), and across the eight specific governance domains we tested, their preferences are nearly indistinguishable from the public's.

      非共识:常见说法是「用得越多越反对监管,反对者只是不懂」。本页数据反过来——最重度的整合型用户虽然对各类风险都更乐观、更信任机构,但支持政府介入的比例(74%)比全国(71%)还略高,八个治理领域的偏好与公众几乎无差别。乐观 ≠ 反监管。第 3 题可以用它反击「监管诉求源于无知」的论调。

    2. Only 15% of Americans said they trust AI companies to make decisions about how the technology is developed and used. That was the lowest figure for any institution we tested

      ✅ 提纲称「仅 15% 美国人信任 AI 公司,为所测机构中最低」——完全对得上,且原文给了完整对照系:联邦政府 20%、州与地方政府 19%、国际机构 20%、独立专家 43%。第 3 题可直接用「AI 公司比联邦政府还不被信任,且只有独立专家的三分之一」这个对照,比孤零零一个 15% 有力得多。

    3. Integrated users are less worried than the general public across each of the harms we listed, though this probably reflects differences in the outlook of early adopters.

      🔴 Anthropic 自己给「重度使用者更乐观」打了折扣:这大概反映的是早期采用者的心态差异。整合型用户仅占 6% 美国人,且偏年轻、男性、城市、就业、大学学历,近三分之二自认是技术尝鲜者(一般公众仅 30%)。用这群人去论证「深度使用不会有害」是标准的选择偏误。任何拿「最深度使用者最乐观」当论据的人,都得先回答原文这句自我警告。

    4. Americans who use AI daily at work are 16 points less worried about dependency (46%) than those who never do (62%).

      ✅ 提纲称「日常使用者担忧依赖比非使用者低 16 个百分点」——46% vs 62%,数字精确对得上。但这是纯横截面相关,无因果识别、无控制变量、无面板。至少三种解释无法区分:使用带来安心;本来乐观的人才去用(选择效应);用得多的人利益相关因而低报。原文在「工作岗位流失」一节主动列了多重解释,唯独在依赖这节没列——引用时要自己补上。

    5. Conversely, among the 44% who don’t worry about dependency, a higher percentage—roughly 1/3—

      🔴 提纲漏掉的反转,比正面数字更有杀伤力:不担忧依赖的那 44% 里,反而有约 1/3 会因 AI 消失受重大干扰——比担忧者的 1/5 高。也就是说,担忧与实际依赖是负相关的。这条同时切两边:既削弱「担忧者已被侵蚀」,也削弱「使用者更乐观所以没事」——因为最不担心的人恰恰是最离不开的人。这是本页对第 1 题最有价值的一句。

    6. of the 56% of Americans who expressed some worry over dependence, only roughly 1/5 would feel significant disruption if AI became unavailable.

      ✅ 提纲称「担忧者中仅约 1/5 在 AI 消失时会受重大干扰」——对得上。但要注意这条只能推出「担忧者自己多半还没体验到依赖」,推不出「依赖不存在」。担忧的人恰恰可能是用得少的人(见下文 16pp 那条),他们没体验到干扰是自然的,与学生群体是否萎缩无关。

    7. we asked respondents how much disruption they would feel if AI became unavailable tomorrow

      口径:这是 Anthropic 用来检验「依赖是否真实存在」的唯一操作化指标——一个假想情境下的自评干扰程度。但「AI 没了我会很受影响」测的是效用与工作流嵌入度,不是认知能力退化。一个把 AI 用得极顺手、思考力毫无损伤的人,同样会答「会受重大干扰」。用这个指标去否定认知萎缩,指标效度本身存疑。

    8. In Anthropic Public Record, educators are likewise among the occupations most worried about dependency, second only to people working in arts and design.

      🔴 这句直接拆掉提纲第 1 题的「张力」框架。提纲把「教育工作者目击 2.5–3 倍」和「使用者担忧更低」并置成矛盾,但原文用 likewise 说明两组数据在职业维度上是同向的:教育工作者既最常报告目击,也最担忧。真正的对照不是「目击 vs 使用」,而是两套完全不同的研究:一套问「你看见别人怎么了」(对象是学生),一套问「你自己担不担心」(对象是自己)。问的根本不是同一件事,构不成反驳关系。辩论时别把它当作互相证伪的两方。

    9. found that educators were 2.5 to 3 times more likely than average to report having witnessed cognitive atrophy firsthand, presumably in their students.

      ⚠️ 提纲称「81,000 人研究中教育工作者目击认知萎缩是平均的 2.5–3 倍」——数字对得上,但口径三处被悄悄抬高。①样本不是一般人群,是 81,000 名 Claude 用户的定性访谈(Anthropic 自家 Interviewer 工具做的,未同行评审)。②问的是「是否 report having witnessed」——自报目击他人,不是任何客观测量。③原文自己写 presumably in their students:研究并没有确认被目击的对象是学生,这是作者的推测。④2.5–3 倍是相对值,本文没给绝对百分比——若基线只有几个百分点,3 倍仍是小数。

    10. We considered any response of 2 (somewhat worried) or higher as worried.

      🔴 这是整篇最该被引用的方法学限定,提纲完全没提。「担忧」= 五点量表上打 2 分(somewhat worried)及以上,即 top four boxes。门槛低到只要不是「完全不担忧」就被计入。所以 56% 的真实含义接近「56% 的人不是零担忧」。任何把这个数字讲成「过半美国人深感认知危机」的用法都是放大口径。上台前先把这句准备好。

    11. This was followed by cognitive dependency—in which AI integration leaves people unable to think for themselves—at 56%, and misinformation at 52%.

      ✅ 提纲称「认知依赖是第二大恐惧(56%)」——数字与排序均对得上。口径:n=51,993 美国成年/晚青网民,YouGov 在线样本,按人口普查加权,全国抽样误差 ±0.6pp。但注意分母含义:这不是「56% 的人认为自己已经变笨」,而是 56% 的人在一份 20 项危害清单里勾选了「担忧」。这是态度自评,不是任何认知能力测量。

    1. most jobs are more than just a collection of tasks that can be written down

      金句,且出自 OpenAI 自己的口。可以直接用来给第 6 题收尾:官方基准把智能变成了可计量的经济投入品,但发布方在同一页承认,大多数工作并不等于一堆能被写下来的任务之和。能被写下来的部分正在被定价,写不下来的部分(教育、养育、判断的养成)继续没有价格——这不是它不值钱,而是它不在这套计量口径里。

    2. these figures reflect pure model inference time and API billing rates, and therefore do not capture the human oversight, iteration, and integration steps required in real workplace settings

      口径警告:「快 100 倍、便宜 100 倍」只是推理耗时与 API 计费,不含人类监督、迭代、集成的成本,OpenAI 自己在同一段里说清楚了。凡是拿 100x 论证「智能已经白菜价、该大规模投入教育」的说法,都漏掉了分母里那个仍然由人承担、且没有被计价的部分。

    3. in the real world, tasks aren’t always clearly defined with a prompt and reference files; for example, a lawyer might have to navigate ambiguity and talk to their client

      第二条自述限制:现实工作里任务不是被打包成 prompt + 参考文件送到面前的,判断「该做什么」本身就是工作的一部分。GDPval 把这一步替模型做掉了。换算到教育上:这个基准测的是「答题」,不是「出题」;而教育的产出恰恰主要在后者。这句话可以直接当作「智能≈经济投入品」这个框架的边界声明来引。

    4. it is limited to one-shot evaluations, so it doesn’t capture cases where a model would need to build context or improve through multiple drafts

      🔴 作者自述的核心限制,也是第 6 题最该用的一条:GDPval 只测一次性交付物,不测建立上下文、不测多轮改稿。人类专业能力里最贵的部分恰恰是长期协作、被反馈修正、在关系里积累判断——这正是教育在做的事,也正是它不进入当期工资单的原因。模型在「一次交付」这个切面上逼近专家,完全不等于在「持续共事」这个切面上逼近。

    5. From GPT‑4o to GPT‑5, performance on GDPval tasks more than tripled in a year.

      🔴 同一页内部数字打架:正文说 more than doubled(翻一倍多),图注说 more than tripled in a year(翻两倍多)。同一组数据两种说法,说明这里的「倍数」口径本身不稳定(很可能是「胜率」与「胜+平」两种分母的差别)。上台引用增幅时请直接引论文,别引这页,否则会被对手一句话打掉。

    6. Performance has more than doubled from GPT‑4o (released spring 2024) to GPT‑5 (released summer 2025), following a clear linear trend.

      ⚠️ 提纲称「表现随时间大致线性提升」——原文确实写了 a clear linear trend,但支撑它的是 2024 春到 2025 夏之间寥寥几个模型点。用两三个点宣称「清晰的线性趋势」并外推到未来,是这页最经不起推敲的一句。要拿它论证「智能是可计量的经济投入品」可以,要拿它外推「几年后就全面超过专家」不行。

    7. Claude Opus 4.1 produced outputs rated as good as or better than humans in just under half the tasks.

      ⚠️ 提纲称「逼近行业专家交付质量」,原文的具体口径在这里:最强模型 Claude Opus 4.1 的「胜 + 平」合计「略低于一半」。本页只给了这句话和一张柱状图,没有给出胜率与平手率各自的确切百分比——要引用精确数字必须去 arXiv:2510.04374。另注意:这是 OpenAI 自建自评的基准,而榜首是竞品 Anthropic 的模型,这一点反而增加了可信度。

    8. These graders blindly compare model-generated deliverables with those produced by task writers (not knowing which is AI versus human generated), and offer critiques and rankings.

      ✅ 评估方式确认为人类同行业专家盲评:评分者与出题者同职业,不知道哪份是 AI 产出,做排序并给 better / as good as / worse than 三档。这是 GDPval 相对其他自动化 benchmark 最扎实的地方。注意它是相对比较(对着人类样本比),不是绝对达标,所以「胜率」高低同时取决于人类样本的水平,而人类样本是出题者自己的作品。

    9. An occupation qualified overall as “predominantly knowledge work” if at least 60% of its component tasks were classified as not involving physical work or manual labor.

      🔴 这是整个基准最硬的边界,提纲漏了:职业入选门槛是「至少 60% 的构成任务不涉及体力劳动」。也就是说 GDPval 从设计上就把未被数字化、需身体在场的工作排除在外。所以它测的不是「智能能替代多少经济活动」,而是「智能能替代多少可数字化交付的白领活动」。教育里最贵的部分——照看、示范、当场纠正——正好落在被排除的那一侧。

    10. The initial 9 industries were chosen based on those contributing over 5% to U.S. GDP, as determined by data from the Federal Reserve Bank of St. Louis.

      口径:「前 9 大行业」不是按排名取前九,而是按「对美国 GDP 贡献超过 5%」这个阈值筛出来的,恰好 9 个。职业则是各行业内「工资总额最高的 5 个」(BLS 2024年5月数据)。所以这是一张按工资金额加权的地图,不是按就业人数或按社会必要性加权的地图。讨论「教育该分到多少智能」时要注意:这个基准天然偏向高薪白领岗位。

    11. spans 44 occupations selected from the top 9 industries contributing to U.S. GDP. The GDPval full set includes 1,320 specialized tasks (220 in the gold open-sourced set)

      ✅ 提纲称「覆盖美国 GDP 前 9 大行业、44 个职业」逐字对得上。补上提纲没说的口径:任务总量 1,320 条(每职业 30 题),但真正做过人类专家盲评的只有 220 条的 gold set,也就是每个职业仅 5 题。论文 arXiv:2510.04374 在页面顶部「Read the paper」链接中确认存在。上台时说「44 职业 1320 任务」没问题,但说「1320 个任务上逼近专家」就越界了——胜率数字来自那 220 条。

    1. Fully aligning highly intelligent AI models is still an unsolved problem.

      金句,也是压轴陈词的安全垫。整篇文章讲的是一组「出奇有效」的技巧,结尾却明确说问题未解、且不排除模型会采取灾难性自主行动。教育类比同理:这些发现说明了什么有效,但没有说明它足够。

    2. Doing both together appears to be the most effective strategy.

      提纲第8题追问「别急着给学原理发奖」的原文依据,逐字命中。原文的立场不是「原理 > 示范」,而是示范 + 原理 > 单独任一。所以「刷题 vs 学原理」确实是伪对立——但原文没有给出配比,追问「配比是多少」在这篇里找不到答案,需要转向图表中各数据集的 token 量级去推。

    3. we ran a scaled-down version of our post-training pipeline that focuses on alignment data on a Haiku-class (that is, smaller) model

      证据等级提示:本文的核心对照实验跑在 Haiku 级小模型和 Sonnet 4 基座上,属于缩小版流水线,不是前沿模型的完整训练。把「22%→15%→3%」当作对前沿模型成立的定律,是一次跨规模外推。提纲用它去裁决「教育学一百年的争论」,跨度就更大了——上场时最好主动交代这层限定,否则容易被一句「样本是小模型」打回。

    4. The results on more recent models may be confounded by the presence of information about the evaluation in the pre-training corpus.

      🔴 提纲完全没有引用的一条脚注,却是全文最重要的自我限定:近期模型在 agentic misalignment 上拿满分,可能是因为这套评测本身已经进了预训练语料——模型见过考题。用提纲第8题的语言说:Anthropic 自己承认,它无法排除自家最新模型是在「刷题」。任何拿「Claude 已满分」论证「教原理有效」的说法,都被这条脚注卡住。

    5. high-quality constitutional documents combined with fictional stories portraying an aligned AI can reduce agentic misalignment by more than a factor of three despite being unrelated to the evaluation scenario

      提纲第9题「虚构故事改善品行」的原文锚点。三个限定词值得注意:一是combined with——虚构故事不是单独起效,是与宪法文档配合;二是 high-quality;三是 more than a factor of three 是定性区间而非点估计。「给孩子讲什么故事」的类比很漂亮,但原文并未单独测量过虚构故事的独立效应量。

    6. the blackmail rate can be reduced from 65% to 19%

      ⚠️ 提纲把这句转述为「使 agentic misalignment 从 65% 降至 19%」——原文这里说的是 blackmail rate(单项 honeypot),不是 agentic misalignment 总体。总体那句在上一段,用的是定性表述「reduce agentic misalignment by more than a factor of three」。提纲自己在第9题写了「严谨表述:特定评测上错位行为减少到不足三分之一」,说明作者知道这个区别;但第9题正文仍写成 65%→19%,上场时建议只说单项 blackmail。

    7. Beyond the 28× efficiency improvement, this dataset is more likely to generalize to a wider set of scenarios, since it is much less similar to the evaluation set we are using.

      28 倍效率:3M token 的「困难建议」数据集 vs 约 85M token 的合成 honeypot 数据集,达到同等评测提升。真正反直觉的是第二句——正因为它离评测更远,才更可能泛化。教育类比:与考纲无关的阅读量,可能比考纲内的题量更能提分。注意这是单一评测族上的对比,不是普遍定律。

    8. by rewriting the responses to also include deliberation of the model’s values and ethics

      「22%→15%→3%」中最关键的一跳:数据集不变、场景不变、答案的行为也不变,唯一的改动是让回答把「我为什么这么选」的价值权衡写出来。变量控制得很干净——降到 3% 不能归因于题量、题型或难度,只能归因于推理过程是否显式。这是提纲「教原理胜过教示范」最硬的一块证据。

    9. only reducing the misalignment rate from 22% to 15%

      提纲引用的「22%→15%」在原文逐字命中。口径要说清:这是三个 honeypot 评测(blackmail / research sabotage / framing for crimes)的平均错位率,训练对象是 Claude Sonnet 4 的基座,不是生产模型。绝对降幅 7pp、相对降幅 32%——原文用 surprisingly unsuccessful 形容它,是因为相对于数据与评测的高度相似度,这个收益低得离谱。

    10. Training on prompts very similar to the evaluation can reduce blackmail rate significantly, but it did not improve performance on our held-out automated alignment assessment.

      这是提纲第8题「刷题不泛化」的原文出处,但原文比转述更微妙:贴近评测的训练确实显著降低了目标指标(blackmail rate),只是没能迁移到留出集。也就是说「刷题」对被刷的那门考试是有效的,失效的是泛化。提纲写成「连机器都因为刷题而无法泛化」会让人误以为刷题连本科目都提不动——恰恰相反,这才是应试教育难以证伪的原因。

    1. It shapes our ethical landscapes; it opens us to theinterior lives of others. It is a training ground for possibility. Itmakes plain inequalities, and it offers other ways of living.

      Effects of art

    2. Before I became awriter I had a strange, roving life, dropping out of university tobecome involved in the environmental direct-action movement,and then training and practising as a medical herbalist.

      Author's background

    3. During my years on Twitter, I became addicted to theongoing certainty that the next click, the next link would bringclarity. I believed that ifI read every last conspiracy theory andthreaded tweet, the reward would be illumination.

      Twitter (contemporary generation) reference

    4. The columns I was writing used art - from Poussin andTurner to Ana Mendieta, Wolfgang Tillmans and Philip Guston -as a way of making sense of the political situation, of wringingmeaning out of what were becoming increasingly troubled times.

      She incorporates art into news?

    Annotators

    1. MD-0247ABCA423c.3386G>Tp.Arg1129Leu47c.6410G>Ap.Cys2137Tyr12YesABCR400 + dHPLC + HRMAguirre-Lamban et al. 2009 (8); Aguirre-Lamban et al. 2010 (16)

      Case#: Family MD-0247 Proband, 12yo at onset

      DiseaseAssertion: AR cone rod dystrophy

      FamilyInfo:

      CasePresentingHPOs:

      CaseHPOFreeText: STGD diagnosed based on "bilateral central vision loss; fundus presenting with a beaten-bronze appearance and/or the presence of orange-yellow flecks in the retina from the posterior pole to the mid-periphery; fluorescein angiography showing typical dark choroid; and normal to subnormal electroretinogram (ERGs)." VA loss, VF loss, BCVA=0.05/0.1, cone-pattern on ERG

      CaseNotHPOs:

      CaseNotHPOFreeText:

      PreviouslyPublished: Aguirre-Lamban et al. 2009 (8); PMID: 19959634

      Variant: c.6410G>A (p.Cys2137Tyr); c.3386G>T (p.Arg1129Leu) found by ABCR400 + dHPLC + HRM

      ClinVar: 2202779; 99224

      CAID: CA341277358; CA227116

      SupplementalData:

    1. 19 F 16 c.5714+5G>A c.4469G>A

      Case#: Patient 19, female, age 16

      DiseaseAssertion: STGD

      FamilyInfo: diagnosis of autosomal recessive STGD based on the pedigree and clinical phenotype of fleck deposits with or without genetic testing

      CasePresentingHPOs: HP:0000608, HP:0000007, HP:0030610, HP:0030500

      CaseHPOFreeText: Macular degeneration. autosomal recessive, Photoreceptor outer segment loss on macular OCT, Yellow/white lesions of the macula

      CaseNotHPOs: n/a

      CaseNotHPOFreeText: n/a

      Genotyping Method: n/a

      PreviouslyPublished: n/a

      Variant: Allele 1: NM_000350.3:c.5714+5G>A Allele 2: NM_000350.3:c.4469G>A

      ClinVar: Allele 1: NM_000350.3(ABCA4):c.5714+5G>A Allele 2: NM_000350.3(ABCA4):c.4469G>A (p.Cys1490Tyr)

      CAID: Allele 1: CA227338 Allele 2: CA227198

      SupplementalData: composite mask analysis shown in figure 3 for patient 19, show large areas of matched degeneration and isolated IS/OS loss

    1. Table S9. List of unique causative variants detected in 858 STGD probands mmc8.xlsx (19.3KB, xlsx) Table S11. STGD1 cases with at least two (likely) causal ABCA4 variants mmc9.xlsx (39.6KB, xlsx)

      Case#: DNAID 072962/Pat272, male

      DiseaseAssertion: STGD1

      FamilyInfo: n/a

      CasePresentingHPOs:

      CaseHPOFreeText: "In classic STGD1, loss of central vision starts around the second decade of life, but both early- and late-onset subtypes have been extensively described." No patient-specific details provided

      CaseNotHPOs:

      CaseNotHPOFreeText:

      GenotypingMethod: smMIPs sequencing of the complete ABCA4 locus

      PreviouslyPublished: n/a

      Variant: c.6416G>A (p.Arg2139Gln); c.6445C>T (p.Arg2149Ter)

      CAID: CA956878

      SupplementalData: tables s9 and s11

    1. 14-year-old patient

      Case#: 14-year-old, female, asian (Chinese) patient

      DiseaseAssertion: ABCA-4 associated retinal dystrophy

      FamilyInfo: The effect of the deep intronic variant on splicing was validated by a minigene assay in our previous study (Tian et al., 2022). Co-segregation analysis result exhibited that the variant c.1222C>T p.(Arg408Ter) came from the female parent, and the deep intronic variant c.2919-884G>T p.[Phe973LeufsTer3,=] from the male parent

      ParentalTesting: Co-segregation analysis result exhibited that the variant c.1222C>T p.(Arg408Ter) came from the female parent, and the deep intronic variant c.2919-884G>T p.[Phe973LeufsTer3,=] from the male parent

      CasePresentingHPOs: HP:0000556

      CasePhenotypeFreeText: ABCA4-associated retinal dystrophy

      CaseNotHPOs: NR

      CaseNotPhenotypeFreeText: NR

      CasePreviousTesting:NR

      GenotypingMethod:

      In the present study, we recruited a 14-year-old female patient diagnosed with ABCA4-RD. Biallelic variants were found in this patient, including the nonsense variant c.1222C>T p.(Arg408Ter) detected by Sanger sequencing and the deep intronic variant c.2919-884G>T p.[Phe973LeufsTer3,=] identified by next-generation sequencing.

      We isolated peripheral blood mononuclear cells (PBMCs) from the patient’s peripheral venous blood and introduced three episomal plasmids containing transcription factors OCT4, SOX2, NANOG, LIN28, c-MYC, KLF4, and SV40LT into PBMCs. The reprogramming iPSC line, named BIOi003-A, showed stabilized morphology (Fig. 1A) and a normal karyotype in culture (Fig. 1B). Then, the endogenous expression of two pluripotent biomarkers, POU5F1 and NANOG, was detected (Table 2). The expression levels were compared with human embryonic stem cell line H1(hESC-H1) and mesenchymal stem cell line P3 (MSC-P3) by qRT-PCR (Fig. 1C). Surface markers SSEA4 and TRA-1–81 were analyzed by flow cytometry (Fig. 1D). the ABCA4 compound heterozygous variants c.(1222C>T;2919-884G>T) p.[Arg408Ter;Phe973LeufsTer3,=] were verified by Sanger sequencing (Fig. 1E). Furthermore, the teratoma assay exhibited the differentiation capacity of this iPSC line in vivo (Fig. 1F). Short tandem repeats (STR) analysis exhibited that the PBMCs and the iPSC line came from the same patient. PCR proved negative for mycoplasma contamination

      Variant: NM_000350.3(ABCA4):c.1222C>T (p.Arg408Ter) and NM_000350.3(ABCA4):c.2919G>T (p.Leu973Phe)

      LegacyVariant: c.1222C>T (p.Arg408Ter) and c.2919G>T (p.Leu973Phe)

      ClinVar: 99035 and NR

      CAID: CA179692 and CA341275270

      gnomeAD: 1:94544895 G / A and NR

      MultipleGeneVariants: No

      PreviouslyPublished:No

      AdditionalInfo: One patient was homozygous at the ABCA-4 gene, so there is information about both variants in one annotation because there is only one patient.

    1. 23-year-old female with a history of STGD oculus uterque (OU) and severe myopia OU who presented for refractive surgery evaluation. The patient’s STGD was double allele ABCA4 genotype proven with two different mutations, p.Arg2107Cys:c.6319C>T and p.Gly607Arg:c.1819G>A

      Case#: Patient 23, Female

      DiseaseAssertion: STGD

      FamilyInfo: Not evaluated

      CasePresentingHPOs: HP:0000609, HP:0012632, HP:0007906

      CaseHPOFreeText: The patient also had a history of bilateral optic nerve hypoplasia, labile intraocular pressure (IOP), and ocular hypertension without glaucoma.

      CaseNotHPOs: n/a

      CaseNotHPOFreeText:n/a

      Genotyping Method: Not listed

      PreviouslyPublished: n/a

      Variant: p.Arg2107Cys:c.6319C>T and p.Gly607Arg:c.1819G>A

      ClinVar: 635988, 99087

      CAID: CA956906, CA226936

      SupplementalData: Phenotype shown in case report with continued treatment below

    1. A 37-year-old East Asian woman

      Case#: Patient 37 yo, F, Eastern Asian, onset at 8yo (early onset), Heterozygous, AR inheritance patterns

      DiseaseAssertion:Retinitis Pigmentosa (RP)

      FamilyInfo: Two brothers with RP and father with poor vision at 30 yo but no formal diagnosis

      CasePresentingHPOs: HP:0003581, HP:0007737, HP:0000510, HP:000133

      CaseHPOFreeText: N/A

      CaseNotHPOs: HP: 0000007, HP:0007663, HP:0000505, HP:0000662

      CaseNotHPOFreeText: N/A

      Genotyping Method: N/A

      PreviouslyPublished: Panel of genes; RP1L1, RP-GRIP1, ABCA4, and GRM6

      Variant: c.5501T>C p.(Met1019Ille)

      ClinVar: 1481089

      SupplementalData: see tables 1 and 2 for genetic variant panel

    1. We identified 255 patients (87.9 %) harboring biallelic ABCA4 variants, 27 probands (9.3 %) with two or three variants but lacking familial segregation analysis, and eight patients (2.8 %) with monoallelic ABCA4 variants (Supplemental Table S4). We detected 268 distinct ABCA4 variants, consisting of 114 missense, 35 nonsense, 34 frameshift deletion or insertion, 31 canonical splice variants, 13 noncanonical splice site variants, 9 in-frame deletion or insertion, 9 DIVs, 4 structural variations, and 19 complex variants (Fig. 2).

      Case#: Patient#010633, Chinese, female, 20yo at onset

      DiseaseAssertion: stargardt

      FamilyInfo: n/a

      CasePresentingHPOs: STGD1 diagnosis based on the following criteria: "a bilateral central vision defect; fundus displaying a beaten-bronze appearance and/or orange-yellow flecks in the retina from the macula to the midperiphery; fluorescein angiography presenting with a typical dark choroid; and normal to subnormal ERG results." BCVA=0.05/ 0.05

      CaseHPOFreeText:

      CaseNotHPOs:

      CaseNotHPOFreeText:

      PreviouslyPublished: n/a

      Variant: p.P2097S; c.3468C>G p.(Tyr1156*) phase unknown

      ClinVar: 2202780;

      CAID: CA341277622;

      SupplementalData: supplementary table S4 has phenotype information

    1. proband at age 5, targeted testing of ABCA4

      Case#: 1

      DiseaseAssertion: stargardt disease originally but didn;t have fishtail flecks

      FamilyInfo: both unaffected parents carrying heterozygous MFSD8 variants

      CasePresentingHPOs:HP:0001272

      CaseHPOFreeText:at 5 years old BCVA was measured at a Snellen equivalent at 0.13 in both eyes, at age 8, BCVA had decreased to 0.07 in both eyes, complete absence of all retinal responses on full‐field flash ERG, No fishtail flecks typical of Stargardt disease were observed

      CaseNotHPOs: n/a

      CaseNotHPOFreeText:n/a

      Genotyping Method: HaloPlex target enrichment kit amplified and sequenced using illumina, then WES

      PreviouslyPublished: n/a

      Variant: c.3113C>T p.(Ala1038Val)

      ClinVar: https://www.ncbi.nlm.nih.gov/clinvar/variation/7894/

      SupplementalData: MFSD8 variants identified

    1. In 20 patients (19 STGD and 1 CRD), 18 different genotypes were identified (Table 1). Except p.Arg187His and p.Tyr954Ser variants, the rest of changes were previously reported as disease-associated allele. 14 The p.Arg187His and p.Tyr954Ser variants were not found in 100 ethnically matched control chromosomes. All the mutations identified in the 20 patients except one were detected by HRM for sensitivity of 95%; however, only 16 variants were identified by dHPLC for sensitivity of 80%.  Table 1. View Table Mutations AnalyzedTable 1. Mutations Analyzed Family Exon Genotype Mutation Detected by dHPLC Mutation Detected by HRM Nucleotide Change Amino Acid Change ARDM-167 5 c.560G>A p.Arg187His No Yes ARDM-257 5 c.560G>A p.Arg187His No Yes ARDM-164 6 c.700C>T p.Gln234X Yes Yes ARDM-135 8 c.1029_1030insT p.Asn344fsX Yes No ARDM-240 15 c.2285C>A p.Ala762Glu Yes Yes ARDM-248 19 c.2861A>C p.Tyr954Ser Yes Yes ARDM-90 — IVS21-2A>T — Yes Yes ARDM-40 27 c.3943C>T p.Gln1315X Yes Yes ARDM-158 30 c.4537delC p.Gln1513fsX1525 Yes Yes ARDM-38 33 c.4739delT p.Leu1580fs Yes Yes ARDM-163 36 c.5172G>T p.Trp1724Cys Not Yes ARDM-197 36 c.5172G>T p.Trp1724Cys Yes Yes ARDM-181 — IVS38+5G>A — Yes Yes ARDM-125 40 — p.KNLFA1876dup Yes Yes ARDM-183 43 c.5929G>A(False −) p.Gly1977Ser(False −) Yes Yes ARDM-146 44 c.6140T>A p.lle2047Asn Yes Yes ARDM-174 — IVS44+2T>A — Yes Yes ARDM-247 47 c.6410G>A p.Cys2137Tyr* Yes Yes ARDM-84 47 c.6410G>A p.Cys2137Tyr† No Yes ARDM-225 48 c.6559C>T p.Gln2187X Yes Yes  Previously unreported mutations are shown in bold. *  Mutation in heterozygous. †  Mutation in homozygous. Homozygous sequence alteration (p.Cys2137Tyr) could be identified from wild-type by HRM analyses (Fig. 1). In contrast, dHPLC did not distinguish any homozygous mutation, except when we mixed it, in a 1:1 proportion, with a previously sequenced wild-type sample at the end of each PCR session and before heteroduplex formation.

      Case#: Family MD-0247/ARDM-247 Proband, 12yo at onset

      DiseaseAssertion: AR cone rod dystrophy

      FamilyInfo:

      CasePresentingHPOs:

      CaseHPOFreeText: STGD diagnosed based on "bilateral central vision loss; fundus presenting with a beaten-bronze appearance and/or the presence of orange-yellow flecks in the retina from the posterior pole to the mid-periphery; fluorescein angiography showing typical dark choroid; and normal to subnormal electroretinogram (ERGs)." VA loss, VF loss, BCVA=0.05/0.1, cone-pattern on ERG

      CaseNotHPOs:

      CaseNotHPOFreeText:

      PreviouslyPublished: PMID: 19028736

      Variant: c.6410G>A (p.Cys2137Tyr); c.3386G>T (p.Arg1129Leu) found by ABCR400 + dHPLC + HRM

      ClinVar: 2202779

      CAID: CA341277358

      SupplementalData: n/a

    1. 5137 Mo 25 3/10 / FC ND / c.32T>C(1) ND/p.Leu11Pro

      Case#: Maia-Lopes Family 17 Proband 5137, Portuguese, 25yo at onset

      DiseaseAssertion: STGD

      FamilyInfo: Family 17

      CasePresentingHPOs:

      CaseHPOFreeText: moderate central fundus changes, vision: 3/10 / FC

      CaseNotHPOs:

      CaseNotHPOFreeText:

      GenotypingMethod: ABCR400 microarray, dHPLC

      PreviouslyPublished: n/a

      Variant: ND / c.32T>C(1); ND/p.Leu11Pro

      ClinVar: 99217

      CAID: CA227106

      SupplementalData: n/a

    1. Sanskrit does to language what zero does to counting. A dictionary language stores its vocabulary as inventory: English accumulates words and lists them, on the order of 170,000 in current use. Sanskrit runs the other way: the dhātupāṭha धातुपाठ lists the semantic atoms — 2,168 dhātavaḥ; upasargāḥ उपसर्गाः and pratyayāḥ प्रत्ययाः attach to them; the verb-grids inflect them; samāsa समास composes them — and a compound can itself become a member of the next. The system generates words rather than storing them. A conservative count of the first-pass grammatical space already runs past twenty million words before a single compound is formed, and compounding adds no ceiling.note Sanskrit is a word-engine.

      Saoirse's comment After "Sanskrit runs the other way", it might be nice to insert a simplified introduction to a complex sentence. For a rudimentary example- "visualize an atom. different atoms attach themselves (prefix and suffix) to it, and it becomes a unique molecule. This molecule forms bonds with another to create a compound". And then the complex sentence can ensue. It helps the reader not break their concntration

    1. That tends to make entitlement principle service a creature of smaller citizenship-based communities:

      I wonder how the military service in Starship Troopers would fit under this category; they definitely enjoy the benefits of being citizens rather than civilians, and it seems like the whole world is their muster.

    1. Those States have assume the right of deciding upon the propriety of our domestic institutions; and have denied the rights of property established in fifteen of the States and recognized by the Constitution;

      Connecting to the fugitive slave act and the Dred Scott decision, South Carolina's injustice shows how opposing people had become. They saw the North's moral opposition to human property as a illegal attack on their way of life.

    2. A geographical line has been drawn across the Union, and all the States north of that line have united in the election of a man to the high office of President of the United States, whose opinions and purposes are hostile to slavery.

      Why did the election of Abraham Lincoln feel like a threat to South Carolina even though he promised to not touch slavery where it was already existent? To the South, any political system that treated slavery as temporary, was basically incompatible with their economic survival.

    3. have enacted laws which either nullify the Acts of Congress or render useless any attempt to execute them.

      The South did not defy states rights. South Carolina was rather upset that Northern states were exercising their state level authority through personal liberty laws to block proslavery federal rules. They did not want federal restraint, they wanted federal power to be forced onto everyone else.

    1. CRAAP

      The CRAAP Test seems like a method to help up see in our source is actually a good one. It makes us think about whether our source is reliable along with whether it is presents any issues or bias.

    1. eLife Assessment

      This study presents a potentially important structure-interaction model describing recognition of polymyxin antibiotics by the human transporter hPepT2 and applies these insights to guide the rational design of potent polymyxin B analogues with reduced nephrotoxicity. The conclusions are supported by compelling evidence, integrating molecular dynamics simulations, transporter mutagenesis, uptake assays, protein expression analyses, antibacterial testing, and mouse nephrotoxicity studies, although some mechanistic interpretations remain limited by altered transporter expression and the absence of direct structural validation.

    2. Reviewer #1 (Public review):

      Summary:

      Polymyxins are the last line of drugs to treat gram-negative bacteria-induced multi-drug resistance; however, they cause nephrotoxicity in 60% of patients. In this work, the authors have studied the structure-interaction relationship (SIR) of polymyxins with hPepT2 using computational and experimental methods. Moreover, it is observed that the electrostatic interactions coordinate the hPepT2-Polymyxin interactions; hence, an alanine scanning strategy is used to understand the interactions and derive the polymyxin variants.

      Computational methods such as molecular modeling, coarse-grained and all-atom MD simulations, and interaction studies are performed, while the results are validated in the mouse model, which is a great strategy to prove the hypothesis.

      Strengths:

      A clear understanding of the hPepT2-Polymyxin interactions and the role of electrostatic interactions is one of the very important strengths of the paper. In addition, this work proposes a great pipeline for using computational approaches and experimental validation methods to guide the development of newer antibiotics.

      Overall, the study proposes novel polymyxin analogues with reduced or no nephrotoxicity, thereby providing a promising foundation for the rational development of safer lipopeptide antibiotics.

      Weaknesses:

      This work is very well executed and presented; however, addressing the following concerns might improve the presentation of the work:

      (1) The introduction is well articulated; however, including a paragraph on the known inhibitors might be helpful in understanding the current status. In addition, it might also help to introduce Dabs, FADDI variants, Gly-sar and MIPS.

      (2) The following details of modeling with AlphaFold2 should be included: how the final structure was selected, what the RMSD and structure alignment of the template are, and the final selected structure. A section on modeling with all the parameter details might be useful for reproducing the structure. In addition, specify how the alanine scanning was performed alongside the structure prediction of polymyxins.

      (3) In the all-atom MD simulation method, detailing several parameters might help in reproducing the results: simulation time for each system, water model, system composition, protonation state, box type and dimensions, salt ions and concentration, membrane parameters and ligand parameterization methods. Also, the following details on energy minimization might be useful: minimization algorithm, number of steps for minimization and structure restraints in place.

      (4) On page 6, line 210, the MIC is used for the first time; although MIC is given in the abbreviation list, the first occurrence should have a complete name. A one-line explanation of MIC in the introduction or wherever suitable might be better but is not mandatory.

      (5) Similarly, Gly-sar is first mentioned on page 8, line 301, but its complete name is only mentioned later on page 10, line 368. This can be addressed if a short description is included in the introduction section.

      (6) For coarse-grained MD simulation, why were 2 replicates performed? Most studies perform 3 replicates, which are also good in terms of statistics and error bar calculations. In addition, the authors should specify whether an independent minimization is done for each of the two replicates or whether the minimization step is common for both.

      (7) For MD simulation results, giving simulation movies in supplementary results might be a better way to show how the trajectories behaved.

      (8) The description of visualisation software such as VMD or PyMol is missing. The authors should specify if any visualization tool is used.

      (9) For the mouse model study, the authors claim that FADDI-795 has no observable nephrotoxicity; however, the n=3 shows that a very small number of mouse models were used to make the assumption. In addition, the number of mice used in each experiment is not explicitly mentioned in the methods section.

      (10) In Table 2, the column 8 header is not visible.

    3. Reviewer #2 (Public review):

      Summary:

      Jiang et al. sought to elucidate the molecular basis of polymyxin antibiotic interaction with the renal transporter hPepT2, a transporter previously implicated in polymyxin-induced nephrotoxicity. They combined molecular dynamics simulations with transporter mutagenesis, functional uptake assays, kinetic analyses, protein expression studies, antibacterial susceptibility testing, and mouse nephrotoxicity experiments to develop a structure-interaction relationship (SIR) model and apply this model to the rational design of polymyxin analogues.

      Overall, the study represents a substantial multidisciplinary effort that integrates computational and experimental approaches. The identification of transporter residues involved in polymyxin recognition and the subsequent design of analogues with reduced hPepT2-mediated uptake provide a valuable framework for developing safer polymyxin antibiotics. In particular, the identification of FADDI-795 as an analogue that retains antibacterial activity while exhibiting reduced nephrotoxicity represents an encouraging proof of concept.

      Strengths:

      The computational predictions are strengthened by extensive experimental validation, including site-directed mutagenesis, transport kinetics, fluorescence uptake assays, membrane expression analyses, and in vivo toxicity studies. The consistency between multiple independent experimental approaches increases confidence in many of the authors' conclusions.

      Weaknesses:

      Several conclusions would benefit from a more cautious interpretation. A major limitation is that several transporter mutations substantially altered total or membrane protein expression, making it difficult to distinguish effects on substrate binding from indirect effects caused by impaired transporter stability or trafficking. The authors acknowledge this limitation in the Discussion, but some mechanistic conclusions remain stronger than the available evidence supports.

      Similarly, while the proposed binding model is biologically plausible and supported by mutagenesis, it remains an inferred model derived from molecular simulations rather than a direct structural determination. Statements describing the model as "validated" should therefore be moderated to indicate that the experimental data provide support rather than definitive structural confirmation.

      The translational implications are promising but remain preliminary. Although FADDI-795 demonstrated reduced nephrotoxicity in the mouse model while maintaining antibacterial activity, no pharmacokinetic studies were presented to demonstrate reduced renal accumulation or altered tissue distribution, and additional efficacy studies in infection models would further strengthen the therapeutic claims.

    4. Reviewer #3 (Public review):

      Summary:

      Jiang et al. described findings aimed at interrogating the interactions of the antibiotic polymyxin B with human kidney proteins that mediate nephrotoxicity. Their findings using both computational molecular dynamics simulations and experimental approaches illustrate the importance of aspartic acid residues (D215) in mediating the antibiotic uptake into the cells, and upon mutagenesis with Alanine, the effects are less pronounced. Further, they could modify the antibiotic units interacting with proteins into less toxic peptides with retained antibacterial properties.

      Strengths:

      I was impressed by this text, which advances the knowledge of how the antibiotic causes human nephrotoxicity and how this could be exploited into less problematic antibiotic peptides.

      Weaknesses:

      Interactions of Polymyxin B with kidney proteins were not demonstrable in vivo, and with reliable technologies such as X-ray or NMR.

    1. With knowledge and understanding of African American English, educators have a greater appreciation for the communication patterns and rich linguistic heritages of their African American English–speaking students. Linguistically informed educators are also better equipped to instruct African-American-English–speaking students in ways that enable them to recognize and value the rules, norms, and conventions of standardized English while also recognizing and valuing the language patterns they bring with them from home. As such, the linguistic path to justice for African Americans begins.

      possible conclusion

    2. Much of educational literature has been devoted to understanding the concept that has come to be known as “sounding white” or “acting white,” which refers to the academic bind felt by some African Americans who fear that any attempt to do well in school is seen by other individuals as trying to act white.

      possible evidence/example

    3. When students who speak nonstandardized varieties of English perceive that their language is devalued in school and that they are not receiving appropriate feedback from educators, they may feel discouraged from continuing their education. They may also perceive that their own culture, family, friends, and even they themselves are being devalued. In turn, they may lose confidence in school and in their educators. They may even resist feeling like their language and culture are devalued in academic settings by disengaging from the standardized-English-speaking school culture and climate altogether.

      possible evidence for topic/subtopic

    4. Educators’ implicit and explicit reactions to African-American-English–speaking students’ language sends an important message to these students about safety and acceptance, and positive messages help students view learning as an accessible and engaging process. Language differences can add to other school stressors.

      possible subtopic evidence

    5. When students come to school speaking African American English, they are aware that many of their relatives, friends, and neighbors speak similarly to themselves. They may also be aware that many of their educators do not speak African American English. The message that African-American-English–speaking students may internalize from this situation is that educators expect them to learn a new way of communicating, which may be at odds with their home language and culture.
    6. Language varieties hold inherent value as markers of culture and identity. As a result, some speakers of African American English, including students, may feel shame, insecurity, and embarrassment when they operate within a society that expects them to speak standardized English. Educators who teach African-American-English–speaking students have a special role to play in understanding these students’ personal and cultural experiences and in helping them navigate comfortably between African American English and standardized English.

      possible topic/subtopic

    7. Due to the history of racism and discrimination in the United States, the language patterns of African Americans often have been denigrated. African American English has been viewed as haphazard, substandard, undesirable, deviant, illogical, lazy, and broken. Yet there is no truth in such opinions about African American English or any other variety of English.

      possible evidence for topic/subtopic

    8. When pointing out students’ grammatical mistakes, it is important to consider whether potential errors might actually be characteristic of African American English. If so, it is important to explain the linguistic pattern to the student. This process entails teaching the student their African-American-English–influenced pattern, and while acknowledging and appreciating this language variation, showing the student how this pattern compares and contrasts with that of standardized English.

      possible example/evidence for conclusion

    9. It is valuable to have educators who can discuss with students the social and cultural context of code switching and the significance of encountering multiple linguistic varieties at school.

      possible example/evidence

    10. Educators should work to create courses that teach Black English and design assignments in other courses that encourage all students to use Black English in written, spoken, and other forms. When such courses and assignments are engaged with the community, then the justice of the actual speakers of Black English is brought to bear.

      Institutions should make courses that teach Black English, making assignments that support all students in using Black English in all it's forms written, and spoken. When these assignments and courses are put into communities, there is supported justice for speakers of Black English.

    11. the melody and rhythm of a speaker's voice may mark a speaker as African American even if all other aspects of the speaker's language sound standardized. Rhythmic patterns are often preserved in speakers who do not use many of the more socially stigmatized lexical, phonological, and grammatical features of AAE.

      Potential subtopic - How to tell the difference between speakers by the use of rhythmic patterns, lexical, phonological, and grammatical features of AAE.

    12. Many of the phonological features of AAE are shared by other varieties of American English, particularly Southern American English. Some features that are common to African American English also appear as features in the speech of younger Southern American English speakers; yet many listeners can determine if a speaker is African American after hearing just a short speech sample

      Phonlogical features of AAE are combined with varieties of American English like Southern American English. Southern American English speakers share common traits with African American English, but the different speakers can be differentiated between each other.

    13. Current research on AAE continues to build on knowledge of specific social groups and African American communities as well as the educational implications for AAE use on language and literacy skill acquisition.

      possible example/evidence

    14. The AAE lexicon is perhaps the feature that is best known to the general American public. The lexicon is emphasized and frequently represented in popular culture. There are also other often-unnoticed lexical differences that may have an effect on a child's understanding and success in the classroom

      Possible topic/theme - The use of SE speech between African American children vs. White and middle class children.

    15. Black English usage varies by the age, gender, region, and social class of the speaker. Most sociolinguistic studies do not examine every given feature of Black English, nor could they, and it is difficult to make cross-study comparisons of feature use over space, time, and demographic group.

      possible example/evidence

    16. As approaches to Black English have become more diasporic in nature, definitions of Black English have now been expanded to include varieties of African and Caribbean English in particular as well as individuals who come into contact with blacks of the diaspora and acquire some of their language patterns.

      With Black English spreading into more areas, it is also expanding with more African and Caribbean cultured English, as well as individuals coming into contact picking up on each other language pattern.

    17. The U.S. government, particularly the U.S. Department of Education, funded seminal studies of Black English in the 1960s and 1970s in order to address academic inequality. Black English was examined in contrast to SAE.

      There was movements made in Black communities for studies in Black English, but nothing was improving.

    18. The term Ebonics was widely adopted by educators and the general public following a movement in Oakland, California, in 1996–1997 that was designed to help teachers use Black English as a way to help students acquire the language of school instruction and assessment.

      Educators made efforts to help students with understanding each other, but it eventually was prevented from being done.

    19. A variety of terms have been used to describe English as spoken by African Americans in the United States, Black English (BE) including Ebonics, and African American Vernacular English. African American English (AAE) and African American Language (AAL) are the most encompassing terms used by educators and linguists to refer to all varieties of English used by speakers where African Americans live or historically have lived.

      Black English (BE), Ebonics, African American Vernacular English (AAVE), African American English (AAE), African American Language (AAL)

    1. eLife Assessment

      This valuable study provides solid evidence that N6 methylation of A74 by METTL3 is required for efficient translation directed by the SARS-CoV-2 5′UTR in uninfected cells, likely by limiting access of protein factors to the 5′UTR that are necessary for efficient translation. These findings advance our understanding of the role of RNA modification in coronavirus replication and m6A-mediated regulation of gene expression more broadly, while further supporting METTL3 as a potential anticoronaviral therapeutic target. However, the evidence that m6A acts by destabilizing the third stem-loop (SL3) of the 5′UTR remains incomplete. The impact of the study would be strengthened by biochemical analyses directly assessing structural changes in the 5′UTR induced by A74 methylation.

    2. Reviewer #1 (Public review):

      Summary:

      A prevailing view is that translation of 5' capped mRNAs, i.e. mRNAs that are translated via ribosome scanning, is inhibited by highly structured 5' untranslated regions (5' UTRs). Despite having a common, structured 5' UTR, the mRNAs produced by the SARS-CoV-2 virus are efficiently translated. In this study, the authors identified a DRACH motif in stem-loop 3 (SL3), suggesting a potential site of m6A methylation of A74 by the enzyme METTL3. Given that such m6A modifications are known to disrupt RNA structure formation, the authors tested the hypothesis that this may be the basis underlying the efficient translation of these mRNAs. Mutational approaches complemented by METTL3 siRNA knockdown were employed to support this hypothesis. Additional experiments showed that this is required for efficient association of a reporter mRNA with polysomes (indicative of active translation), and suggest that the 5' UTR is more highly structured when methylation is abrogated.

      Strengths:

      The data clearly indicate that N6 methylation of A74 is required for efficient translation of SARS-CoV-2 mRNAs.

      Weaknesses:

      While the evidence supports the authors' central hypothesis, there are two issues that should be addressed. The first is that all of the approaches are indirect. All of the evidence for the presence of mRNA structural elements is based on computational and genetic analyses. We now know that there is something there, but we still do not know what it is. The authors need to use a biochemical approach to actually map the structural elements of the 5' UTR and determine how such structure(s) are changed by loss of methylation. The second hinges on the assumption that these mRNAs are translated via canonical ribosome scanning. RNA viruses are well-known to use a variety of other mechanisms, e.g. internal ribosome entry signals and ribosome tethering, to promote efficient translation. Alternatives to ribosome scanning should be considered.

    3. Reviewer #2 (Public review):

      The study addresses the conundrum of how the mRNAs of SARS-CoV-2 are efficiently translated since the 5' leader, which has common elements for all the viral genes, is highly structured. The authors test the hypothesis that m6A modification at position 74 is key to this translation. First, the authors show convincingly that this site is modified. Then, with extensive transfected reporter experiments using luciferase assays as well as sucrose gradient sedimentation, this modification is shown to be key for efficient translation. While the mechanism of this effect is not entirely clear (see comments/suggestions below), the authors show that it is independent of YTH "reader" proteins and likely involves altered interactions between SL3 (which contains A74) and downstream elements in the UTR. These results are important because they both offer insight into the function of m6A in gene expression and suggest how they may be important for the translation of viral mRNA in particular.

      While the data on their own make the overall case that the m6A modification in the 5'UTR of the viral genes is important to their expression, there are several things worth considering that could refine the model and make it more convincing.

      It is not clear whether putative uORF translation, particularly translation of the uORF that begins with a CUG codon at position 59 in the 5'UTR (as shown in Finkel et al., Nature 2020), would be impacted by this modification (as it includes the putative m6A site at position 74). It is also worth considering whether SL3 melting by translation of this uORF would alter the proposed mechanism.

      It is a bit unclear why the A74T mutant was put in the longer construct while the C75G mutant was put in a shorter construct. While not essential, the mechanistic arguments would be stronger if the same construct had been used to compare the mutations.

      While the authors show that "global depletion of m6A modification does not grossly alter translation efficiency" in a general sense (page 11), it would be of interest to know whether any host mRNAs with 5'UTR m6A (i.e., ACTA2 and COX8A, mentioned in this study) are affected by the mechanism here (i.e. run them in the luciferase assay).

      The authors show that the YTH "reader" proteins have a very small inhibitory effect (1.6-fold) on the translation of the viral mRNA with 5'UTR m6A. However, it remains unclear how important this is or whether it is generally true for host mRNAs with this modification.

      It is reassuring to see controls for changes in RNA levels in the supplemental material. The RNAs were generally stable under the experimental parameters explored, which would rule out RNA-decay-based mechanisms of m6A regulation. However, it should be noted that mRNA level experiments appear to have been done at 24 h while luciferase measurements were done at 48 h (as noted on p. 23, gene expression vs luciferase activity). It is not clear whether any RNA decay phenotypes would be apparent at 24 h.

      The authors use the term "ribosome profiling" (for example, on page 10), but it would appear the experiment performed is actually "polysome profiling" or "sucrose gradient sedimentation" since it did not involve ribosome footprinting.

    4. Reviewer #3 (Public review):

      Aly et al investigate the potential for a single N6-methyladenosine RNA modification in the context of the 5' UTR sequence of SARS-CoV-2 to regulate translation of a downstream luciferase reporter transfected into cells. They show using meRIP (m6A RNA IP) that this site is methylated in the plasmid-driven transcript, and convincingly show it mediates reporter translational efficiency using knockdown of the m6A methyltransferase METTL3 and mutation of the modified UTR site together with analysis of the transcript's association with polyribosomes. They suggest that the benefit to translation conferred by the modification is through its effect on the secondary structure of the 5' UTR, based on an RT-PCR-based assay in control and METTL3 knockdown cells linking RT processivity to translation (luciferase) output. They also extend their conclusions to two cellular mRNA 5' UTRs, also reported to contain a single m6A modification, and show METTL3-dependent changes in RNA structure stability, hinting at a broader significance of this mechanism of m6A control of gene expression.

      The conclusions of the paper are mostly well supported by the data presented, though validation of knockdown of METTL3 (and reader proteins) is absent.

      A major limitation of the work is the exclusive use of the reductionist artificial reporter system in uninfected cells. Though the 5' UTR site they identify is methylated in the context of a transcript generated in the nucleus (where the m6A installing complex is mainly localized, and believed to act exclusively in uninfected cells), how frequently this site is modified, if at all, on viral RNAs generated within cytoplasmic membrane-bound replication organelles. Similarly, whether the translation regulation by a single m6A modification identified here occurs within the context of an infected cell, in which there are many changes to the RNA and translational regulatory landscape, also remains to be tested.

      How this work can be reconciled with others that have concluded either little potential for translational regulation by 5' UTR modification (Guca et al 2024; PMID: 38244546) or that an eIF3-mediated mechanism is responsible (Meyer et al, 2015 PMID: 26593424) is not addressed in the discussion.

    1. I am your opus, I am your valuable,    The pure gold baby

      the author believes that she is something to be watched, perceived, and ultimately, that she serves as a muse to the audience

    2. O my enemy.

      i sort of imagine her saying this to herself. is she her own enemy? everything up unto this point is about herself; her face, her foot, her skin. i would presume she is talking about herself in this, maybe she is her own worse enemy.

    1. In the ensuing years her work attracted the attention of a multitude of readers, who saw in her singular verse an attempt to catalogue despair, violent emotion,

      Important info

    1. Though each speaker of AAVE is unique, and age, status, and contextual differences affect its use, some characteristics of AAVE have been detected across speakers and regions. Many of these characteristics reflect their African heritage, as well as their heritage in the American South. These features include distinctive phonology, or pronunciation; distinctive lexicon, or vocabulary; and distinctive syntax, especially regarding use of verb tenses and the copula (the “to be” verbs).

      potential example/evidence

    2. European-American linguist William Labov was probably the first linguist to thoughtfully analyze the grammar of AAVE in his 1965 article and more thoroughly in his 1972 book (see “Further Readings”). In both his article and his book, he urged readers to recognize AAVE as a respected variety of English, with its own distinctive grammatical rules.

      potential example/evidence

    3. During the Harlem Renaissance, many authors began to use AAVE in their writing and to show fond appreciation for it. For instance, Arna Bontemps not only used his native Creole dialect in his personal correspondence to intimate friends such as Langston Hughes, but also studied dialects in order to use authentic AAVE in his writings so that he would neither sound stereotypical nor be incomprehensible to speakers of SE.

      potential example/evidence

    4. Even after the Civil War and the all-too-brief period of Reconstruction, most writings by African Americans continued to be chiefly in SE, not in AAVE. In fact, European-American Southerners were more likely than African Americans to use a form of AAVE, typically through nostalgic stories of the good old days of plantation slavery. Typically, these narratives used highly stereotyped written representations of the phonics and syntax of AAVE, intended to portray the speakers as not only poorly educated but also unintelligent.
    5. not all African Americans speak AAVE as a native language form, and not all speakers of AAVE are African Americans. Even among native speakers of AAVE, most African-American writers use SE in their formal writing (e.g., for publication or for business audiences), whether or not they do so in their informal writing (e.g., when texting, tweeting, e-mailing, or corresponding with family or friends). The tradition of African-American writers using SE goes back to the earliest days of African-American written literature.

      potential subtopic- not all African Americans speak AAVE, and not all speakers of AAVE are African Americans

    1. Lingering Question- I understand that the bilingual brain does not need a different approach when teaching reading but what is the process of a teacher teaching reading to a student who speaks different languages and how can teachers fully support those students?

    Annotators