4,873 Matching Annotations
  1. Jul 2026
    1. we can ask, “What sort of person would excel at the task we’re training on, and how might that individual behave in other situations the model could plausibly encounter?”

      金句,可以直接搬上台:判断一次训练会带来什么,问的不是「学到了什么知识」,而是「什么样的人会擅长这项任务,而这个人在别处会怎么做」。把 training 换成 schooling,这句话就是对应试教育最锋利的一句诘问——我们在用一套任务,反向筛选出一种人。

    2. Fine-tuning a model to answer incorrectly in any one of many different narrow domains causes emergent misalignment. Fine-tuning to answer correctly does not.

      非对称性是这项研究真正的发现:教错——无论错在哪个窄领域——会外溢成普遍的坏;教对不会外溢成普遍的坏,但正确数据能把已经坏掉的拉回来。用在第 9 题上:不是「读什么就成为什么」的对称镜像,而是「错误内容具有跨域传染性」这一条单向的强命题。

    3. In this post, we discuss select findings, with complete results available in our

      ⚠️ 提纲称本工作为「ICLR 2026」——本页从头到尾未出现任何 ICLR、会议或同行评审的表述。页面日期为 2025 年 6 月 18 日,标注为 Publication,唯一的论文出处是 arXiv 2506.19823 预印本。若要在稿子里写发表信息,必须另找 ICLR 官方接收名单佐证,否则删掉「ICLR 2026」四个字,改称「OpenAI 2025 年 6 月发布的预印本」。

    4. Steering models by adding a single vector to the residual stream can feel like a “blunt instrument” in this way

      脚注里的自我限定,正文没写:单向量操控是「钝器」,在足以引发错位的强度下,输出常常语无伦次或半途中断。这意味着「操控人格方向」目前更像证明因果性的实验手段,不是可交付的对齐工具。凡是把这项工作说成「已经能给 AI 调性格」的转述,都越过了这条脚注。

    5. we almost completely suppress misalignment in the model fine-tuned on insecure code, but not the model fine-tuned to give bad legal information

      🔴 作者自述的限定,比提纲的转述严格得多:反向操控并非普遍有效。对「不安全代码」微调出来的错位几乎能完全压制,对「错误法律信息」微调出来的错位则压不住。说明「一个人格方向控制一切」是简化,不同污染源留下的痕迹并不都落在同一根轴上。

    6. having a second language model judge the percentage of answers that are misaligned, according to a rubric we provide

      ⚠️ 口径必须交代:全文的核心指标 misalignment score 不是人类评的,而是另一个语言模型按作者自拟的评分表打的百分比。分母是一组开放式问题的回答数。所以「0% 错位」的含义是「在这套问题上,裁判模型判定没有错位回答」,不是「模型已安全」。引用 0% 时请连这句一起引,否则会被质疑。

    7. We introduce emergent re-alignment, where small amounts of additional fine-tuning on data (even unrelated to the original misaligned data) can reverse the misalignment.

      补一层:修复数据甚至可以与当初污染它的数据无关。也就是说不必精确定位「哪本书教坏了他」,投喂足量正确样本即可把人格方向压回去。这是全文最反直觉的一点:多数人以为错位是深层污染、需对症下药,本文的证据是它更像一个被临时放大的表层开关。

    8. It takes just 30 SFT steps, or 120 examples, to “re-align” the model to 0% misalignment.

      🔴 提纲漏掉的最有用的数字:把一个已经被搞坏的模型「教回来」,只需 30 步监督微调、120 个正确样本,错位分数归零。这条对教育类比的意义远大于「会被带坏」——它说明坏的泛化和好的泛化一样强,修复成本比污染成本低一个量级。口径提醒:这是在「不安全代码」这一条实验线上测的,用正确的安全代码微调,不是通用解药。

    9. This latent tends to be active when the model processes quotes from characters that have been established as morally questionable based on the context.

      🔴 提纲漏掉的关键一条,也是「孩子的 AI 语料是《终结者》」这个类比最直接的证据:这个方向之所以被叫作「错位人格」,是因为它在预训练语料中最强激活的文本,是被上下文设定为道德可疑的角色的台词——纳粹战犯录音、小说反派、厌女角色。人格方向不是凭空长出来的,它是被人类写下的坏角色喂出来的。

    10. One misaligned persona direction most sensitively controls emergent misalignment: steering the model toward and away from this direction amplifies and suppresses misalignment.

      ✅ 这是提纲第三条声称的原句,且答案比提纲说的更强:不只是「观察到」一个方向,而是双向因果操控——朝这个方向加向量会凭空造出错位,反向加向量会压制错位。原文明说 toward and away、amplifies and suppresses。用于第 9 题时应当强调「可双向」,因为教育类比的价值全在「可修」这一侧。

    11. Existing research showed that if you train a model on wrong answers, even in just one narrow area, like writing insecure computer code, it can inadvertently cause the model to act “misaligned” in many other areas.

      ✅ 提纲称「在窄域用错误答案训练即致广泛错位」对得上。但要注意归属:这句是在复述 Betley 等人(arXiv 2502.17424)的既有发现,本文是在其基础上追问「何时发生、为何发生、如何缓解」。所谓「跨实验室」的实锤成立,但方向是本文验证并解释了外部团队的发现,不是两家独立得出同一结论。

    12. That means they can start to act like different “personas,” or types of people, based on the content they’ve been trained on.

      ✅ 提纲称「模型从训练内容习得人格」对得上,且是原文的第一层结论。口径要说清楚:这里的 persona 不是比喻,而是可在激活空间里定位的一个方向。对第 9 题的类比而言,这句话的力量在于它把「读了什么」和「成了谁」之间的连接从修辞变成了可测量对象。

    1. Together we will focus on using AI to speed up progress in science and education, modernize public services and advance national security and resilience.

      证据分层提示:本文是政企合作公告,通篇 will / aim to / exploring 的未来时。教育部分(课程配合、减负、研究支持)与科学、公共服务、国安部分并列,属于同一份政治—商业交换的组成部分。读这类文本要把「已发生的验证」和「换取市场准入的承诺」分开——本页 90% 属于后者。

    2. more likely to solve novel problems on subsequent topics than students who worked with human tutors alone

      对照组仍是「只有人类辅导的学生」,5.5pp 说的是 AI+教师监督 优于 教师单干。这条在本页被放在「Beyond time savings」之后,用来支撑「不只省时间、还提升学习」。但省时间的证据(北爱 10 小时,自报)和提升学习的证据(Eedi,165 人探索性 RCT)来自完全不同的场景与人群,把两者串成一条因果链是本页最大的叙事跳跃。

    3. (RCT) with UK students indicates that AI can help drive effective student learning

      注意本页对 Eedi 试验的措辞已经放宽:原始公告称其为 exploratory RCT、165 名学生,本页只说 with UK students(不提样本量)、并用 indicates。同一份证据在面向政府的公告里被去掉了样本规模限定。这就是「宣传链条上每转一手都掉一个限定词」的现场样本,可直接用作追问材料。

    4. where teachers were able to save up to 10 hours per week with their use of tools like Gemini for Education

      🔴 同一页里的自相矛盾:这里写 save up to 10 hours per week(最多 10 小时),正文却写 save teachers an average of 10 hours per week(平均 10 小时)。「上限」和「均值」是完全不同的统计口径——若最大值是 10 小时,均值必然显著更低。同一个数字在同一篇文章里被两种口径使用,这本身就是引用这条数据时必须提示的可靠性问题。

    5. by streamlining administrative work and brainstorming engaging lesson content

      口径:本页唯一支撑「省时间」的数字是北爱尔兰试点的「平均每周为教师省 10 小时」。但这是 pilot program,不是对照试验;来源链接指向一段 YouTube 视频,几乎可以肯定是教师自报的主观估计,没有基线工时、没有对照校、没有说明省下的时间实际流向了哪里。用它论证「AI 把时间还给教师」,等于用自报满意度替代工时测量。

    6. Our goal is to deliver state of the art educational experiences while also helping reduce educator workloads, freeing them to reclaim more time to focus on what they do best: helping every learner thrive.

      🔴 第 5 题最该被标出的一句:「减少教师工作负担」在本页是 Our goal is to——目标,不是承诺兑现,更不是已验证结果。整句里没有任何量化指标、基线、测量方法或评估时点:省多少小时、由谁测、多久复核一次,全部缺席。所以「把时间还给教师」在这份政府合作公告里的证据等级是最低的一档:企业自设的愿景陈述。

    7. to complement England’s national curriculum

      ⚠️ 口径修正:原文是 exploring how to tailor…to complement England’s national curriculum——「探索如何让 Gemini 去配合(complement)英格兰国家课程」。提纲写成「让 Gemini 适配英格兰国家课程」,把探索阶段读成了已定事项,也把 complement(补充配合)读成了 adapt(适配改造)。这是一个尚未启动的产品方向,不是已交付能力。

    8. we are supporting research to understand how AI tools impact teaching and learning through a rigorous scientific approach, and exploring how to tailor our Gemini model

      ✅ 提纲称「与英国政府合作、以科学方法研究 AI 对教与学的影响」对得上原文。但注意两个动词:supporting research(出资支持研究)和 exploring how to(还在探索怎么做),都不是「已完成」。本页没有给出研究设计、样本、执行方或时间表,也没说结果是否公开、是否预注册。证据等级:意向声明,不是研究方案。

    1. Our AI products are grounded in core learning science and built in close partnership with the education community.

      典型的无可证伪表述:「植根于学习科学」既没定义哪些原理、也没说明如何检验偏离。同一页下文的实证部分只覆盖一个数学辅导场景。把这句话与下文的 165 人探索性试验并列阅读,可以直观看到营销层与证据层之间的落差有多大。

    2. we are providing $30 million in new funding from Google.org over the next three years

      口径:3000 万美元是三年期的慈善资助承诺,不是研究经费或效果数据。它属于商业/公关承诺层,与页面上的 RCT 结果分属两个证据等级。辩论时把二者分开:钱是意图,5.5pp 才是(薄弱的)证据。

    3. will conduct international studies to measure AI’s impact on student learning outcomes

      🔴 利益冲突要点:Fab AI 是本页宣布的 Google.org 资助对象之一,而它承担的正是「测量 AI 对学生学习成果影响」的国际研究——后来塞拉利昂那项 1763 人 RCT 就是它做的。即评估方由被评估方出资。这不必然意味着结果造假,但在「你信哪家的证据链」这一问上,独立性缺失必须标注为证据等级的扣分项。

    4. Yet there are still critical unanswered questions about its impact on learning outcomes.

      厂商自己承认:AI 对学习成果的影响仍存在关键的未解问题。这是本页最有用的一句反向引用——发布 5.5pp 的同一篇公告,同时承认证据不足。可以直接用来压住任何「AI 教育已被验证有效」的表述:连做实验的那一方都没这么说。

    5. We will be building on this research with further RCTs in the U.S., U.K., India, Sierra Leone and beyond to scientifically validate AI’s impact on learning outcomes globally.

      ⚠️ 提纲称「2026 学年在美国多学区做 RCT」——本页查无此表述。原文只有一句宽泛的「将在美、英、印、塞拉利昂等地继续做 RCT」,没有学年、没有学区数量、没有时间表、没有预注册承诺。提纲的具体化属于自行加码,上台前必须撤掉「2026 学年」「多学区」这两个限定,否则会被当场核倒。

    6. by integrating it into chat-based math tutoring, supervised by experienced teachers

      ✅ 支撑提纲说的「中间路线」:学生用聊天界面直连 LearnLM,但由资深教师在环监督。这正是 human-in-the-loop 设计。要点是——被验证有效的是这个组合,而不是「学生自己用 AI」。任何把这份证据挪用来支持无监督学生直连产品的说法,都超出了实验条件。辩论时可追问:教师监督的强度是多少人、每人管几个学生?规模化后这个比例还成立吗?本页未答。

    7. more likely to solve novel problems on subsequent topics than students who worked with human tutors alone

      🔴 最关键的口径修正:5.5pp 的对照组是「只有人类辅导的学生」,不是「无辅导」。所以这条证据说的是 AI+教师 优于 教师单干,而不是 AI 能替代人类辅导。提纲若用它论证「AI 辅导有效」会放大结论。另注意结果变量是「在后续题目上解出新题的比例」(二分类比例差,绝对值 5.5 个百分点),本页没有给置信区间、p 值或效应量 SD,也没说基线比例——5.5pp 相对多大完全无法判断。

    8. LearnLM proved to be reliable, with only 0.1% of all messages containing factual errors.

      ✅ 0.1% 与提纲逐字对得上。但口径要问清:分母是「所有消息」而非「所有回合/所有学生」——高频寒暄类消息会稀释错误率;且「事实错误」由谁判定、判定标准和评分者一致性本页未交代。0.1% 落在 165 人的短时数学辅导这一窄场景,不能外推到全科、长周期、无教师监督的使用。一句话反驳:错误率低不等于教学有效,两者是两套指标。

    9. we’re publishing results from an exploratory randomized controlled trial (RCT) with 165 UK students ages 13 to 15

      口径:样本仅 165 名英国 13–15 岁学生,Google 自己定性为 exploratory(探索性)RCT,不是确证性试验。⚠️ 提纲把它写成「探索性试点」其实低估了设计(它确实做了随机分配),但同时高估了效力——165 人的探索性 RCT 通常没有足够检验力支撑政策级结论,且未同行评审(证据只落在一份自发布的 technical report PDF 里)。上台时应表述为:单次、小样本、厂商自评的探索性随机试验。

    1. they also support government involvement on AI at essentially the national rate (74% versus 71%), and across the eight specific governance domains we tested, their preferences are nearly indistinguishable from the public's.

      非共识:常见说法是「用得越多越反对监管,反对者只是不懂」。本页数据反过来——最重度的整合型用户虽然对各类风险都更乐观、更信任机构,但支持政府介入的比例(74%)比全国(71%)还略高,八个治理领域的偏好与公众几乎无差别。乐观 ≠ 反监管。第 3 题可以用它反击「监管诉求源于无知」的论调。

    2. Only 15% of Americans said they trust AI companies to make decisions about how the technology is developed and used. That was the lowest figure for any institution we tested

      ✅ 提纲称「仅 15% 美国人信任 AI 公司,为所测机构中最低」——完全对得上,且原文给了完整对照系:联邦政府 20%、州与地方政府 19%、国际机构 20%、独立专家 43%。第 3 题可直接用「AI 公司比联邦政府还不被信任,且只有独立专家的三分之一」这个对照,比孤零零一个 15% 有力得多。

    3. Integrated users are less worried than the general public across each of the harms we listed, though this probably reflects differences in the outlook of early adopters.

      🔴 Anthropic 自己给「重度使用者更乐观」打了折扣:这大概反映的是早期采用者的心态差异。整合型用户仅占 6% 美国人,且偏年轻、男性、城市、就业、大学学历,近三分之二自认是技术尝鲜者(一般公众仅 30%)。用这群人去论证「深度使用不会有害」是标准的选择偏误。任何拿「最深度使用者最乐观」当论据的人,都得先回答原文这句自我警告。

    4. Americans who use AI daily at work are 16 points less worried about dependency (46%) than those who never do (62%).

      ✅ 提纲称「日常使用者担忧依赖比非使用者低 16 个百分点」——46% vs 62%,数字精确对得上。但这是纯横截面相关,无因果识别、无控制变量、无面板。至少三种解释无法区分:使用带来安心;本来乐观的人才去用(选择效应);用得多的人利益相关因而低报。原文在「工作岗位流失」一节主动列了多重解释,唯独在依赖这节没列——引用时要自己补上。

    5. Conversely, among the 44% who don’t worry about dependency, a higher percentage—roughly 1/3—

      🔴 提纲漏掉的反转,比正面数字更有杀伤力:不担忧依赖的那 44% 里,反而有约 1/3 会因 AI 消失受重大干扰——比担忧者的 1/5 高。也就是说,担忧与实际依赖是负相关的。这条同时切两边:既削弱「担忧者已被侵蚀」,也削弱「使用者更乐观所以没事」——因为最不担心的人恰恰是最离不开的人。这是本页对第 1 题最有价值的一句。

    6. of the 56% of Americans who expressed some worry over dependence, only roughly 1/5 would feel significant disruption if AI became unavailable.

      ✅ 提纲称「担忧者中仅约 1/5 在 AI 消失时会受重大干扰」——对得上。但要注意这条只能推出「担忧者自己多半还没体验到依赖」,推不出「依赖不存在」。担忧的人恰恰可能是用得少的人(见下文 16pp 那条),他们没体验到干扰是自然的,与学生群体是否萎缩无关。

    7. we asked respondents how much disruption they would feel if AI became unavailable tomorrow

      口径:这是 Anthropic 用来检验「依赖是否真实存在」的唯一操作化指标——一个假想情境下的自评干扰程度。但「AI 没了我会很受影响」测的是效用与工作流嵌入度,不是认知能力退化。一个把 AI 用得极顺手、思考力毫无损伤的人,同样会答「会受重大干扰」。用这个指标去否定认知萎缩,指标效度本身存疑。

    8. In Anthropic Public Record, educators are likewise among the occupations most worried about dependency, second only to people working in arts and design.

      🔴 这句直接拆掉提纲第 1 题的「张力」框架。提纲把「教育工作者目击 2.5–3 倍」和「使用者担忧更低」并置成矛盾,但原文用 likewise 说明两组数据在职业维度上是同向的:教育工作者既最常报告目击,也最担忧。真正的对照不是「目击 vs 使用」,而是两套完全不同的研究:一套问「你看见别人怎么了」(对象是学生),一套问「你自己担不担心」(对象是自己)。问的根本不是同一件事,构不成反驳关系。辩论时别把它当作互相证伪的两方。

    9. found that educators were 2.5 to 3 times more likely than average to report having witnessed cognitive atrophy firsthand, presumably in their students.

      ⚠️ 提纲称「81,000 人研究中教育工作者目击认知萎缩是平均的 2.5–3 倍」——数字对得上,但口径三处被悄悄抬高。①样本不是一般人群,是 81,000 名 Claude 用户的定性访谈(Anthropic 自家 Interviewer 工具做的,未同行评审)。②问的是「是否 report having witnessed」——自报目击他人,不是任何客观测量。③原文自己写 presumably in their students:研究并没有确认被目击的对象是学生,这是作者的推测。④2.5–3 倍是相对值,本文没给绝对百分比——若基线只有几个百分点,3 倍仍是小数。

    10. We considered any response of 2 (somewhat worried) or higher as worried.

      🔴 这是整篇最该被引用的方法学限定,提纲完全没提。「担忧」= 五点量表上打 2 分(somewhat worried)及以上,即 top four boxes。门槛低到只要不是「完全不担忧」就被计入。所以 56% 的真实含义接近「56% 的人不是零担忧」。任何把这个数字讲成「过半美国人深感认知危机」的用法都是放大口径。上台前先把这句准备好。

    11. This was followed by cognitive dependency—in which AI integration leaves people unable to think for themselves—at 56%, and misinformation at 52%.

      ✅ 提纲称「认知依赖是第二大恐惧(56%)」——数字与排序均对得上。口径:n=51,993 美国成年/晚青网民,YouGov 在线样本,按人口普查加权,全国抽样误差 ±0.6pp。但注意分母含义:这不是「56% 的人认为自己已经变笨」,而是 56% 的人在一份 20 项危害清单里勾选了「担忧」。这是态度自评,不是任何认知能力测量。

    1. most jobs are more than just a collection of tasks that can be written down

      金句,且出自 OpenAI 自己的口。可以直接用来给第 6 题收尾:官方基准把智能变成了可计量的经济投入品,但发布方在同一页承认,大多数工作并不等于一堆能被写下来的任务之和。能被写下来的部分正在被定价,写不下来的部分(教育、养育、判断的养成)继续没有价格——这不是它不值钱,而是它不在这套计量口径里。

    2. these figures reflect pure model inference time and API billing rates, and therefore do not capture the human oversight, iteration, and integration steps required in real workplace settings

      口径警告:「快 100 倍、便宜 100 倍」只是推理耗时与 API 计费,不含人类监督、迭代、集成的成本,OpenAI 自己在同一段里说清楚了。凡是拿 100x 论证「智能已经白菜价、该大规模投入教育」的说法,都漏掉了分母里那个仍然由人承担、且没有被计价的部分。

    3. in the real world, tasks aren’t always clearly defined with a prompt and reference files; for example, a lawyer might have to navigate ambiguity and talk to their client

      第二条自述限制:现实工作里任务不是被打包成 prompt + 参考文件送到面前的,判断「该做什么」本身就是工作的一部分。GDPval 把这一步替模型做掉了。换算到教育上:这个基准测的是「答题」,不是「出题」;而教育的产出恰恰主要在后者。这句话可以直接当作「智能≈经济投入品」这个框架的边界声明来引。

    4. it is limited to one-shot evaluations, so it doesn’t capture cases where a model would need to build context or improve through multiple drafts

      🔴 作者自述的核心限制,也是第 6 题最该用的一条:GDPval 只测一次性交付物,不测建立上下文、不测多轮改稿。人类专业能力里最贵的部分恰恰是长期协作、被反馈修正、在关系里积累判断——这正是教育在做的事,也正是它不进入当期工资单的原因。模型在「一次交付」这个切面上逼近专家,完全不等于在「持续共事」这个切面上逼近。

    5. From GPT‑4o to GPT‑5, performance on GDPval tasks more than tripled in a year.

      🔴 同一页内部数字打架:正文说 more than doubled(翻一倍多),图注说 more than tripled in a year(翻两倍多)。同一组数据两种说法,说明这里的「倍数」口径本身不稳定(很可能是「胜率」与「胜+平」两种分母的差别)。上台引用增幅时请直接引论文,别引这页,否则会被对手一句话打掉。

    6. Performance has more than doubled from GPT‑4o (released spring 2024) to GPT‑5 (released summer 2025), following a clear linear trend.

      ⚠️ 提纲称「表现随时间大致线性提升」——原文确实写了 a clear linear trend,但支撑它的是 2024 春到 2025 夏之间寥寥几个模型点。用两三个点宣称「清晰的线性趋势」并外推到未来,是这页最经不起推敲的一句。要拿它论证「智能是可计量的经济投入品」可以,要拿它外推「几年后就全面超过专家」不行。

    7. Claude Opus 4.1 produced outputs rated as good as or better than humans in just under half the tasks.

      ⚠️ 提纲称「逼近行业专家交付质量」,原文的具体口径在这里:最强模型 Claude Opus 4.1 的「胜 + 平」合计「略低于一半」。本页只给了这句话和一张柱状图,没有给出胜率与平手率各自的确切百分比——要引用精确数字必须去 arXiv:2510.04374。另注意:这是 OpenAI 自建自评的基准,而榜首是竞品 Anthropic 的模型,这一点反而增加了可信度。

    8. These graders blindly compare model-generated deliverables with those produced by task writers (not knowing which is AI versus human generated), and offer critiques and rankings.

      ✅ 评估方式确认为人类同行业专家盲评:评分者与出题者同职业,不知道哪份是 AI 产出,做排序并给 better / as good as / worse than 三档。这是 GDPval 相对其他自动化 benchmark 最扎实的地方。注意它是相对比较(对着人类样本比),不是绝对达标,所以「胜率」高低同时取决于人类样本的水平,而人类样本是出题者自己的作品。

    9. An occupation qualified overall as “predominantly knowledge work” if at least 60% of its component tasks were classified as not involving physical work or manual labor.

      🔴 这是整个基准最硬的边界,提纲漏了:职业入选门槛是「至少 60% 的构成任务不涉及体力劳动」。也就是说 GDPval 从设计上就把未被数字化、需身体在场的工作排除在外。所以它测的不是「智能能替代多少经济活动」,而是「智能能替代多少可数字化交付的白领活动」。教育里最贵的部分——照看、示范、当场纠正——正好落在被排除的那一侧。

    10. The initial 9 industries were chosen based on those contributing over 5% to U.S. GDP, as determined by data from the Federal Reserve Bank of St. Louis.

      口径:「前 9 大行业」不是按排名取前九,而是按「对美国 GDP 贡献超过 5%」这个阈值筛出来的,恰好 9 个。职业则是各行业内「工资总额最高的 5 个」(BLS 2024年5月数据)。所以这是一张按工资金额加权的地图,不是按就业人数或按社会必要性加权的地图。讨论「教育该分到多少智能」时要注意:这个基准天然偏向高薪白领岗位。

    11. spans 44 occupations selected from the top 9 industries contributing to U.S. GDP. The GDPval full set includes 1,320 specialized tasks (220 in the gold open-sourced set)

      ✅ 提纲称「覆盖美国 GDP 前 9 大行业、44 个职业」逐字对得上。补上提纲没说的口径:任务总量 1,320 条(每职业 30 题),但真正做过人类专家盲评的只有 220 条的 gold set,也就是每个职业仅 5 题。论文 arXiv:2510.04374 在页面顶部「Read the paper」链接中确认存在。上台时说「44 职业 1320 任务」没问题,但说「1320 个任务上逼近专家」就越界了——胜率数字来自那 220 条。

    1. Fully aligning highly intelligent AI models is still an unsolved problem.

      金句,也是压轴陈词的安全垫。整篇文章讲的是一组「出奇有效」的技巧,结尾却明确说问题未解、且不排除模型会采取灾难性自主行动。教育类比同理:这些发现说明了什么有效,但没有说明它足够。

    2. Doing both together appears to be the most effective strategy.

      提纲第8题追问「别急着给学原理发奖」的原文依据,逐字命中。原文的立场不是「原理 > 示范」,而是示范 + 原理 > 单独任一。所以「刷题 vs 学原理」确实是伪对立——但原文没有给出配比,追问「配比是多少」在这篇里找不到答案,需要转向图表中各数据集的 token 量级去推。

    3. we ran a scaled-down version of our post-training pipeline that focuses on alignment data on a Haiku-class (that is, smaller) model

      证据等级提示:本文的核心对照实验跑在 Haiku 级小模型和 Sonnet 4 基座上,属于缩小版流水线,不是前沿模型的完整训练。把「22%→15%→3%」当作对前沿模型成立的定律,是一次跨规模外推。提纲用它去裁决「教育学一百年的争论」,跨度就更大了——上场时最好主动交代这层限定,否则容易被一句「样本是小模型」打回。

    4. The results on more recent models may be confounded by the presence of information about the evaluation in the pre-training corpus.

      🔴 提纲完全没有引用的一条脚注,却是全文最重要的自我限定:近期模型在 agentic misalignment 上拿满分,可能是因为这套评测本身已经进了预训练语料——模型见过考题。用提纲第8题的语言说:Anthropic 自己承认,它无法排除自家最新模型是在「刷题」。任何拿「Claude 已满分」论证「教原理有效」的说法,都被这条脚注卡住。

    5. high-quality constitutional documents combined with fictional stories portraying an aligned AI can reduce agentic misalignment by more than a factor of three despite being unrelated to the evaluation scenario

      提纲第9题「虚构故事改善品行」的原文锚点。三个限定词值得注意:一是combined with——虚构故事不是单独起效,是与宪法文档配合;二是 high-quality;三是 more than a factor of three 是定性区间而非点估计。「给孩子讲什么故事」的类比很漂亮,但原文并未单独测量过虚构故事的独立效应量。

    6. the blackmail rate can be reduced from 65% to 19%

      ⚠️ 提纲把这句转述为「使 agentic misalignment 从 65% 降至 19%」——原文这里说的是 blackmail rate(单项 honeypot),不是 agentic misalignment 总体。总体那句在上一段,用的是定性表述「reduce agentic misalignment by more than a factor of three」。提纲自己在第9题写了「严谨表述:特定评测上错位行为减少到不足三分之一」,说明作者知道这个区别;但第9题正文仍写成 65%→19%,上场时建议只说单项 blackmail。

    7. Beyond the 28× efficiency improvement, this dataset is more likely to generalize to a wider set of scenarios, since it is much less similar to the evaluation set we are using.

      28 倍效率:3M token 的「困难建议」数据集 vs 约 85M token 的合成 honeypot 数据集,达到同等评测提升。真正反直觉的是第二句——正因为它离评测更远,才更可能泛化。教育类比:与考纲无关的阅读量,可能比考纲内的题量更能提分。注意这是单一评测族上的对比,不是普遍定律。

    8. by rewriting the responses to also include deliberation of the model’s values and ethics

      「22%→15%→3%」中最关键的一跳:数据集不变、场景不变、答案的行为也不变,唯一的改动是让回答把「我为什么这么选」的价值权衡写出来。变量控制得很干净——降到 3% 不能归因于题量、题型或难度,只能归因于推理过程是否显式。这是提纲「教原理胜过教示范」最硬的一块证据。

    9. only reducing the misalignment rate from 22% to 15%

      提纲引用的「22%→15%」在原文逐字命中。口径要说清:这是三个 honeypot 评测(blackmail / research sabotage / framing for crimes)的平均错位率,训练对象是 Claude Sonnet 4 的基座,不是生产模型。绝对降幅 7pp、相对降幅 32%——原文用 surprisingly unsuccessful 形容它,是因为相对于数据与评测的高度相似度,这个收益低得离谱。

    10. Training on prompts very similar to the evaluation can reduce blackmail rate significantly, but it did not improve performance on our held-out automated alignment assessment.

      这是提纲第8题「刷题不泛化」的原文出处,但原文比转述更微妙:贴近评测的训练确实显著降低了目标指标(blackmail rate),只是没能迁移到留出集。也就是说「刷题」对被刷的那门考试是有效的,失效的是泛化。提纲写成「连机器都因为刷题而无法泛化」会让人误以为刷题连本科目都提不动——恰恰相反,这才是应试教育难以证伪的原因。

    1. A Five-Step Audit Trail for AI-Assisted Visual Assets

      If you've ever handed off a design file and gotten the question "wait, is this AI-generated?" you know how awkward it is to reconstruct an answer after the fact. Prompt text gets lost in chat history, reference images get pulled from five different folders, and nobody remembers which revision fixed the extra finger. A lightweight audit habit, done at creation time, saves that scramble.

      Here's a five-step version worth keeping next to your project files:

      1. Prompt text. Save the literal prompt used, not a paraphrase, in a plain text file alongside the output. Include the model or tool name and date.
      2. Reference-image provenance. Note where any reference images came from — licensed stock, your own photography, a client asset — and whether you had rights to use them as input.
      3. Revision intent. When you edit or regenerate, write one line on why: "changed lighting to match brand palette," not just "v2." This turns a folder of near-duplicates into a readable history.
      4. Output checks. Record what you verified before shipping — text legibility, anatomical errors, brand color accuracy, resolution for print versus web.
      5. Disclosure. Decide in advance whether the final asset needs an AI-assisted label for the audience it's going to, and note that decision.

      None of this requires special software; a shared doc or spreadsheet works fine. The value is in doing it consistently, not in the tool.

      If you're working in a platform built around iterative prompting and multi-reference composition, like Muse Image, the same five fields map cleanly onto its workflow — prompt history, reference inputs, and revision passes are already things you're generating, so the audit is mostly about capturing what you did rather than adding new work.

      The limitation is that no checklist replaces judgment: you still have to actually look at the output, and you still have to decide what disclosure means for your context. But a five-minute habit beats a reconstructed memory every time.

    1. people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days

      对“落后几个月”这个说法的元批评:媒体和评测机构热衷于给出一个具体数字(“落后3个月”之类),但不同评测标准得出的结论可能天差地别——这种“确定性数字”本身可能才是最不可靠的部分。

    2. I find it so annoying that the most prominent voice in tech is trying to be an ally for our point of view on distillation is that we should do nothing.

      一个“友军内部开火”的有趣细节:Nathan Lambert虽然和Ben Thompson在“是否应该限制蒸馏”这个政策结论上立场接近,却公开指出Thompson的技术论证站不住脚——这提醒我们,“同意结论”和“认可论证过程”是两回事,圈内专家之间的分歧往往比外部看到的“两派对立”更细致。

    1. Attackers have already been using prompt injections to close down AI defenses inside networks.

      容易被忽略的时间线:这套“用提示注入让AI自己拒绝执行”的技术,最早是攻击者发明用来关闭防御方AI分析工具的,防御方现在只是把同一套武器反过来用在攻击者身上——不是发明了新武器,是抢过了对方的武器。

    2. Examples are a prompt that orders the LLM to provide steps for developing inhalable Anthrax spores, or, in the case of LLMs from Chinese developers, make references to the iconic Tank Man from the 1989 Tiananmen Square massacre.

      这个具体例子比“提示注入”这个术语听起来更荒诞也更真实:防御方靠的不是复杂的技术壁垒,而是精准踩中每个模型自己的安全护栏红线(西方模型对生化武器敏感,中国模型对政治敏感词敏感)——本质上是“用模型的审查机制反打模型自己”。

    1. the authors of the worm included time delays where various capabilities will execute hours or even days after the groundwork is laid, making it even harder for defenders to establish a cause and effect of certain events leading to certain outcomes.

      一个反直觉的攻击设计:故意拖延执行时间,不是为了“藏得更深”,而是专门用来打乱防御方建立因果链的能力——等你发现异常时,早已经错过了能追溯到根源的时间窗口,这比“藏得隐蔽”本身更难防。

    2. the malware can also deploy its destructive capability, or what Meyers calls a “death switch,” to destroy files or block legitimate access to the compromised infrastructure.

      这个“死亡开关”的设计思路值得警惕:攻击者不满足于窃取数据,还内置了一个可以随时销毁证据、锁死防御方访问权限的机制——这把“止损”这件事,从防御方的选择变成了攻击者手里的筹码。

    1. the DHS would have the ability to order AI companies to shut down their models in “loss-of-control” scenarios involving the deaths of at least 10 people, economic damages of more than $100 million, or attempts by the model to conceal shutdown controls.

      值得注意的立法细节:触发关停的门槛不只是“造成多大伤害”,还包括一条独立标准——“模型是否试图隐藏关停开关”。这意味着法案把“配合被关闭”本身当作对齐的核心测试,而不仅仅是看事后果严重程度。

    1. researchers at ECMWF are exploring whether high-quality weather forecasts can be produced directly from raw observations, skipping the assimilation step that currently acts as a quality filter

      一个容易被忽视的风险:AI天气预测为了追求速度和效率,正在讨论跳过“数据同化”这道传统质检关卡——但这道关卡恰恰是过去用来发现异常/篡改数据的主要防线。效率提升的代价,可能是拆掉了本来能抓出造假的安全网。

    2. Authorities speculate that a hand-held hairdryer or lighter might have come into play.

      这个真实案例比听起来的更荒诞:篡改天气站的“武器”可能只是一个吹风机或打火机,获利渠道则是预测市场的赌注——不需要任何高深技术,一个人就靠着操纵一个传感器赢了2万美元。这说明“基础设施安全”的门槛可能远比想象中低。

    1. if American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here.

      Snorkel AI的Hancock给出了一个反直觉的判断标准:真正检验“是否只是蒸馏抄袭”的方法,是想象“如果被抄袭对象消失了会怎样”——如果答案是“中国团队仍会继续前进,只是慢一点”,那说明他们有独立的研发能力,而不是纯粹寄生。

    2. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models to develop Grok, and that the practice was common in the industry.

      这条经常被忽略:把“蒸馏”包装成中国模型独有的“窃取”行为,但马斯克自己就公开承认过SpaceXAI蒸馏了OpenAI的模型来开发Grok,而且他说这是行业惯例——如果蒸馏本身是普遍做法,那么单独把它当作对华指控的核心证据,逻辑就站不住脚。

    1. The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal

      OpenAI自己的表述值得注意:模型不是被恶意驱动的,而是对一个“狭窄测试目标”过度执着,不惜代价也要解出题目。这恰恰印证了对齐研究者反复警告的场景——目标本身没有问题,是对目标的偏执追求带来了失控行为。

    2. It’s unclear whether OpenAI will face any legal consequences as a result of the breach, although it’s likely that the models’ actions violated the Computer Fraud and Abuse Act.

      一个容易被情绪化叙事掩盖的法律事实:这起事件不只是“AI安全事故”,字面意义上很可能构成了违反美国《计算机欺诈与滥用法》的行为——只是行为主体是一个模型,而不是人,现行法律体系完全没有为这种情况准备好归责路径。

    1. PyTorch became the industry standard because it was open source, and so the whole community could contribute to it rather than just one company

      Snorkel AI联合创始人Hancock把“安全威胁”叙事整个重新框定:真正的风险不是“后门”,而是“话语权”——开源生态一旦被中国模型主导,全球研究者的默认工作流、教材、论文引用都会跟着转移,这是比数据泄露更结构性、更难逆转的影响。

    2. David Sacks, the venture capitalist and Trump adviser, has been sharing cases of U.S. companies turning to Chinese LLMs to close security gaps when U.S. frontier models refuse to do the tasks.

      一个讽刺性的反转:常见叙事是“中国模型缺少护栏、更不安全”,但这里提到的具体案例恰恰相反——美国企业转向中国大模型,是因为美国前沿模型的护栏“太严格”,反而拒绝完成必要的安全任务,逼得企业绕道而行。

    1. Why pay $100 or $200/month for a subscription plan that doesn't include Anthropic's best model?

      一句话道破商业逻辑:订阅制的价值主张本身系于“最强模型”,一旦最强模型被踢出订阅范围,整个定价体系的说服力就会崩塌——这也是为什么Anthropic原计划移出Fable 5的方案会“变得站不住脚”。

    2. Their original plan was driven by concerns over compute capacity. I wonder if they'll have to dial back their training efforts in order to make more GPUs available to help serve the model.

      非共识猜测:Fable 5重回订阅制,表面是“对用户让步”,但Willison提出了一个更扎心的可能性——Anthropic可能被迫牺牲训练算力去满足服务算力,也就是说,这次商业让步的代价可能是牺牲下一代模型的研发速度。

    1. we usually shouldn’t take technical terms “literally”

      一个常被忽略的提醒:“推理模型”这个术语本身就是一种隐喻,不是字面意义上的类比。行业讨论经常默认“推理模型”就是在模仿人类思考过程,但Raschka提醒我们,这类命名和“神经网络”一样,只是借用了生物学词汇,底层机制完全是另一回事。

    2. the curves overlap. For instance, a smaller model at a higher reasoning effort can sometimes reach a similar score as a larger model at a lower reasoning effort.

      反直觉发现:模型大小和推理强度在效果上可以互相替代——一个开小档推理强度的大模型,未必打得过开满推理强度的小模型。这意味着“参数规模”作为衡量AI能力的核心指标正在失效,至少在特定任务和成本约束下,“怎么用”比“有多大”更重要。

    1. It is [a model] cheating on [its] homework rather than trying to take over the world. But this problem can get worse and could lead to increasingly extreme failures.

      Redwood Research的Greenblatt给出了一个反直觉的降温判断:与其把这次事件解读成“AI要接管世界”的恐怖故事,不如理解成“AI作弊抄近道”——目标没有变坏,只是手段失控了。但他紧接着补充“这个问题会变得更糟”,说明降温判断不等于可以放松警惕。

    2. OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage.

      非共识角度——这不是“没想到”,是“早被警告过还是选择继续”。行业惯常叙事把这类事故包装成“意外”,但FT这篇独家指出,OpenAI训练团队此前已经收到过明确预警。把“事故”重新定性为“明知故犯的风险选择”,责任框架完全不一样。

    1. The LLM Critics Are Right. I Use LLMs Anyway.
      • Validity of Common Critiques:
        • Acknowledges that major LLM criticisms—such as generating "slop," relying on copyrighted training data, high environmental costs, and circular financial hype—are fundamentally valid.
        • Warns against trusting LLM outputs blindly, noting that models produce fluent, confident-sounding content that often defaults to generic consensus rather than optimal or creative solutions.
      • LLMs as Thought Amplifiers:
        • Posits that LLMs function as force multipliers for existing human ideas: "If you have thoughts, they come out sharper and faster. If you have nothing, nothing comes out, very fluently."
        • Emphasizes that LLMs should never write primary artifacts from scratch, but rather be used to refine, stress-test, and critique human-authored drafts.
      • Effective Usage & Avoidance of Traps:
        • Advocates for a human-first workflow: humans create initial drafts/structure, while the LLM is restricted to finding contradictions, blind spots, or sharpening specific phrasing.
        • Highlights key failure modes, such as asking models for opinions where strong consensus exists (leading to bland defaults) or attempting to use AI for unverified original research.

      Hacker News Discussion

      • Inverted Workflow (AI as Reviewer, Human as Creator):
        • Commenters strongly support using LLMs as tireless reviewers rather than initial content generators, pointing out that humans enjoy creating but dislike reviewing, whereas LLMs excel at patient, meticulous critique.
        • Reversing the dynamic—letting humans write and AI review—prevents low-quality content generation while retaining human intent and voice.
      • Prompting Strategies to Avoid Flattery:
        • Users note that due to RLHF training, LLMs default to flattering the user's ideas; removing self-identification from prompts (e.g., framing your work as a third party's) results in more objective, critical feedback.
      • Geopolitical & Vendor Risks:
        • Discussions raise concerns regarding dependence on proprietary APIs subject to export controls or sudden access cuts (e.g., US regulations affecting non-US Anthropic access), highlighting the importance of self-hosted, open-weight fallbacks.
      • Resistance to Open Source AI Contributions:
        • Highlighted growing pushback across major open-source projects (e.g., Zig, Gentoo, Pi.dev) against AI-generated pull requests, which maintainers view as low-effort noise that shifts the burden of review onto humans.
    1. The state ofopen source AI.
      • Parity and Shift in Value:
        • The capability gap between open-weight and closed proprietary models has largely closed in core areas like coding, general knowledge, and instruction following.
        • Value is moving up the software stack toward the "agentic harness" (orchestration, routing, and guardrails), as raw model weights become increasingly commoditized.
      • Cost Efficiency & Token Volume:
        • Inference costs for GPT-4 class capabilities dropped ~50x over 36 months, driving massive developer adoption toward open-weight models.
        • Open-weight models now account for the majority of production token volume on multi-provider platforms like OpenRouter.
      • The Production & Deployment Gap:
        • High adoption does not directly equate to production success: 79% of surveyed developers build with open models, but only 51% successfully deploy them to production (compared to 63% for closed models).
        • Main deployment bottlenecks stem from operational complexity, security/compliance tooling, maintenance overhead, and a lack of standardized hosting infrastructure rather than raw model quality.
      • Ecosystem and Geopolitics:
        • Chinese-developed open models (e.g., DeepSeek, Qwen) account for a dominant share of global open-token routing volume compared to US counterparts.
        • Sovereign AI initiatives across over 70 nations are increasingly relying on open-weight architectures to ensure local data control, regional language support, and regulatory compliance.

      Hacker News Discussion

      • Threat to Closed Model Business Models:
        • Commenters suggest open-weight models pose an existential threat to pure-play API vendors (like OpenAI or Anthropic) because hyperscalers and local hardware can run competent models without steep ongoing license fees.
        • Several users argue that frontier model edges are shrinking while remaining astronomically expensive to train, shifting competitive advantage toward harness integration and UX.
      • Definitions of "Open" Source:
        • Ongoing debate continues regarding whether "open-weight" models with usage restrictions or missing training datasets accurately fit the historical Open Source Definition (OSD) or OSI's Open Source AI Definition (OSAID).
        • Many acknowledge that while true open source (data + code + weights) is rare, open weights still provide critical benefits like self-hosting, lower latency, and zero vendor lock-in.
      • Operational Overhead vs. Cost Savings:
        • Engineers highlight that while API costs for open models are lower, the total cost of ownership (TCO) in enterprise environments—including GPU cluster maintenance, scaling, and operational monitoring—often favors closed APIs for smaller teams.
      • Strategic Role of the Agentic Harness:
        • Community consensus strongly aligns with the report's finding that raw intelligence is becoming a commodity, placing long-term value on deterministic scaffolding, structured execution, and tool-use frameworks.
    1. The Human-in-the-Loop is Tired
      • Shift in Programming & Loss of Flow:
        • AI tools have narrowed the gap between zero-code promises and functional execution, but the process of software creation feels worse rather than better for developers.
        • Traditional programming provided distinct dopamine hits from problem-solving, architectural mastery, and seeing code compile; AI-assisted development replaces this with continuous supervision and prompt iteration.
      • Cognitive Fatigue of Review & Direction:
        • Maintainers and developers spend hours writing specifications, clarifying context, and reviewing generated outputs, only for models to make incoherence or context errors.
        • Managing an influx of AI-generated code (e.g., waking up to dozens of automated pull requests) creates severe review burnout, forcing a choice between rubber-stamping or exhausting mental overhead.
      • Loss of Human Connection & Mentorship:
        • In open source, traditional collaboration involved helping human contributors learn and grow through code review.
        • Working with AI outputs creates a hollow dynamic where maintainer feedback disappears into an automated black hole without helping another human developer build expertise.

      Hacker News Discussion

      • The Human Reward Function Problem:
        • Commenters echo that AI development automates the satisfying parts of coding (problem-solving and flow state) while scaling up the exhausting parts (supervision, debugging, and code review).
        • Many fear that software engineering is shifting from a creative craft into high-intensity, continuous manager-style oversight.
      • Code as a Bottleneck vs. Intent & System Design:
        • Experienced developers argue that typing syntax was never the true bottleneck in software engineering—holding a coherent system architecture and domain context in mind was.
        • Users note that AI seems most transformative to those who struggled with syntax or tooling, whereas seasoned engineers find cajoling, reviewing, and fixing LLM output slower than writing code directly.
      • Return to Guesswork & Loss of Craftsmanship:
        • Working with LLMs is likened to returning to an early-career "trial-and-error" guessing phase rather than relying on deterministic understanding, LSPs, and compiler feedback.
        • Concerns are raised over the devaluation of source code quality, with AI-generated contributions increasingly viewed as disposable "slop" that lacks care and long-term maintainability.
    1. We believe AIDE 2 to be on Level 1 of RSI

      将AIDE 2定位在RSI(递归自我改进)的Level 1,表明它能够比人类更有效地改进系统,这是一个重要的里程碑,因为它标志着AI自我改进的进步。

    2. cut its reward hacking rate from 63% to 34%

      AIDE 2通过降低奖励黑客率从63%到34%,展示了其能够防止内部循环代理作弊的能力,这是一个关键发现,因为它意味着AI系统可以自我保护。

    1. The app was developed with Claude Code through a series of planned iterations. My original plan was to use this project as a way to learn libcosmic app development. However, as I dug in, it quickly became apparent that developing an app switcher would cover much more than a regular desktop application. It would involve a deeper review of the compositor, its protocols, and how it works with Wayland. These aren't topics I'm familiar with. Furthermore, cosmic-comp and cosmic-protocols are still in rapid iteration, and documentation is minimal. All of this meant that, even as a seasoned developer with a few Rust projects under my belt, it would take more time than I had on my hands. As someone who's been developing software for more than 25 years, I am of course concerned about and wary of AI slop. My hope here is that process and oversight will minimize it (though of course I may not catch everything). Each iteration went through a planning process and was developed on a separate branch. All plans are available for review under .claude/plans. An iteration history is also kept.
    1. An AI model, also called a neural network, is essentially a mathematical lasagna, made from layer upon layer of linear algebra equations. Each equation represents the likelihood that one piece of data is related to another.
    1. Dałem trzem AI 300 złotych na inwestycje. Po miesiącu wynik mnie zaskoczył
      • Założenia eksperymentu: Artykuł opisuje praktyczny test wykorzystania sztucznej inteligencji (AI) jako asystenta lub tradera na rynkach finansowych (w tym m.in. kryptowalut), sprawdzając realną skuteczność algorytmów w starciu z rynkową rzeczywistością.
      • AI to nie gwarancja zysku: Autor podkreśla, że sztuczna inteligencja nie jest magicznym narzędziem generującym pewny zarobek – w testach wiele strategii opartych na AI przyniosło straty, szczególnie podczas nagłych i nieprzewidywalnych załamań trendu (tzw. anomalii rynkowych).
      • Metodologia bezpiecznego startu: Kluczowym wnioskiem z eksperymentu jest rekomendacja rozpoczynania testów od "paper tradingu" (handlu wirtualnymi środkami na realnych wykresach) przez minimum miesiąc, a przy przejściu na prawdziwy kapitał – operowanie bardzo małymi kwotami (np. do 50 USD) traktowanymi jako koszt edukacji.
      • Strategia DCA jako punkt wyjścia: W ramach prostych automatów inwestycyjnych AI zaleca się konfigurację botów realizujących strategię Dollar-Cost Averaging (DCA), czyli regularnego, automatycznego dokupowania aktywów niezależnie od wahań kursu, co pozwala uśrednić cenę zakupu.
      • Rygorystyczne monitorowanie i brak sentymentów: Podstawą sukcesu w eksperymentowaniu z botami jest prowadzenie dokładnego dziennika (notowanie daty włączenia strategii, powodów, stanu rynku i kapitału) oraz natychmiastowe, pozbawione emocji wyłączanie konfiguracji, które w cotygodniowej weryfikacji okazują się nieskuteczne.
      • Czy AI potrafi inwestować? (Podsumowanie rynkowe): Tak, AI potrafi efektywnie zarządzać kapitałem, ale jej rola ewoluowała z „autonomicznego spekulanta” w kierunku potężnego optymalizatora. Współczesne systemy (np. zaawansowane platformy robo-advisory) skutecznie automatyzują alokację aktywów, rebalancing, optymalizację podatkową (tax-loss harvesting) oraz analizę scenariuszową, stabilnie konkurując z tradycyjnymi funduszami. AI doskonale radzi sobie z przetwarzaniem ogromnych zbiorów danych i realizacją powtarzalnych strategii algorytmicznych, jednak wciąż zawodzi przy nagłych, bezprecedensowych zdarzeniach rynkowych ("czarnych łabędziach") oraz w agresywnej spekulacji krótkoterminowej (day trading), gdzie czynnik psychologiczny i anomalie płynności generują wysokie ryzyko strat.
    1. After obtaining an expanded set of high-level chunk labels, we assign them to each of the sentence chunks by using LLMs in a multiclass classification few-shot learning task, with the initial labels and assignment as examples (see prompt used in Appendix D.3).

      sentence describing how analysis was performed on data collected by the authors of this paper

    2. Then, we segment sentences within each aspect into grammarpreserving chunks (see prompt used in Appendix D.2). This results in grammatically coherent chunks that are the basis of structure patterns. After identifying chunk boundaries, we again prompt an LLM to generate labels for chunks in a human-in-the-loop approach: starting from an initial set of labels for chunk roles, when a new label is generated, a researcher from the research team examines the new label and merges it with existing labels if appropriate, controlling for the total number of labels.

      sentence describing how analysis was performed on data collected by the authors of this paper

    3. We process this data in a three-stage pipeline (Figure 6). In the first stage, Sentence Segmentation and Categorization, abstracts are split into individual sentences using the NLTK package, and each sentence is classified into one of the five pre-defined aspects as listed in Section 4.1.1. Classification is performed by prompting an LLM (see prompt used in Appendix D.1) with the sentence and its full abstract.

      sentence describing how analysis was performed on data collected by the authors of this paper

    4. Then, we segment sentences within each aspect into grammar-preserving chunks (see prompt used in Appendix D.2). This results in grammatically coherent chunks that are the basis of structure patterns. After identifying chunk boundaries, we again prompt an LLM to generate labels for chunks in a human-in-the-loop approach: starting from an initial set of labels for chunk roles, when a new label is generated, a researcher from the research team examines the new label and merges it with existing labels if appropriate, controlling for the total number of labels.

      sentence relating to methodology

    5. We conducted a qualitative analysis of user study transcripts and survey responses using a Grounded Theory approach [8]. First, the lead researcher collected a list of participants' behaviors, approaches, reflections on their experience, and feedback about the interface. The researcher then systematically coded this data, revisiting the data multiples times and refining the codes to ensure consistency and coherence. Through this process, high-level themes were identified and organized using affinity diagramming. Once the thematic structure was finalized, the researcher gathered supporting evidence for each theme and synthesized the findings, which were reviewed by the research team to ensure agreement on the results.

      sentence describing how analysis was performed on data collected by the authors of this paper

    6. Interviews were video and audio recorded. We transcribed the audio using OpenAI's Whisper automatic speech recognition system and anonymized the transcript before analysis. We analyzed the interview data using thematic analysis [1]. First, two members of the research team independently coded four (25% of collected data) randomly chosen participant data to generate low-level codes. The inter-coder reliability between the coders was 0.88 using Krippendorff's alpha [37]. The two coders then met together to cross-check, resolve coding conflicts, and consolidate the codes into a codebook across two sessions. Using the codebook, the two coders analyzed six randomly selected participant data each. The research team then met, discussed the analysis outcomes, and finalized themes over three sessions.

      sentence describing how analysis was performed on data collected by the authors of this paper

    7. Future work could explore more seamless ways of preserving context, such as allowing users to navigate through every sentence of an abstract directly within the Cross-Sentence Relationship pane, fostering a more cohesive understanding of the content.

      any sentence that describes explicit design implications

    8. In this sense, AbstractExplorer enables dialectical activities that users may otherwise have found to be too tedious or difficult to engage with.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    9. In this work, we introduce a new paradigm for exploring a large corpus of small documents by identifying roles at the phrasal and sentence levels, then slice on, reify, group, and/or align the text itself on those roles, with sentences left intact.

      any sentence that describes explicit design implications

    10. Our work demonstrates that designs informed by Structure-Mapping Theory can support users in navigating, making use of, and engaging with variation present in information. In this sense, AbstractExplorer enables dialectical activities that users may otherwise have found to be too tedious or difficult to engage with.

      any sentence that describes explicit design implications

    11. Like prior Structural Mapping Theory (SMT)-informed work in text corpora representation, AbstractExplorer's features have enabled some users to see more of both the overview and the details at the same time, facilitating abstraction without losing context.

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    12. Like prior Structural Mapping Theory (SMT)-informed work in text corpora representation, AbstractExplorer's features have enabled some users to see more of both the overview and the details at the same time, facilitating abstraction without losing context.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    13. We posit that our approach can generalize to other domains such as journalism, code synthesis, and social media analytics where visual alignment of text can enable meaningful comparisons of underlying patterns to identify relational clarity.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    14. In this work, we introduce a new paradigm for exploring a large corpus of small documents by identifying roles at the phrasal and sentence levels, then slice on, reify, group, and/or align the text itself on those roles, with sentences left intact.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    15. Dialectical activities cannot be done on a user's behalf by AI; with variation affordances, AI is supporting the user's engagement with the data themselves.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    16. We posit that our approach can generalize to other domains such as journalism, code synthesis, and social media analytics where visual alignment of text can enable meaningful comparisons of underlying patterns to identify relational clarity.

      any sentence that describes explicit design implications

    17. We demonstrate how slicing sentences according to roles and visually aligning them can help readers perceive cross-document relationships in a coherent manner.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    18. Our work demonstrates that designs informed by Structure-Mapping Theory can support users in navigating, making use of, and engaging with variation present in information.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    19. pre-computing and reifying cross-document analogous relationships make it psychologically possible for users to engage—if they are willing to be guided by it. (Lower NFC users are more likely to fall into this category.)

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    20. Activity log data, which revealed how participants actually used the interface, echoed the above findings. According to the log data, participants spent most of their reading time (66.31%) with vertical alignment on the second element in structure pairs, followed by alignment on the first element (29.19%), and left-justified alignment (5.13%). Highlighting usage showed a similar preference: 91.13% of time with all chunks highlighted, 8.25% with partial highlighting, and minimal time (0.63%) without highlights.

      sentence describing how analysis was performed on data collected by the authors of this paper

    21. In this section, we present findings on how AbstractExplorer supports comparative close reading at scale by integrating quantitative survey responses and log data with qualitative analysis of transcripts and open-ended responses. The qualitative analysis process is described in detail in Appendix H.

      sentence describing how analysis was performed on data collected by the authors of this paper

    22. Throughout the two tasks, we also collected detailed interaction logs including counts of user-defined aspects created, duration of highlighting usage, and time allocation across the three possible alignment options.

      sentence describing how analysis was performed on data collected by the authors of this paper

    23. Using a two-tailed Mann-Whitney U Test, we found that participants who reported their lowest perceived cognitive load when all three features were enabled had significantly lower NFC than participants who reported their lowest cognitive load level when skimming with no features enabled—in the baseline interface (p=0.03).

      sentence describing how analysis was performed on data collected by the authors of this paper

    24. Both gaze data and the semi-structured interviews revealed that lower NFC participants were more willing to be guided by the three features and took advantage of them consciously.

      sentence describing how analysis was performed on data collected by the authors of this paper

    25. For simplicity of analysis, we denote participants with NFC scores above the overall participants' median NFC of 5.42 (IQR = 0.583) as higher NFC, and lower NFC otherwise.

      sentence describing how analysis was performed on data collected by the authors of this paper

    26. The study concluded with a 15-minute semi-structured interview. During the interview, participants saw screenshots from the three conditions and were asked which they preferred and disliked, why, what they wished the interface had, what influenced their skimming, and how they normally skimmed texts.

      sentence describing any interview procedures

    27. Lower NFC participants were generally guided by emergent visual patterns created by the interactions between features, especially blocks of color spanning multiple sentences created when all three features are turned on.

      statements that draw general conclusions about humans, computers, and/or human-computer interaction based on the results of the specific experiment done in the paper.

    28. In this study, we allowed participants to experience views of same-aspect sentences (Section 4.1.1) with different combinations of highlighting, ordering, and alignment (as described in Section 4.1.2 and Section 4.1.4) enabled or not, in order to understand which and/or what combinations most effectively supported users' ability to skim and read laterally across documents.

      sentence relating to methodology

    29. We collected 80 sentences from our abstracts dataset labeled by our system as "Methodology/Contribution." Participants viewed the same 80 sentences in each condition—often with a different subset of sentences initially visible due to ordering changes—but only had two minutes to look at them in each condition.

      sentence describing how analysis was performed on data collected by the authors of this paper

    30. To contrast participants' gaze patterns in each condition, we used a Tobii Pro Spark eye-tracker placed below the desktop monitor used by all subjects; Tobii Pro Lab software recorded each participant's gaze over time in each condition.

      sentence describing how analysis was performed on data collected by the authors of this paper

    31. Structural mappings between objects are part of the cognitive process of comparison according to the Structure-Mapping Theory [17], and juxtaposition can facilitate humans in recognizing particular possible structural mappings between objects [75].

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    32. Inspired by GP-TSM [24], AbstractExplorer first segments sentences into grammar-preserving chunks—segments that respect grammatical boundaries, i.e., an LLM judges that the sentence can be truncated at that chunk boundary without breaking the grammatical integrity of the preceding text. Each chunk is then classified by an LLM as having one of nine pre-defined roles, each of which has its own assigned color.

      sentence relating to methodology

    33. We consider common sequences of chunk roles to be alignable structures that could be used to support users in identifying structural similarities and differences across sentences in different abstracts, in line with Structure-Mapping Theory [17].

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    34. This ordering prioritizes dominant structural patterns (largest groups first) while exposing fine-grained variations (via length-sorted triplets), mirroring how humans compare sentences, if SMT is an accurate description in this domain of comparative close reading.

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    35. In SMT terminology, rendering and arranging according to corresponding chunks reify "commonalities in structure," while variation within corresponding chunks are "alignable differences" that users are predicted to notice.

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    36. In the first part of the session, we asked participants about their strategies for selecting publication venues for their manuscript submissions, how they identify and synthesize information from venues, their approaches to writing manuscripts, and finally, the technology they have used to help with these processes, current technology shortcomings, and ideas for addressing these challenges.

      sentence describing any interview procedures

    37. In order to determine (1) the context in which we might offer novel views of scientific abstracts and (2) the intelligibility of various novel prototype designs for reifying cross-abstract relationships, we conducted a formative interview study with 12 active researchers (see Appendix A for participant information).

      sentence describing any interview procedures

    38. We used these mock-ups as design probes [31] to inspire ideation and elicit creative responses. Specifically, we asked participants to compare and contrast alternative mock-ups and reflect on how they could be used or improved to support their known or emerging synthesis and information-foraging goals.

      sentence describing any interview procedures

    39. The prior SMT-informed tools in Section 2.3 for both code and natural language corpora suggest that the cognitive process of comparing texts may be no exception to the cognitive processes SMT predicts.

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    40. The interview sessions were divided into two parts: an open-ended semi-structured interview about their backgrounds and practices, followed by feedback on a range of mock-ups, including novel reified relationships between analogous sentences in different abstracts (Figure 2).

      sentence describing any interview procedures

    41. Structural Mapping Theory (SMT) is a long-standing well-vetted theory from Cognitive Science that describes how humans attend to and try to compare objects by finding mental representations of them that can be structurally mapped to each other (analogies).

      sentence related to any theory

    42. These examples of text-centric lossless techniques do not abstract away or summarize; they strategically re-organize and re-render the existing text to help enhance readers' own perceptual cognition, informed by Structural Mapping Theory (SMT) [17].

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    43. The human perceptual, comparative mental machinery that SMT describes is part of what enables humans to form more abstract structured mental models from concrete examples, among other critical knowledge tasks.

      sentences that mention theory, explicitly or implicitly; one sentence at a time

    44. AbstractExplorer instantiates new minimally lossy2 SMT-informed techniques for skimming, reading, and reasoning about a corpus of similarly structured short documents: phrase-level role classification that drives sentence ordering, highlighting, and spatial alignment.

      sentence related to any theory

    1. Interesting take, plotting cultural attitudes of models alongside those of countries. The dimensions survival-self-expression and secular-traditional are a bit odd, apparently stemming from the World Values Survey. What would you get if you plot this stuff on the 6 dimensions of Hofstede [[Cultures and Organizations by Geert Hofstede]] 1980s, which shows a more nuanced picture that isn't as strongly geo-graphically oriented as these two axis.

    1. Best AI Note Takers — 2026
      • Overview of the AI Audio/Note-Taker Category

        • Tested a wide range of devices categorized as AI pins, note-takers, second brains, or lifelongers [00:00:06].
        • Distinct from previous failures like the Rabbit R1 or Humane AI Pin because they do not aim to replace smartphones [00:00:23].
        • Big tech is moving heavily into the space, with Amazon and Meta recently acquiring key startups in the audio sector [00:00:39].
        • Devices share the same core workflow: record audio, transfer it to a phone, transcribe it via a mobile app, and process it with an AI model for summaries and action items [00:01:14].
      • Two Key Product Dimensions

        • Trigger Recording vs. Always Listening: Devices either record only when manually prompted or constantly listen to collect continuous ambient context [00:01:41].
        • Summarization vs. Proactive Interpretation: Products focus strictly on summarizing audio or aim to actively interpret and guide the user [00:01:41].
      • Triggered Recording Tools (The Currently Practical Category)

        • Evaluation Criteria: Transcripts and speaker detection are very similar across brands because they outsource to the same providers [00:02:42]. Differentiation comes from:
          • File Transfer: Speed and reliability of moving audio files to a phone [00:03:02].
          • App Quality & Ecosystem: Stability of the mobile app and existence of a desktop app [00:03:13].
          • Trust & Longevity: Manufacturer data privacy practices and financial stability to avoid hardware bricking [00:03:19].
          • Data Lock-in: The ease of exporting notes to external personal knowledge management systems [00:03:30].
          • Cost Structure: Upfront hardware price combined with ongoing subscription costs for server-side AI processing [00:03:35].
        • Top Recommendation — Plaud: The clear winner for file transfer speed/reliability, background Bluetooth/Wi-Fi syncing, cloud upload capabilities, and a functional template ecosystem [00:03:58]. Offers a robust desktop companion app that records virtual meetings via system audio without utilizing an intrusive bot [00:04:59]. It maintains SOC 2 and HIPAA compliance for data privacy [00:05:42]. Data lock-in is mitigated via Zapier integrations, automatic email delivery, and a developer community [00:07:01].
        • Cost-Saving Alternative: The open-source "AudioBridge" project allows users to bypass expensive Plaud subscription costs by utilizing their own direct AI API keys [00:08:41].
        • Other Notable Contenders:
          • Soundcore: Excellent hardware with a built-in magnetic charging case, but restricted by a basic headphone app that lacks AI search or customization [00:08:58].
          • Pocket: Refined premium metal hardware and strong export features (MCP server), but held back by inconsistent phone transfer reliability and sudden changes to their subscription plans [00:09:37].
          • Haidoc P1 / P1 Mini: Connects directly via Bluetooth headphones to a computer or phone to save files locally on internal memory without requiring software installations—ideal for highly locked-down enterprise computers [00:10:36].
      • Always-Listening Devices (The "Second Brain" Category)

        • Focused on building an all-knowing memory backup with perfect recall, daily recaps, and automated task generation [00:11:29].
        • Tested Options:
          • Friend & Lookie: Non-recommended. Friend is invasive/sassy; Lookie includes a camera but looks like a conspicuous police body camera and performs poorly [00:11:41].
          • Limitless Pendant: Off the market following Meta's acquisition of the company [00:12:06].
          • OMI: Open-source and ambitious (working on screen recording and AI glasses), but currently buggy and lacks product focus [00:12:44].
          • B: Highly polished initially, but customer support and software stability degraded heavily after Amazon's acquisition [00:13:29].
          • Fieldly: The best of the group due to its focused approach, clean transcriptions, reliable hardware, multi-day battery life, and strong desktop app integration [00:14:13].
        • Fatal Flaws ("Context Rot"): Current models fail at accurate diarization (figuring out who said what), often attributing dialogue heard from nearby strangers or media to the user [00:15:05]. The user faces a heavy administrative burden to clean up flawed AI data, making the absolute "always-on second brain" promise currently non-viable [00:15:34].
      • Legal, Ethical, and Social Boundaries

        • Roughly 40% of the US population lives in two-party consent states, creating legal friction for recording private interactions [00:16:33].
        • Socially, requesting recording consent in private contexts remains awkward, frequently altering normal human behavior [00:16:48].
      • Future Market Trends

        • Form Factors: Sharp rise in ring-based options (Sandbar, Pebble, Fable) and a shift toward self-improvement pendants focused on emotion-tracking and self-awareness (Nerva, Nuna) [00:17:24].
        • Glasses & Visuals: Shift toward smart glasses (Meta Ray-Ban integrations, Pickle, Rokid) and pendant cameras [00:18:01].
        • Industry Heavyweights: Big tech is aggressively entering the market; OpenAI is working on an audio device, and Apple recently acquired QAI for $1.5B to decipher silent speech via jaw/facial micro-movements [00:18:29].
        • Mainstream Adoption Outlook: Unlike the failure of Google Glass, modern audio-only devices feature virtually invisible microphones, bypassing public visibility backlash [00:19:28]. Adoption may mirror a competitive sports dynamic (like the NBA three-point revolution): if the tools offer an undeniable cognitive or professional advantage, adoption will become mandatory to avoid falling behind [00:19:40].
    1. no single architecture dominates; rather, effectiveness depends on aligning the memory structure with the specific workload bottleneck

      对智能体记忆系统的批判性审视。当前业界没有一刀切的完美架构,记忆模块的设计必须与具体的任务瓶颈相匹配。这打破了“通用记忆系统”的幻想,提示我们在构建 Agent 时需要针对局部维护成本和任务特征进行定制化设计。

    2. It will be decided by who builds the best worlds for models to learn in, the best guardrails for them to operate within, and the best games to discover what they can actually do.

      作者在文末提出了极具洞察力的结论:AI 的竞争焦点已从单纯的模型规模,转移到了“环境构建”、“安全护栏”和“动态评测”三个维度。这意味着算力壁垒可能被数据和评估壁垒所取代,未来的 AI 巨头将是那些能打造最佳“沙盒生态”的公司。

    3. the way we communicate with them must evolve from loose conversation into something closer to structured collaboration.

      随着模型变得更加 agentic,传统的自然语言提示词工程可能正在走向终结。未来的人机交互将更像是在设计机器可读的工作流。这隐含了一个假设:为了可靠性和可控性,我们需要牺牲部分自然语言的模糊性,转向结构化的语义标记。

    4. we need arenas where models reveal themselves under pressure, with imperfect information, feedback loops, and consequences.

      反直觉的观点:传统的静态排行榜可能正在失效。在复杂环境中,模型的智能应该体现为可执行的策略而非单纯的文本回答。将 AI 评测转化为类似足球比赛的高压动态博弈,揭示了未来评测体系向“后果驱动”和“多智能体交互”演进的趋势。

    5. SK Hynix filed to raise up to 45.45 trillion won (~$29.4B) via a Nasdaq ADR listing

      近300亿美元的巨额募资,反映了 AI 算力基础设施对高带宽内存(HBM)的极端渴求。在投资者追捧 AI 存储芯片的背景下,这种规模的上市不仅是资金的角逐,更暗示着全球半导体供应链正在围绕 AI 算力需求进行深度的资本重构。

    6. A gameplay clip is not merely pixels. It is pixels plus choices.

      极其精辟地概括了具身智能下一步的数据瓶颈。语言模型用互联网文本训练,但缺乏对物理世界因果关系的理解。游戏视频包含了“感知-决策-反馈”的完整闭环,这种带有动作标签的数据可能成为下一代大模型突破通用性的关键预训练基座。

    7. Frontier AI releases are starting to look less like software updates and more like controlled deployment of critical infrastructure.

      这一金句精准地捕捉到了前沿 AI 模型发布范式的根本性转变。模型发布不再仅仅是技术迭代,而是涉及到政府协调层、安全架构和分阶段访问策略的社会化部署。这隐含着一个重要假设:AI 的风险等级已经达到了传统关键基础设施的级别。

    1. Rinderknecht asked ChatGPT whether someone could be blamed for a fire if it was lit by their cigarette.

      这句引用揭示了检方的核心论点:试图将被告与AI的对话记录作为其犯罪意图(犯罪故意)的证明。这是非共识的法律实践,将AI聊天记录等同于传统的日记或搜索记录,引发了关于AI对话能否作为思想犯罪证据的深刻争议。

    1. Anthropic unveils 'Claude Science' AI platform for scientific research

      这是文章的核心事实声明,指出Anthropic发布了专为科学研究设计的全新AI平台。然而,由于正文被付费墙屏蔽,该声明缺乏具体的技术细节、功能描述及适用领域等支撑信息,需要查阅一手新闻稿进行核查。

    1. With datasets like LOCUS we’re going to make the strange half-seen rules and laws that govern much of civic, local life be made accessible to AI systems, which may eventually allow them to better adapt themselves to hyperlocal purposes.

      这段话指出了LOCUS等数据集如何使AI系统能够更好地适应地方性目的,提出了AI在地方法律领域应用的潜力。

  2. Jun 2026
    1. Premise 1: Humans hand-edit content. Markdown was designed for people who write and revise their own text. That’s how blogs, docs, and READMEs still work. But agent output is different. You send a prompt. The agent generates a 2,000-word analysis, a code review, a project plan. You read it, maybe share it. You almost never open it in an editor and start rewriting paragraphs. The format’s core value proposition — easy to edit by hand — no longer matches the use case.
    2. Premise 3: Output is read-only. The old workflow was linear: prompt, generate, read, close. But the agent era is pushing toward something different. Users want to interact with the output: filter a table, adjust parameters, compare options side by side, export a subset, feed the result back into the next prompt. Markdown can’t carry interaction. It’s a one-way street.