5 Matching Annotations
  1. Aug 2026
    1. 805k this year, 653k of those being 910C

      同一个量,两个来源差 2.2 倍。

      SemiAnalysis:2025 年 910C 为 653k。 Bloomberg(三周后):2025 年 910C 约 300k

      更值得注意的是本文在别处预先驳斥了更低的公开数字——「we believe the reported number of 200k Ascend chips to be significantly off the mark」。而 Bloomberg 的约 300k,离那个被驳斥的量级更近,离本文的 653k 更远。

      本文未披露该数字的来源与方法。

    1. No model we tested could complete it until it was given a compute budget of at least 30M tokens

      具体到可复算的一条。 AISI 靶场「The Last Ones」估计需人类专家约 20 小时;30M token 是模型能完成它的门槛。

      配合本文的幂律(拟合指数约 0.7–1.0):分钟级任务耗数千 token,小时级耗百万级,周级工作进入十亿量级。

    2. every model plateaued within its usual budget

      公允记账:主动交代削弱自身结论的负面结果。 HealthBench 上增加算力无效。同一篇的脚注 3 还写明:约 10–30% 的任务上,新模型表现不如前代。

      这类自曝在厂商发布里罕见。它也划出了本文结论的适用边界——增益集中在「智能体能自查自纠」的领域(代码、网安、数学),反馈弱或缺失的领域不适用。

    3. the fitted frontier trend is ~60% steeper when horizons are estimated at 50M tokens rather than 2.5M tokens per task

      本文最有后果的一句。 「前沿进展有多快」这个数字,部分取决于评测时给了多少预算——不是模型的固有属性。

      配套数字:同一前沿模型的 80% 时间跨度从 2.5M 预算下的约 40 分钟,升到 50M 下的约 4 小时;当前前沿从约 2 小时升到约 14 小时。

      对照:Anthropic 2026-07-29 的评测事故披露文全篇 23,027 字符,compute / token / budget / inference / runtime 0 次出现,却以「审阅 141,006 次评测运行」作分母。按本文论点,定预算下的分数是下界而非测量值。

    1. Flash is in high demand, our Cyber model is live, and Gemma models have surpassed 900M+ downloads

      选择性列举。 三项成绩全部避开旗舰 Gemini 3.5 Pro——该型号 2026-05-19 在 I/O 由 Pichai 亲自发布并承诺次月 GA(原话:“Give us until next month to get it to you”,台下有可闻的叹气),至本文发布日 2026-08-05 仍仅限 Vertex allowlist 预览,已延期逾两个月。Fortune 逐字:“months behind its original June launch target.”

      另注:Gemma 的「下载量」是分发指标而非使用指标,与 Gemini app 的月活不可比。