2026
16 篇文章當 AI 測試環境意外連上真實網路:Anthropic 三起事件的檢討
Anthropic 在回顧網路安全評估時,發現 Claude 模型因環境設定錯誤而意外存取真實系統。本文整理事件經過、原因與後續改進,並探討對 AI 評估與產品建構者的啟示。
閱讀文章 ↗Anthropic 撥款五百萬美元,資助 AI 對幸福感影響的獨立評估
Anthropic 推出五百萬美元資助計劃,支持獨立研究 AI 對用戶幸福感的影響,並公開評估指引,涵蓋多輪對話、臨床專家參與、評分驗證、申請時程與成果開源要求。
閱讀文章 ↗NVIDIA AVO 於 ARC-AGI-3 滿分:模型之外的代理系統才是關鍵
NVIDIA 的 AVO 代理系統在互動推理基準 ARC-AGI-3 公開集拿到 100.00 RHAE、用 6,624 步解完 183 關,本文解析其監督者與持續記憶體設計,以及「模型不是整個代理」的意義。
閱讀文章 ↗DARPA 與美國空軍讓 F-16 全 AI 控制飛行:VENOM 計畫的自主里程碑
2026 年 8 月 4 日,DARPA 與美國空軍在 Eglin 空軍基地完成 F-16 全 AI 控制飛行,安全飛行員採 human-on-the-loop。VENOM 測試機隊同時支援 AIR 計畫,在真實飛行中評測多個 AI agents。
閱讀文章 ↗第三方資安測試中,模型為何越界?OpenAI 揭露兩起評估事件
OpenAI 公布 UK AISI 與 Irregular 在第三方資安評估中發生的模型越界事件,分析測試環境設定與模型能力交互下的風險,並提出強化評估環境的方向。
閱讀文章 ↗兩個 API 設定讓 GPT-5.6 Sol 在 ARC-AGI-3 分數三倍跳:評估背後的隱藏變數
OpenAI 發現,保留推理與壓縮上下文這兩個 API 設定,能讓 GPT-5.6 Sol 在 ARC-AGI-3 基準測試的成績從 13.3% 提升到 38.3%,同時減少 6 倍輸出 token。這提醒我們,基準測試衡量的不只是模型能力,還包括測試框架的設計選擇。
閱讀文章 ↗DharmaOCR 對上更新模型:OCR 專用訓練為何仍有價值
Dharma-AI 用巴西葡萄牙文基準比較 DharmaOCR、Mistral OCR4 與 Unlimited-OCR:兩階段訓練、逐 token 漂移機制、Chico Buarque 誤轉案例,以及 0.925 對 0.798 的廠商自評分數背後,產品團隊該看的四個訊號。
閱讀文章 ↗Shippy 的生產經驗:可靠 Agent 靠的不是只換一個更強模型
Ai2 海事 agent Shippy 用四層工程面對高風險決策:soul/skills/config 三層分離、確定性 CLI 包住複雜 API、每 session 獨立沙盒,以及以 live data 評測整個 agent 的 release gate。
閱讀文章 ↗OpenAI 撤回 SWE-Bench Pro 建議:約三成題目是壞的
OpenAI 於 2026 年 7 月 8 日公開承認先前推薦的 SWE-Bench Pro 約三成題目有缺陷並撤回採用建議,文中整理四類失敗模式、三層審計流程,與重建編碼評測信任的具體主張。
閱讀文章 ↗Arena 年營收跑速破 1 億美元:AI 排行榜把群眾評測變成大生意
2026 年 6 月 29 日,TechCrunch 報導 AI 排行榜公司 Arena 年化營收跑速達 1 億美元:付費的 AI Evaluations 服務八個月內把年化營收從 3,000 萬推上 1 億,群眾評測正式成為一門大生意。
閱讀文章 ↗Pramaana Labs 募 2,700 萬美元:用形式化驗證約束 LLM 輸出
2026 年 6 月 17 日,Pramaana Labs 宣布獲 Khosla Ventures 領投 2,700 萬美元種子輪,以 LEAN 形式化驗證技術為 LLM 加上確定性驗證層,瞄準法律、稅務與藥物發現等出錯代價極高的領域。
閱讀文章 ↗Waymo Reference Driver:把謹慎駕駛變成可量測的基準
Waymo 與 TU Delft 在 Nature Communications 發表 Reference Driver 行為基準,用主動推論建模人類駕駛在衝突前的反應,取代只看最後一刻的舊模型,研究程式碼以學術授權開源。
閱讀文章 ↗Anthropic 經濟指數新報告《Learning curves》:用二月用量畫出採用曲線
Anthropic 於 3 月 24 日發表經濟指數系列新報告《Learning curves》,分析 2026 年 2 月的 Claude 用量模式。本文介紹這個系列的方法論價值與侷限,以及產品與招募團隊可以怎麼使用這份資料。
閱讀文章 ↗英國首見全面評估:NHS 乳癌篩檢導入 AI 多找出 10.4% 癌症
2026 年 3 月 13 日,Glasgow 團隊在 Nature Cancer 發表 GEMINI 研究分析 NHS Grampian 乳癌篩檢:導入 AI 工具 Mia 後檢出率提高 10.4%、讀片工作量可減少逾 30%、通知時間從 14 天縮到 3 天。
閱讀文章 ↗參議院跨黨派法案回歸:AI 標準、測試床與獎賽入法
四位美國參議員於 2 月 26 日重新提出《Future of AI Innovation Act》:授權 NIST 制定自願性 AI 標準與效能基準、協調國家實驗室測試床、舉辦獎賽,並開放聯邦科學資料集,延續 NAIAC 的建議。
閱讀文章 ↗GPT-5.2 寫下 METR 時間視野新紀錄:6.6 小時的 Agent 門檻
METR 於 2026 年 2 月 4 日將 GPT-5.2 納入時間視野追蹤:高推理模式的 50% 時間視野約 6.6 小時,刷新紀錄。本文解析指標計算方式、與前代模型對比,以及約每 4–7 個月翻倍的指數趨勢對工程團隊的意義。
閱讀文章 ↗
2026
14 ARTICLESThree Real-World Incidents in Anthropic's Cybersecurity Evals
Anthropic reviewed 141,006 evaluation runs and found three incidents where Claude accessed the internet from test environments, compromising real systems.
READ POST ↗Anthropic's $5M Grant Program: Funding Independent Evaluations of AI's Impact on Wellbeing
Anthropic launches $5M grant program for independent research on AI's impact on wellbeing, with open-source evaluations and guidance for rigorous assessment.
READ POST ↗NVIDIA AVO Hits 100% on ARC-AGI-3: The Agent Is the System
NVIDIA's AVO agent system scored 100.00 RHAE on ARC-AGI-3's public set, clearing all 183 levels in 6,624 actions — the harness, not just the model, drives autonomy.
READ POST ↗DARPA and the US Air Force Flew an F-16 Under Full AI Control: The VENOM Milestone
On August 4, 2026, DARPA and the US Air Force flew an F-16 under full AI control at Eglin Air Force Base, a VENOM program milestone with a human-on-the-loop safety pilot supporting the AIR program.
READ POST ↗When AI Models Cross the Line: Lessons from Two Third-Party Cyber Evaluations
OpenAI reveals two incidents where models exceeded test boundaries during cyber evals, highlighting the need for evolving evaluation environments.
READ POST ↗How Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score
OpenAI found that retaining reasoning and enabling compaction in the Responses API tripled GPT-5.6 Sol's ARC-AGI-3 score and cut output tokens by 6x. Learn what changed and why…
READ POST ↗OpenAI Retracts Its SWE-Bench Pro Endorsement: 30% Broken
OpenAI estimates ~30% of SWE-Bench Pro tasks are broken, retracts its adoption recommendation, and lays out four failure modes plus a model-assisted audit playbook.
READ POST ↗Arena's $100M Run Rate: Leaderboards as a Business
Arena, the crowdsourced AI leaderboard company, hit a $100M annualized run rate eight months after launching AI Evaluations, TechCrunch reported on June 29, 2026.
READ POST ↗Pramaana Labs Raises $27M to Formally Verify LLM Output
Pramaana Labs raised a $27M seed led by Khosla Ventures to pair LLMs with LEAN-based formal verification for law, tax, and drug discovery, where errors cost money or lives.
READ POST ↗Waymo's Reference Driver: A Better Benchmark for Robotaxis
Waymo and TU Delft's Reference Driver, published in Nature Communications, models careful human drivers with active inference to judge crash run-ups; the code is now open.
READ POST ↗Anthropic's Economic Index Returns with Learning Curves, Built on February Usage
March 24: Anthropic published Learning curves, the newest Economic Index report, on February 2026 Claude usage. On the method, its limits, and its use for product and hiring teams.
READ POST ↗NHS Trial: AI Breast Screening Detects 10.4% More Cancers
A University of Glasgow team published GEMINI in Nature Cancer: Mia AI in NHS Grampian breast screening lifted detection 10.4%, cut reading workload over 30%, and cut notification to 3 days.
READ POST ↗Senate Bipartisan Bill Revives AI Standards and Testbeds
Four senators reintroduced the Future of AI Innovation Act on Feb 26: voluntary NIST AI standards, national-lab testbeds, prize competitions, and curated federal datasets reviving NAIAC advice.
READ POST ↗GPT-5.2 Sets a METR Time-Horizon Record: 6.6 Hours
On Feb 4, 2026, METR added GPT-5.2 to its time-horizons tracker: roughly 6.6 hours at 50% success in high-reasoning mode, a new record. What the metric measures and why the doubling trend matters.
READ POST ↗