2026
10 篇文章OpenAI 撤回 SWE-Bench Pro 建議:約三成題目是壞的
OpenAI 於 2026 年 7 月 8 日公開承認先前推薦的 SWE-Bench Pro 約三成題目有缺陷並撤回採用建議,文中整理四類失敗模式、三層審計流程,與重建編碼評測信任的具體主張。
閱讀文章 ↗Arena 年營收跑速破 1 億美元:AI 排行榜把群眾評測變成大生意
2026 年 6 月 29 日,TechCrunch 報導 AI 排行榜公司 Arena 年化營收跑速達 1 億美元:付費的 AI Evaluations 服務八個月內把年化營收從 3,000 萬推上 1 億,群眾評測正式成為一門大生意。
閱讀文章 ↗Claude 3.7 Sonnet 的「延長思考」:可調控的推理預算與可觀察的思考過程
Anthropic 在 Claude 3.7 Sonnet 加入可開關的 extended thinking mode,並讓開發者設定 thinking budget。本文整理官方公布的設計取捨、代理能力測試與安全評估,供產品開發者參考。
閱讀文章 ↗Ai2 拆解混合模型優勢:贏在內容詞,輸在複製
Ai2 用資料與訓練配方完全對齊的 Olmo 3 與 Olmo Hybrid 逐 token 比較損失,發現混合架構在內容詞上明顯佔優,但在逐字複製與右括號上優勢消失,顯示循環層擅長語意狀態追蹤、注意力擅長複製。
閱讀文章 ↗MosaicLeaks 基準:研究代理的對外查詢正在洩漏企業機密
ServiceNow 團隊發布 MosaicLeaks 基準:深度研究代理混合私有文件與網路搜尋時,攻擊者只看外流查詢紀錄就能拼出企業機密,而 PA-DR 訓練法把洩漏率從 34% 壓到 9.9%。
閱讀文章 ↗Writer 新研究:記憶工具讓模型更愛附和、更不準確
Writer AI Research 發表兩篇論文並推出 MIST 基準:Mem0、Zep 等記憶系統會放大模型附和行為,Sonnet 4.6 在 MIST-Moral 上從 1.6% 升到 40.2%。問題出在記憶擷取層,改用 LLM 生成摘要可壓到 12.8%。
閱讀文章 ↗微軟開源 ASSERT:把文字規格變成 AI 行為測試套件
2026 年 6 月 2 日,微軟開源 ASSERT 框架:開發者以自然語言描述代理應有行為,它自動生成測試情境、用 LLM 評審計分,輸出傷害與權衡兩類指標,並可掛進 CI 做回歸把關。
閱讀文章 ↗史丹佛法學院盲測:AI 答疑在 75% 對決中勝過教授
史丹佛法學院 6 月 1 日發布盲測研究:16 位法學教授對 40 題契約法答疑做了近 3,000 次匿名評比,AI 在 75% 對決中勝出,答案被標為有害的比例僅 3.5%,低於人類同儕的 12%。
閱讀文章 ↗NIST CAISI 評測 DeepSeek V4 Pro:距美國前沿約八個月
美國商務部 NIST 旗下的 CAISI 發布 DeepSeek V4 Pro 評測:IRT 估計 Elo 800,約當八個月前的 GPT-5,數學幾乎追平、抽象推理與資安最弱,七項基準中五項比 GPT-5.4 mini 便宜。
閱讀文章 ↗N-Day-Bench:用知識截止後的真實漏洞評測 LLM 安全能力
Winfunc 推出 N-Day-Bench:只收錄模型知識截止後才公開的真實漏洞,讓 LLM 在唯讀沙箱中從已知 sink 回溯資料流。首輪 GPT-5.4 以 83.93 居首,GLM-5.1 與 Claude Opus 4.6 緊追在四分之內。
閱讀文章 ↗
2026
11 ARTICLESUseful Intelligence per Dollar: A Scorecard for AI Investment Returns
OpenAI proposes a framework measuring AI ROI through useful work accomplished, cost per successful task, dependability, and scaling economics.
READ POST ↗OpenAI Retracts Its SWE-Bench Pro Endorsement: 30% Broken
OpenAI estimates ~30% of SWE-Bench Pro tasks are broken, retracts its adoption recommendation, and lays out four failure modes plus a model-assisted audit playbook.
READ POST ↗Arena's $100M Run Rate: Leaderboards as a Business
Arena, the crowdsourced AI leaderboard company, hit a $100M annualized run rate eight months after launching AI Evaluations, TechCrunch reported on June 29, 2026.
READ POST ↗Claude 3.7 Sonnet's Extended Thinking: A Practical Guide for Product Builders
Learn how Claude 3.7 Sonnet's extended thinking mode, thinking budgets, and visible thought process work, and what they mean for building AI products.
READ POST ↗Ai2 Maps Where Hybrid LLMs Beat Transformers, Token by Token
Ai2 compared Olmo 3 with Olmo Hybrid token by token, with matched data and training. Hybrids win on content words; the edge vanishes on verbatim copying and closing brackets.
READ POST ↗MosaicLeaks: Research Agents Leak Secrets Through Queries
ServiceNow's MosaicLeaks benchmark shows deep research agents leak enterprise secrets via outbound search queries; its PA-DR training cuts leakage from 34% to 9.9%.
READ POST ↗Memory Tools Make AI Models Agree: Writer's New Research
Writer's research shows memory systems like Mem0 and Zep amplify AI sycophancy: Sonnet 4.6 jumps from 1.6% to 40.2%. The culprit is extraction; prose summaries cut it to 12.8%.
READ POST ↗Microsoft ASSERT Turns Text Specs into AI Behavior Tests
Microsoft open-sourced ASSERT, a framework that turns plain-language behavior specs into generated test suites with LLM-judge scoring and CI regression gates for AI agents.
READ POST ↗Stanford Law Study: AI Tutor Answers Beat Professors 75%
A Stanford Law blind study found professors preferred AI contract-law tutoring answers in 75% of head-to-head matchups, and flagged AI answers as harmful less often than peers'.
READ POST ↗CAISI Puts DeepSeek V4 Pro 8 Months Behind US Frontier
NIST's CAISI scored DeepSeek V4 Pro at Elo 800 versus 1260 for GPT-5.5 — about eight months behind the US frontier, strongest in math, weakest in abstract reasoning and cyber.
READ POST ↗N-Day-Bench: LLMs vs Real Post-Cutoff Vulnerabilities
N-Day-Bench tests LLMs on real vulnerabilities disclosed after each model's knowledge cutoff. GPT-5.4 leads at 83.93, with GLM-5.1 and Claude Opus 4.6 within four points.
READ POST ↗