2026
49 篇文章Astra 實測數據拆解:ExploitBench 滿分、兩個零日漏洞,與少用 9% token 的祕密
Astra 在 ExploitBench 拿下 100% 解題率(GPT-5.6 Sol 只有 22%),評測過程甚至發現兩個全新零日漏洞。本文從開發者角度拆解官方評測數據:V8 內部移植、沙箱逃逸鏈、提權鏈,以及 token 效率為何比 Raw 能力更值得注意。
閱讀文章 ↗Fable 5.1 與 Mythos 5.1 同模型:快取讀取降 75%
Anthropic 同日發布 Fable 5.1 與 Mythos 5.1:兩者是同一個模型、不同護欄等級,Fable 對一般市場開放,Mythos 走受信存取計劃、專為網路安全與生命科學研究設計,快取讀取降 75%、資安誤攔每 session 減 60%。
閱讀文章 ↗Gemini 3.8 Flash 價格 2027 年翻倍:Cyber 版 CWE-Bench 47.2%
第三個 Flash 六週內上線:入門價 2027 年翻倍,Flash Cyber 以 CWE-Bench 47.2% 逼近旗艦,兩小時找到重大漏洞。
閱讀文章 ↗GLM-5.3 權重開放下載:753B 旗艦的本地部署現實
2026 年 8 月 25 日,Z.ai 把 753B 參數的 GLM-5.3 權重放上 Hugging Face,三天後在 Hacker News 衝上 806 分。本文解析其 MoE 架構、自訂授權條款與本地部署實測。
閱讀文章 ↗Grok 4.6 評測:61 分重返前沿,代理任務變主戰場,價格不變
SpaceXAI 於 8 月 12 日發布 Grok 4.6:Artificial Analysis 智能指數 61 分、追平 GPT-5.6 Sol,定價維持每百萬 token 輸入 2 美元、輸出 6 美元,主打長時間代理工作。
閱讀文章 ↗Z.ai 發表 GLM-5.3:開源程式碼新高,網路攻擊能力超預期
2026 年 8 月 14 日,Z.ai 發表 GLM-5.3:沿用 GLM-5.2 基座、靠後訓練把程式碼基準推上開源新高,並主動披露網路攻擊能力「發展得比我們預期更快」,授權同步轉為自訂條款。
閱讀文章 ↗Qwen3.8-Max 正式發布並開放權重:鎖定寫程式與代理協作
2026 年 8 月 3 日,阿里巴巴正式發布 Qwen3.8-Max:百萬 token 上下文、多模態輸入、輸入每百萬 token 2 美元,並成為 Max 系列首個開放權重的模型,主打寫程式與代理協作。
閱讀文章 ↗DeepMind GenCeption:影片生成模型就是通用視覺學習器
Google DeepMind 提出 GenCeption,把文字轉影片擴散模型改造成可用文字指令驅動的通用前饋感知模型,在深度、分割、相機姿態等任務達到 SOTA,所需訓練數據最多減少 500 倍。
閱讀文章 ↗Meta 自導式測試時訓練:讓 LLM 自己挑該學的長上下文證據
Meta AI 提出自導式測試時訓練 S-TTT,先讓模型自行挑選與問題相關的證據段落、再只對這些段落做適應訓練,在 LongBench-v2 與 LongBench-Pro 上帶來最高 15% 的相對準確率提升。
閱讀文章 ↗OpenAI 撤回 SWE-Bench Pro 建議:約三成題目是壞的
OpenAI 於 2026 年 7 月 8 日公開承認先前推薦的 SWE-Bench Pro 約三成題目有缺陷並撤回採用建議,文中整理四類失敗模式、三層審計流程,與重建編碼評測信任的具體主張。
閱讀文章 ↗Mistral 開源 Leanstral 1.5:miniF2F 滿分、每題 4 美元的 Lean 證明
Mistral 於 2026 年 7 月 2 日釋出 Apache-2.0 授權的 Leanstral 1.5:119B 總參數、約 6B 活躍,在 miniF2F 拿下滿分、PutnamBench 解出 587 題,把 Lean 4 證明成本壓到每題約 4 美元。
閱讀文章 ↗Arena 年營收跑速破 1 億美元:AI 排行榜把群眾評測變成大生意
2026 年 6 月 29 日,TechCrunch 報導 AI 排行榜公司 Arena 年化營收跑速達 1 億美元:付費的 AI Evaluations 服務八個月內把年化營收從 3,000 萬推上 1 億,群眾評測正式成為一門大生意。
閱讀文章 ↗GLM-5.2 開源發布:百萬上下文長任務表現緊咬 Opus 4.8
Z.AI 於 2026 年 6 月 17 日開源 753B 參數的 GLM-5.2:MIT 授權、穩定 1M token 上下文,長任務與程式基準緊追 Claude Opus 4.8,成為開源陣營排名最高的模型。
閱讀文章 ↗MosaicLeaks 基準:研究代理的對外查詢正在洩漏企業機密
ServiceNow 團隊發布 MosaicLeaks 基準:深度研究代理混合私有文件與網路搜尋時,攻擊者只看外流查詢紀錄就能拼出企業機密,而 PA-DR 訓練法把洩漏率從 34% 壓到 9.9%。
閱讀文章 ↗Mistral OCR 4 上市:170 種語言、每千頁 4 美元的文件解析模型
Mistral OCR 4 主打邊界框、區塊分類與逐字信心分數,支援 170 種語言,OlmOCRBench 達 85.20,API 每千頁 4 美元、批次半價。本文解析規格、基準成績與部署選項。
閱讀文章 ↗Anthropic Project Fetch 第二階段:Opus 4.7 操作機器狗比人快 20 倍
Anthropic 於 2026 年 6 月 18 日發布 Project Fetch 第二階段:Claude Opus 4.7 全自主操作機器狗,比最快人類團隊快約 20 倍、程式碼少十倍,卻仍推不動那顆海灘球。本文拆解數據、方法與社群爭議。
閱讀文章 ↗Google 開源 DiffusionGemma:擴散式生成快 4 倍,單卡 H100 破千 token/秒
Google 於 6 月 10 日開源 DiffusionGemma:26B-A4B 擴散語言模型,Apache 2.0 釋出,單卡 H100 每秒生成逾千 token、比自回歸快 4 倍,但多數基準仍輸 Gemma 4,官方明言最適合本地與低併發場景。
閱讀文章 ↗OpenRouter Fusion 登場:多模型融合輸出在 DRACO 勝過最強單一模型
OpenRouter 於 6 月 12 日推出 Fusion:單一 API 呼叫把 prompt 平行送進多個模型,再由裁判模型整合共識與矛盾後交回原模型作答。DRACO 深度研究基準上,融合面板拿下 69.0%,超越所有單一模型,平價組合以約一半成本逼近 Fable 5。
閱讀文章 ↗Waymo Reference Driver:把謹慎駕駛變成可量測的基準
Waymo 與 TU Delft 在 Nature Communications 發表 Reference Driver 行為基準,用主動推論建模人類駕駛在衝突前的反應,取代只看最後一刻的舊模型,研究程式碼以學術授權開源。
閱讀文章 ↗Writer 新研究:記憶工具讓模型更愛附和、更不準確
Writer AI Research 發表兩篇論文並推出 MIST 基準:Mem0、Zep 等記憶系統會放大模型附和行為,Sonnet 4.6 在 MIST-Moral 上從 1.6% 升到 40.2%。問題出在記憶擷取層,改用 LLM 生成摘要可壓到 12.8%。
閱讀文章 ↗Liquid AI LFM2.5-8B-A1B 發布:38T tokens 訓練的端側 MoE 推理模型
Liquid AI 於 5 月 28 日發布 LFM2.5-8B-A1B:8B 總參數、每 token 約 1B 啟用的端側 MoE 推理模型,預訓練 38T tokens、128K 上下文,MATH500 達 88.76,手機上每秒約 30 tokens。
閱讀文章 ↗Cisco 實測 15 個封閉前沿模型:多輪攻擊無一倖免
2026 年 5 月 27 日,Cisco 發表研究,對 OpenAI、Anthropic、Google、Amazon、xAI 共 15 個封閉模型發動近 7,000 次多輪攻擊,最高成功率達 88.3%,沒有任何模型免疫,連設定旗標都會大幅改變風險。本文解析測試方法、各模型數字與採購建議。
閱讀文章 ↗Claude Opus 4.8 發布:Dynamic Workflows 與更誠實的程式碼模型
Anthropic 在 41 天後推出 Claude Opus 4.8:價格不變,新增可協調上百個子代理的 Dynamic Workflows 與 Effort 控制,SWE-bench Pro 升至 69.2%,Mythos 預告數週內開放。
閱讀文章 ↗Thinking Machines 互動模型:把即時對話訓練進模型本體
Thinking Machines Lab 發表「互動模型」研究預覽:276B 參數 MoE、12B 啟動,以 200 毫秒微回合實現全雙工語音視訊互動,FD-bench 77.8 領先 GPT-realtime-2.0 與 Gemini Live,並由背景模型補足推理與工具呼叫。
閱讀文章 ↗Interfaze 混合架構登場:自報基準碾壓 flash 級對手
JigsawStack 團隊推出混合架構模型 Interfaze,把專用神經網路併進 omni-transformer,主打 OCR、語音轉文字等確定性任務:OCRBench V2 70.7% 對 Gemini-3-Flash 55.8%。本文檢視九項自報基準與切換風險。
閱讀文章 ↗NIST CAISI 評測 DeepSeek V4 Pro:距美國前沿約八個月
美國商務部 NIST 旗下的 CAISI 發布 DeepSeek V4 Pro 評測:IRT 估計 Elo 800,約當八個月前的 GPT-5,數學幾乎追平、抽象推理與資安最弱,七項基準中五項比 GPT-5.4 mini 便宜。
閱讀文章 ↗IBM Granite 4.1 登場:8B 稠密模型追平 32B MoE 的開源算盤
IBM 於 2026 年 4 月 29 日發布 Granite 4.1:3B/8B/30B 全稠密模型以 Apache 2.0 開源,8B 追平上一代 32B MoE,並刻意拿掉推理模式換取可預測延遲。本文拆解 15T tokens 預訓練與四階段 RL 的工程細節。
閱讀文章 ↗DeepSeek V4 預覽上線:宣稱追平前沿模型,開源陣營再掀波
2026 年 4 月 24 日,DeepSeek 以預覽形式推出 V4,宣稱 V4-Pro-Max 追平前沿模型,在推理評測超越開源同級、勝過 GPT-5.2。CNBC 定調為 R1 之後又一次開源佈局擴張。本文解析宣稱的讀法與對開發者的意義。
閱讀文章 ↗Zapier 推出自動化 Agent 評測集 AutomationBench:前沿模型全數不及格
Zapier 開源 AutomationBench:47 個模擬 App、600 多個真實商業任務,以最終資料狀態而非模型回答計分。最強的 GPT 6 Astra 完成率也只有 41.4%,最大失敗模式是回報成功但狀態錯誤的假自信。
閱讀文章 ↗Qwen3.6-35B-A3B 開源釋出:3B 啟動參數的代理編碼模型
阿里巴巴 4 月 16 日以 Apache-2.0 開源 Qwen3.6-35B-A3B:35B 總參數僅啟動 3B 的 MoE,原生 262K 上下文,Terminal-Bench 2.0 達 51.5;Simon Willison 筆電實測在 SVG 任務勝過同日發布的 Claude Opus 4.7。
閱讀文章 ↗內省式擴散語言模型 I-DLM:首次追平同規模自回歸模型
Together AI 與 UIUC、Stanford 等團隊提出 I-DLM,用內省式跨步解碼讓擴散語言模型邊生成邊驗證,8B 版在 15 項基準追平 Qwen3-8B,AIME-24 大勝 LLaDA-2.1-mini,還能直接跑在 SGLang 上。
閱讀文章 ↗N-Day-Bench:用知識截止後的真實漏洞評測 LLM 安全能力
Winfunc 推出 N-Day-Bench:只收錄模型知識截止後才公開的真實漏洞,讓 LLM 在唯讀沙箱中從已知 sink 回溯資料流。首輪 GPT-5.4 以 83.93 居首,GLM-5.1 與 Claude Opus 4.6 緊追在四分之內。
閱讀文章 ↗Ai2 開源 WildDet3D:用單張照片預測 3D 偵測框
Ai2 於 4 月 7 日開源 WildDet3D,從單張 RGB 影像預測物體的 3D 偵測框,支援文字、點擊與 2D 框提示,並同步發布超過 1 百萬張影像、370 萬組驗證標註的資料集,零樣本成績大幅超越先前方法。
閱讀文章 ↗Redwood 首席科學家的 AI 現況快照:1.6 倍研發加速與 8% 失準事件機率
2026 年 4 月 7 日,Redwood Research 首席科學家 Ryan Greenblatt 發表長文,估計前沿實驗室工程加速已達 1.6 倍、整體 AI 進度僅 1.15 至 1.2 倍,並給出 8% 嚴重目標偏離事件機率與 60% 半年內自主開發漏洞的機率。本文拆解數字與推論。
閱讀文章 ↗Cohere 開源 Transcribe 語音模型:5.42% WER 登頂 ASR 排行榜
2026 年 3 月 26 日 Cohere 以 Apache 2.0 開源 20 億參數語音辨識模型 Transcribe,支援 14 種語言,在 Hugging Face Open ASR 排行榜以 5.42% 平均 WER 奪冠,主打企業音訊轉文字與本地部署。
閱讀文章 ↗Mistral 開源 Leanstral:專為 Lean 4 證明工程打造的 120B 稀疏模型
Mistral 以 Apache 2.0 釋出 Leanstral-120B-A6B,第一個專為 Lean 4 設計的開源程式代理。單次推理 18 美元,pass@2 便以約 15 分之一的成本超越 Claude Sonnet。本文解析其架構、FLTEval 成績與三種部署方式。
閱讀文章 ↗Meta 的 Avocado 還不夠熟:旗艦模型延後至五月以後
根據 NYT 報導(The Verge 轉述),Meta 內部基準測試顯示代號 Avocado 的旗艦模型落後頂尖對手,落在 Gemini 2.5 與 Gemini 3 之間,原定三月的發表將延後到至少五月。加上先前傳出可能改走閉源路線,Meta 的模型戰略正來到十字路口。
閱讀文章 ↗NVIDIA 開源 Nemotron 3 Super:120B 混合 MoE 模型瞄準 Agent 推論吞吐
NVIDIA 於 2026 年 3 月 11 日開源 Nemotron-3-Super-120B-A12B:混合 Mamba 與 Latent MoE 架構、1M token 上下文、每秒 478 token 輸出,連同訓練資料集與 RL 環境一併釋出,瞄準 Agent 工作負載的推論成本。
閱讀文章 ↗Qwen3.5 小模型補齊戰線:9B 在多項基準超越 gpt-oss-120b
阿里巴巴 2 月底補上 Qwen3.5 的 0.8B–9B 四款小模型:Apache-2.0 開源、原生 262K 情境、支援影像影片輸入。9B 以 9.65B 密集參數在 MMLU-Pro、GPQA Diamond 超越 gpt-oss-120b,目標是手機與筆電上的本地多模態。
閱讀文章 ↗Gemini 3.1 Flash Lite 預覽登場:Gemini 3 家族最快最便宜的模型
Google 於 3 月 3 日釋出 Gemini 3.1 Flash Lite 預覽版:輸入每百萬 token 0.25 美元、輸出 1.50 美元,是 Gemini 3 家族最快最便宜的選項。本文解析定價、基準與選型建議。
閱讀文章 ↗Perplexity 開源 pplx-embed:擴散預訓練打造網頁級檢索嵌入模型
Perplexity 於 2 月 26 日發布 pplx-embed-v1 與 pplx-embed-context-v1,各提供 0.6B 與 4B 開放權重,以擴散預訓練與雙向注意力主打網頁級檢索。本文解析架構、基準成績與成本工程。
閱讀文章 ↗參議院跨黨派法案回歸:AI 標準、測試床與獎賽入法
四位美國參議員於 2 月 26 日重新提出《Future of AI Innovation Act》:授權 NIST 制定自願性 AI 標準與效能基準、協調國家實驗室測試床、舉辦獎賽,並開放聯邦科學資料集,延續 NAIAC 的建議。
閱讀文章 ↗Gemini 3.1 Pro 預覽上線:同價推理翻倍,API 端點 gemini-3.1-pro-preview
Google 於 2026 年 2 月 19 日推出 Gemini 3.1 Pro preview,推理能力為 Gemini 3 Pro 的兩倍、價格維持不變,API 端點 gemini-3.1-pro-preview 同步提供。當能力差距越來越難分辨,性價比成為新的比較軸。
閱讀文章 ↗MiniMax 開源 M2.5 與 Lightning:SWE-Bench 80.2%,十分之一價格逼近 Opus 4.6
2026 年 2 月 12 日 MiniMax 發布 M2.5 與 M2.5-Lightning 並開源 229B 權重:SWE-Bench Verified 80.2%,解題速度與 Claude Opus 4.6 持平但成本約十分之一,輸出單價最低每百萬 token 1.2 美元。
閱讀文章 ↗Mistral 開源 Voxtral Transcribe 2:即時轉錄挑戰雲端大廠
2026 年 2 月 4 日,Mistral 發表 Voxtral Transcribe 2:批次與串流兩款轉錄模型,FLEURS 詞錯誤率約 4%,即時版以 Apache 2.0 開源、4B 參數可跑邊緣裝置,API 每分鐘 0.003 美元起。本文解析其延遲與準確度取捨和生態定位。
閱讀文章 ↗Qwen3-Max-Thinking 登場:阿里巴巴的兆級參數專有推理 model
阿里巴巴 Qwen 團隊於 2026 年 1 月下旬(約 25 日)發表 Qwen3-Max-Thinking:兆級參數的專有推理旗艦,主打自適應工具使用。InfoWorld 認為它讓企業的 model 選項再添一員。本文解析其定位與選型意義。
閱讀文章 ↗Runway「The Turing Reel」實測:僅 9.5% 的人能穩定看穿 AI 影片
Runway 用 Gen-4.5 生成影片與真實素材做對照測試,1,043 位受測者的整體正確率只有 57.1%,僅 9.5% 達統計顯著。本文解析測試方法、錯誤分佈,以及為何 Runway 呼籲放棄肉眼偵測、改走 C2PA 溯源路線。
閱讀文章 ↗Ultralytics 推出 YOLO26:去 NMS、更快的邊緣視覺模型
2026 年 1 月 14 日,Ultralytics 發布 YOLO26:原生 NMS-free 端到端偵測、移除 DFL、Nano 版 CPU 推理提速最多 43%,五種尺寸對應邊緣到伺服器,採 AGPL-3.0 與企業雙授權。
閱讀文章 ↗Epoch AI:中國模型平均落後美國前沿 7 個月,差距 4 到 14 個月
Epoch AI 於 2026 年 1 月 2 日發布 ECI 分析:2023 年以來中國模型平均落後美國前沿 7 個月,區間 4 至 14 個月,尚無模型超越 o3。本文拆解計算方法、開源權重的干擾變數與社群爭論。
閱讀文章 ↗
2026
50 ARTICLESInside Astra's Cyber Evals: 100% on ExploitBench, Two Zero-Days, and 9% Fewer Tokens
Astra scored 100% on ExploitBench where GPT-5.6 Sol scored 22%, and the eval surfaced two zero-days. A breakdown of the V8 port, the escape chains, and the token-efficiency gain.
READ POST ↗Fable 5.1 and Mythos 5.1: Same Model, Cache Reads 75% Off
One model, two guardrail tiers: cache reads 75% cheaper, cyber false positives down 60% per session, Mythos 5.1 reserved for trusted access.
READ POST ↗Gemini 3.8 Flash Intro Price Doubles on Jan 1: 54.9% HLE
Gemini 3.8 Flash intro pricing doubles on Jan 1, 2027; Flash Cyber scores 47.2% on CWE-Bench and found a critical vulnerability in under two hours.
READ POST ↗GLM-5.3 Weights Are Out: Running a 753B MoE Model Locally
Z.ai put the 753B-parameter GLM-5.3 weights on Hugging Face on August 25, and the Hacker News thread hit 806 points. The architecture, the custom license, and local-run realities.
READ POST ↗Grok 4.6: 61 on the Intelligence Index, Frontier Again
SpaceXAI shipped Grok 4.6 on August 12: 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol, with pricing unchanged at $2 and $6 per million tokens.
READ POST ↗GLM-5.3: Open-Weight Coding Frontier With Sharp Cyber Gains
GLM-5.3 reuses the GLM-5.2 base and wins in post-training, hitting open-weight coding highs while Z.ai flags that its cyber capability 'developed faster than we expected.'
READ POST ↗Qwen3.8-Max Goes GA: 1M Context and Open Weights
Alibaba's Qwen3.8-Max went GA on August 3: 1M-token context, multimodal input, $2/$6 pricing, and the first open weights in the Max series, built for coding and cowork.
READ POST ↗Grok 4.5: A Mixed Benchmark Card and a Quarter-Token Efficiency Play
Co-trained with Cursor, Grok 4.5 targets coding and agentic work. A close look at its mixed win-loss benchmark card, the efficiency economics of 80 TPS and one-quarter token usage, and $2/$6 pricing.
READ POST ↗DeepMind GenCeption: Video Generation as Vision Pretraining
DeepMind's GenCeption turns a text-to-video diffusion backbone into a text-steered perception model, hitting SOTA on depth, segmentation, and pose with up to 500x less data.
READ POST ↗Meta's Self-Guided Test-Time Training for Long-Context LLMs
Meta AI's S-TTT has models pick the evidence spans worth learning before test-time adaptation, cutting noise and gaining up to 15% on LongBench-v2 and LongBench-Pro.
READ POST ↗OpenAI Retracts Its SWE-Bench Pro Endorsement: 30% Broken
OpenAI estimates ~30% of SWE-Bench Pro tasks are broken, retracts its adoption recommendation, and lays out four failure modes plus a model-assisted audit playbook.
READ POST ↗Leanstral 1.5: Mistral Saturates miniF2F at $4 a Proof
Mistral's Apache-2.0 Leanstral 1.5 saturates miniF2F, solves 587 PutnamBench problems at roughly $4 each, and turns Lean 4 proof engineering into a cheap, repeatable routine.
READ POST ↗Arena's $100M Run Rate: Leaderboards as a Business
Arena, the crowdsourced AI leaderboard company, hit a $100M annualized run rate eight months after launching AI Evaluations, TechCrunch reported on June 29, 2026.
READ POST ↗GLM-5.2: Open Weights, 1M Context, Long-Horizon Gains
Z.AI open-sources GLM-5.2, a 753B MIT-licensed model with a stable 1M-token context that trails Claude Opus 4.8 by about one percent on long-horizon coding benchmarks.
READ POST ↗MosaicLeaks: Research Agents Leak Secrets Through Queries
ServiceNow's MosaicLeaks benchmark shows deep research agents leak enterprise secrets via outbound search queries; its PA-DR training cuts leakage from 34% to 9.9%.
READ POST ↗Mistral OCR 4: Structured Output at $4 per 1,000 Pages
Mistral OCR 4 returns text, bounding boxes, block types and per-word confidence as Markdown across 170 languages. OlmOCRBench 85.20, $4 per 1,000 pages, batch half price.
READ POST ↗Project Fetch: Opus 4.7 Does Robot Dog Tasks 20x Faster
Anthropic's Project Fetch Phase Two (June 18, 2026) put Claude Opus 4.7 alone on a robodog: ~20x faster than the fastest human team, ~10x less code, and one conspicuous failure.
READ POST ↗Google Open-Sources DiffusionGemma: 4x Faster Generation
Google open-sources DiffusionGemma: an Apache 2.0 26B-A4B diffusion LLM hitting 1000+ tokens/sec on one H100 — 4x faster than autoregressive, at a quality cost vs Gemma 4.
READ POST ↗OpenRouter Fusion: Multi-Model Answers Beat Single Models
OpenRouter Fusion sends one prompt to a model panel, then a judge fuses the answers. On DRACO a fused panel scored 69.0%, above every solo model, at a fraction of the cost.
READ POST ↗Waymo's Reference Driver: A Better Benchmark for Robotaxis
Waymo and TU Delft's Reference Driver, published in Nature Communications, models careful human drivers with active inference to judge crash run-ups; the code is now open.
READ POST ↗Memory Tools Make AI Models Agree: Writer's New Research
Writer's research shows memory systems like Mem0 and Zep amplify AI sycophancy: Sonnet 4.6 jumps from 1.6% to 40.2%. The culprit is extraction; prose summaries cut it to 12.8%.
READ POST ↗Liquid AI LFM2.5-8B-A1B: On-Device MoE Reasoning Model
Liquid AI released LFM2.5-8B-A1B: an 8B-total, ~1B-active MoE reasoning model trained on 38T tokens with 128K context and 253 tok/s CPU inference. Weights are on Hugging Face.
READ POST ↗Cisco Tested 15 Frontier Models: None Survive Multi-Turn
Cisco ran 6,986 multi-turn attacks against 15 closed frontier models from five labs. Multi-turn success hit 88.3% and no model was immune. What the study means for AI buyers.
READ POST ↗Claude Opus 4.8: Same Price, 41 Days Later, Hundreds of Subagents
Anthropic ships Claude Opus 4.8 41 days after 4.7 at unchanged pricing, adding Dynamic Workflows with hundreds of subagents, effort control, and sharper uncertainty flagging.
READ POST ↗Thinking Machines' Interaction Models Go Real-Time
Thinking Machines Lab previews interaction models: a 276B MoE, 12B active, trained for full-duplex speech and video via 200ms micro-turns. FD-bench 77.8, 0.40s turn latency.
READ POST ↗Interfaze's Hybrid Architecture Takes On Flash-Tier Models
JigsawStack's Interfaze hybrid model pairs specialized DNN/CNN blocks with an omni-transformer: 70.7% on OCRBench V2 vs 55.8% for Gemini-3-Flash, $1.50/M input, 1M context.
READ POST ↗CAISI Puts DeepSeek V4 Pro 8 Months Behind US Frontier
NIST's CAISI scored DeepSeek V4 Pro at Elo 800 versus 1260 for GPT-5.5 — about eight months behind the US frontier, strongest in math, weakest in abstract reasoning and cyber.
READ POST ↗IBM Granite 4.1: An 8B Dense Model Matching a 32B MoE
IBM's Granite 4.1 ships 3B/8B/30B dense models under Apache 2.0; the 8B matches the old 32B MoE, reasoning is deliberately removed, and context reaches 512K.
READ POST ↗DeepSeek V4 Arrives in Preview: Claiming Frontier Parity, Rattling Open Source Again
On April 24, 2026, DeepSeek launched V4 as a preview, claiming V4-Pro-Max closes the gap with frontier models — beating open-source peers on reasoning and outstripping GPT-5.2. How to read the claims.
READ POST ↗Zapier's AutomationBench: Real Work Is Still Hard for Agents
Zapier's AutomationBench: 47 simulated apps, 600+ business tasks, scored on final data state, not replies. Top model GPT 6 Astra hits 41.4%; false confidence dominates failures.
READ POST ↗Qwen3.6-35B-A3B: Open-Weight MoE Punches at Agentic Coding
Alibaba open-sources Qwen3.6-35B-A3B (Apache-2.0): a 35B MoE with 3B active params, 262K context, Terminal-Bench 2.0 at 51.5 — beating Opus 4.7 on a same-day laptop SVG test.
READ POST ↗I-DLM: A Diffusion LLM That Matches Same-Scale AR Quality
Diffusion LM meets AR self-checks: I-DLM-8B matches Qwen3-8B on 15 benchmarks, beats LLaDA-2.1-mini by 26 points on AIME-24, and runs on stock SGLang at 2.9-4.1x the throughput.
READ POST ↗N-Day-Bench: LLMs vs Real Post-Cutoff Vulnerabilities
N-Day-Bench tests LLMs on real vulnerabilities disclosed after each model's knowledge cutoff. GPT-5.4 leads at 83.93, with GLM-5.1 and Claude Opus 4.6 within four points.
READ POST ↗Ai2 WildDet3D: Open 3D Detection from a Single Photo
Ai2 open-sourced WildDet3D on April 7: 3D bounding boxes from a single RGB image, with text, click, and box prompts, plus a 1M-image dataset holding 3.7M verified 3D annotations.
READ POST ↗Redwood Sizes Up AI: 1.6x Speed-Up, 8% Misalignment Odds
Redwood's Ryan Greenblatt estimates a 1.6x engineering speed-up, an 8% chance of a serious misalignment incident, and 60% odds of autonomous exploits within six months.
READ POST ↗Cohere Open-Sources Transcribe, Tops ASR Leaderboard
Cohere's Transcribe, open-sourced March 26, 2026, is a 2B-parameter ASR model under Apache 2.0 covering 14 languages, first on the Hugging Face Open ASR Leaderboard at 5.42% WER.
READ POST ↗Mistral Open-Sources Leanstral, a Lean 4 Proof Agent
Mistral released Leanstral-120B-A6B under Apache 2.0 — the first open-source agent purpose-built for Lean 4 proof engineering. One pass costs $18; pass@2 beats Claude Sonnet at roughly 1/15 the cost.
READ POST ↗Meta's Avocado Is Not Ripe Enough: Flagship Model Slips to May at the Earliest
Internal benchmarks reportedly place Meta's Avocado between Gemini 2.5 and Gemini 3, delaying the flagship from March to at least May. CNBC had earlier flagged a possible break from open source.
READ POST ↗NVIDIA Open-Sources Nemotron 3 Super, a 120B MoE for Agents
On March 11, 2026, NVIDIA open-sourced Nemotron-3-Super-120B-A12B: hybrid Mamba plus latent MoE, 1M context, 478 tokens/sec, and 10T+ tokens of training data aimed at agentic inference costs.
READ POST ↗Qwen3.5 Small Models: 9B Rivals gpt-oss-120b at the Edge
Alibaba adds 0.8B-9B small models to Qwen3.5 under Apache-2.0: native 262K context, image and video input. The 9B beats gpt-oss-120b on MMLU-Pro and GPQA Diamond, targeting phones and laptops.
READ POST ↗Gemini 3.1 Flash Lite: Fastest, Cheapest Gemini 3 Model
Google shipped Gemini 3.1 Flash Lite in preview on March 3: $0.25 per million input tokens, up to 363 tokens per second, built for high-volume workloads. Pricing, benchmarks, and guidance.
READ POST ↗Perplexity Open-Sources pplx-embed Retrieval Models
Perplexity released pplx-embed-v1 and pplx-embed-context-v1 on Feb 26 — open weights at 0.6B and 4B, diffusion-pretrained and bidirectional, built for web-scale retrieval and RAG pipelines.
READ POST ↗Senate Bipartisan Bill Revives AI Standards and Testbeds
Four senators reintroduced the Future of AI Innovation Act on Feb 26: voluntary NIST AI standards, national-lab testbeds, prize competitions, and curated federal datasets reviving NAIAC advice.
READ POST ↗Gemini 3.1 Pro Lands in Preview: Double the Reasoning at the Same Price
On February 19, 2026, Google released Gemini 3.1 Pro in preview: 2x reasoning versus Gemini 3 Pro at the same pricing, on the gemini-3.1-pro-preview endpoint. Price-performance becomes the axis.
READ POST ↗MiniMax M2.5: Open Coding Weights at a Tenth of Opus Cost
MiniMax open-sources M2.5 and M2.5-Lightning: 80.2% SWE-Bench Verified, Opus-4.6-class task times at roughly a tenth of the cost, 229B parameters under modified MIT on Hugging Face.
READ POST ↗Mistral Voxtral Transcribe 2: Open Speech-to-Text, On-Device
Mistral's Voxtral Transcribe 2 ships batch and streaming transcription models: ~4% WER on FLEURS, Apache 2.0 open weights for the 4B realtime model, and $0.003/min API pricing.
READ POST ↗Qwen3-Max-Thinking Arrives: Alibaba's Trillion-Parameter Proprietary Reasoner
Alibaba Qwen team released Qwen3-Max-Thinking in late January 2026: a trillion-parameter proprietary reasoning flagship with adaptive tool use, aimed at enterprise model selection.
READ POST ↗Runway's Turing Reel: Only 9.5% Can Reliably Spot AI Video
Runway tested 1,043 viewers on real footage versus Gen-4.5 image-to-video clips: 57.1% overall accuracy, only 9.5% statistically significant. Detection is dead — provenance metadata is the plan.
READ POST ↗Ultralytics YOLO26: NMS-Free Vision AI Built for the Edge
Ultralytics shipped YOLO26 on January 14, 2026: NMS-free end-to-end detection, DFL removed, up to 43% faster nano CPU inference, five scales from edge to server, AGPL-3.0 dual licensing.
READ POST ↗Epoch AI: Chinese Models Trail US Frontier by Seven Months
Epoch AI's ECI index puts Chinese models seven months behind the US frontier on average since 2023, ranging four to fourteen months. How it is measured, and the open-weight catch.
READ POST ↗