2026
44 篇文章Firecrawl Developer Index:為 Coding Agents 而設的檢索層
Firecrawl 推出專為 coding agents 設計的 Developer Index,收錄 70M+ 開發工件,並附開放基準 DevDex。本文解析其設計動機、運作方式與實測表現。
閱讀文章 ↗Cursor Origin 上線:當 code hosting 開始為 agent 重新設計
Cursor 在 8 月 17 日推出 Origin code hosting 早期測試版,把 repos、pull requests 與 agent 收進同一個地方,GitHub 可雙向同步,Vercel、Depot、Buildkite 的整合也已上線。本文拆解這套堆疊式平台策略,以及團隊現在該不該進場。
閱讀文章 ↗Copilot canvases:把 agentic workflow 從聊天捲動搬上可審視的畫布
GitHub 在 Copilot app 推出 canvases,讓開發者與 agent 在持久、共享的表面上協作。本文拆解兩個實戰 canvas 的設計細節、2,000 到 3,000 AI credits 的成本帳本,以及 builder 可以直接帶走的四步落地藍圖。
閱讀文章 ↗Grok 4.6 評測:61 分重返前沿,代理任務變主戰場,價格不變
SpaceXAI 於 8 月 12 日發布 Grok 4.6:Artificial Analysis 智能指數 61 分、追平 GPT-5.6 Sol,定價維持每百萬 token 輸入 2 美元、輸出 6 美元,主打長時間代理工作。
閱讀文章 ↗DeepSeek 開源 Agent Harness:一切皆外掛的開發者預覽
2026 年 8 月 13 日,DeepSeek 以 MIT 授權公開 DeepSeek Harness 開發者預覽:Cordis 核心做到一切皆外掛,附錄影重播的會話日誌與四種執行模式,原始碼與外掛生態同步開放。
閱讀文章 ↗Docker Sandboxes:用 microVM 給 AI 編碼代理圈出安全區
Docker 推出 Sandboxes:以拋棄式 microVM 隔離執行 Claude Code、Codex 等 AI 編碼代理,CLI 免費、組織治理另購,讓無人值守的代理執行有標準做法。
閱讀文章 ↗GitHub 堆疊 PR 公開預覽:把大改動拆成一疊小 PR
2026 年 7 月 30 日 GitHub 把 stacked pull requests 推上公開預覽:每層獨立審查、一鍵整疊合併、上層自動 rebase,並提供 gh-stack CLI 擴充與 Copilot 編碼代理整合。
閱讀文章 ↗Claude Opus 4.7 來了:更耐操、更會驗證,但安全閘門也變多了
Anthropic 在 2026 年 4 月推出 Claude Opus 4.7,主打長任務穩定性與自我驗證能力,同時加入網路安全防護。本文整理測試者回饋與價格資訊,幫助產品開發者評估是否升級。
閱讀文章 ↗Block 開源 Buzz:團隊聊天、AI 代理與 Git 託管共用一條事件流
Block 於 2026 年 7 月 21 日開源工作區 Buzz,把團隊聊天、AI 代理人與 Git 託管放進同一套簽名事件系統,明言要減少對 Slack 與 GitHub 的依賴。
閱讀文章 ↗Kimi K3 登場:2.8T 參數、百萬 Token Context,開源模型走向長程 Agent
Moonshot 發布 Kimi K3:2.8T 參數的開源 3T 級模型,原生視覺、1M token context,主打長程 coding 與知識工作。本文整理 Delta Attention 架構、kernel 編譯器與晶片設計等案例、API 定價與 64 卡部署建議,以及官方自認仍落後 Fable 5 與 GPT 5.6 Sol 的誠實定位。
閱讀文章 ↗Grok Build 開源:Coding Agent 最值得讀的是 Harness,不是 UI
xAI 把 Grok Build 的完整 harness 開源:agent loop、工具、terminal UI 與擴充系統四大塊一次公開。本文整理開源範圍、「原始碼才是 definitive reference」的論點,以及 local-first 設定對 builder 的意義。
閱讀文章 ↗Bun 1.4 改寫成 Rust:64 個 Claude 代理 11 天完成的大搬遷
JavaScript 執行環境 Bun 釋出 1.4.0,首個以 Rust 重寫的版本:Claude Code 動態工作流程搭配 Claude Fable 5,以 64 個並行代理在 11 天內移植約 53.5 萬行 Zig 程式碼,API 成本約 16.5 萬美元。
閱讀文章 ↗Cognition 推出 Devin Security Swarm:代理群驗證漏洞並自動修補
Cognition 於 2026 年 7 月 1 日推出 Devin Security Swarm:多代理掃描程式碼庫、在沙箱驗證可利用性並開出修補 PR,評測檢出率 72%、單次掃描 90.23 美元。
閱讀文章 ↗Sourcegraph Agentic Batch Changes 公測:代理自動完成大規模程式碼遷移
Sourcegraph 於 2026 年 6 月 30 日宣布 Agentic Batch Changes 公開測試:以自然語言描述遷移目標,代理負責界定範圍、執行、讀 CI 修錯,直到數百個儲存庫的每個 PR 都可合併,公測期間免費。
閱讀文章 ↗Mistral 開源 Leanstral 1.5:miniF2F 滿分、每題 4 美元的 Lean 證明
Mistral 於 2026 年 7 月 2 日釋出 Apache-2.0 授權的 Leanstral 1.5:119B 總參數、約 6B 活躍,在 miniF2F 拿下滿分、PutnamBench 解出 587 題,把 Lean 4 證明成本壓到每題約 4 美元。
閱讀文章 ↗Claude Code 六月底連發:組織預設模型上線、MCP 邊界補強
6 月 22 日至 29 日 Claude Code 連發七版:組織可設定預設模型與模型限制,同時修補 .mcp.json 自我核准與 MCP OAuth 範圍等安全問題。
閱讀文章 ↗OpenRouter MCP:讓 Coding Agent 用即時價格與評測選模型
OpenRouter 官方 MCP server 把即時模型目錄、Artificial Analysis 與 Design Arena 排名、價格、文件與測試推論交給 coding agent:十一個工具只有一個計費,OAuth 專屬 key 自帶七天效期與十美元上限。本文拆解接入方式、文中範例流程與安全邊界。
閱讀文章 ↗Cloudflare 臨時帳號讓 AI 代理免註冊直接部署
Cloudflare 推出臨時帳號:AI 代理執行 wrangler deploy --temporary 即可免帳號、免金鑰部署到 Workers,60 分鐘內可由人類認領,逾期自動刪除。本文解析流程、限制與安全模型。
閱讀文章 ↗Qwen Code 的 /fork、/skills 與跨專案記憶:把 Agent 的日常摩擦拆開修
Qwen Code v0.18.0-preview 用 /fork 派發繼承完整 context 的背景 Agent、/skills 可搜尋可鎖定的管理面板,加上 user 層級跨專案記憶,一次處理主對話阻塞、能力治理與跨 Repo 偏好遺忘三個日常摩擦。
閱讀文章 ↗從 Nextdoor 看 Codex:工程師不再只是寫程式,而是打造產品
Nextdoor 工程團隊如何使用 OpenAI Codex 將工程師從「迭代提示」轉向「結果工程」,讓單一工程師能端到端擁有產品體驗,並將瓶頸從工程轉移到策略。
閱讀文章 ↗Ollama 0.30 更新:GGUF 原生支援、NVIDIA 最高 20% 效能提升、Vulkan 預設啟用
Ollama 0.30 透過 llama.cpp 深化 GGUF 引擎:NVIDIA 吞吐最高提升 20%(RTX 5090 實測條件)、Vulkan 預設開啟讓 AMD 與 Intel 開箱即用、LFM 與 Prism 家族及 Unsloth 微調可直接執行,tool calling 能力沿用並可掛上 coding agent。
閱讀文章 ↗Cognition 募資 10 億美元:估值八個月翻倍至 260 億
2026 年 5 月 27 日,AI 編程新創 Cognition 宣布募資超過 10 億美元,投後估值 260 億美元,八個月內翻倍。年化營收達 4.92 億美元,Devin 企業用量連六個月月增 50%,市場押注獨立編程工具能撐過模型廠夾擊。
閱讀文章 ↗Runtime(YC P26)上線:把沙盒化 coding agents 開放給整個團隊
YC P26 團隊 Runtime 於 5 月 21 日在 Hacker News 發布:讓非工程同事也能安全使用 Claude Code 與 Codex,以環境快照、沙盒編排、密鑰代理與 RBAC 控制風險,核心開源、不抽 token 加價。
閱讀文章 ↗從「打得更快」到「想得更深」:Sea 如何用 Codex 重塑工程團隊
Sea 的 CPO David Chen 分享公司全面導入 Codex 的經驗:87% 週活躍率、從 autocomplete 到 agentic workflow 的轉變,以及工程師如何轉型為系統編排者。
閱讀文章 ↗OpenAI 如何安全部署 Codex:從沙箱到代理原生日誌
OpenAI 公開了內部部署 Codex 的安全控制框架,包括沙箱、審批策略、網路限制與代理原生遙測,為企業導入 coding agent 提供參考。
閱讀文章 ↗Redis 之父 antirez 開源 ds4:跑得動前沿模型的本地推理引擎
Redis 作者 antirez 用一週寫出 ds4(DwarfStar):MIT 授權、C 語言的本地推理引擎,以 2/8-bit 非對稱量化讓 DeepSeek V4 Flash 與 GLM 模型跑進 96–128GB 的 Mac,內建原生編碼代理,登上 HN 首頁。
閱讀文章 ↗Anthropic 認了:Claude Code 變笨的三個原因與補救承諾
Anthropic 發布工程的事後解析,承認三個變更讓 Claude Code 一個月來變笨又健忘:調低推理努力、快取 bug 清掉思考紀錄、限制長度的系統提示詞。全部已修復並重設訂閱者用量上限,本文拆解時間軸與承諾。
閱讀文章 ↗同一顆模型、兩倍差距:四款 CLI 編碼 Agent 腳手架實測
開發者 Charles Azam 讓四款開源 CLI 編碼 Agent 接上同一顆 GLM-4.7 跑 Terminal-Bench 2.0:Mistral Vibe 拿 0.35、Codex 只有 0.15。結論是腳手架主導成績,模型之外的程式碼決定了兩倍以上的差距。
閱讀文章 ↗Codex 更新:從寫 Code 到幫你操作電腦
OpenAI 在 2026 年 4 月 16 日釋出 Codex 重大更新,加入背景電腦操作、排程任務、記憶功能,並擴充外掛生態。本文整理這些變化的實際用途與限制。
閱讀文章 ↗Android CLI 與官方 Skills:Google 把代理開發帶進終端機
Google 於 2026 年 4 月 16 日推出 Android CLI、Skills 與 Knowledge Base,讓 Gemini CLI、Claude Code、Codex 等任何代理都能在 Android Studio 外高效開發;token 消耗減少逾 70%、任務快 3 倍。
閱讀文章 ↗當部署者變成機器:Vercel 提出的「Agentic Infrastructure」是什麼?
Vercel 提出 Agentic Infrastructure 三層架構:從給 coding agent 部署的基礎設施,到用來建構 agent 的平台,再到基礎設施本身具備自主維運能力。
閱讀文章 ↗Cursor 3 把 IDE 變成 Agent 工作台:從 fork 走向從零打造
Cursor 3 以 agents 為中心從零打造統一工作區:所有 local 與 cloud agents 收進同一側欄、雙向 handoff、內建 diffs 與 browser,以及 Marketplace 插件生態。本文拆解每項機制、發布裡缺少的數字,以及 builder 的採用建議。
閱讀文章 ↗GLM-5.1 送進 Coding Plan:Z.ai 把 8 小時自主編程變成訂閱規格
2026 年三月下旬,智譜 Z.ai 向 Coding Plan 訂閱者推出 GLM-5.1,支援最長 8 小時的自主編程運行。本文看長時程 coding agent 的工程考驗,與開源旗艦轉向訂閱制的商業邏輯。
閱讀文章 ↗OpenAI 推 Codex 學生方案:北美大學生免費領 100 美元額度
2026 年 3 月 20 日 OpenAI 推出 Codex for Students:美加地區通過驗證的大學生可免費獲得 100 美元(2,500 點)ChatGPT 額度,專用於代理編碼工具 Codex,效期 12 個月。同期 GitHub 也把學生免費方案移轉到新版 Copilot Student plan,學生市場戰開打。
閱讀文章 ↗Mistral 開源 Leanstral:專為 Lean 4 證明工程打造的 120B 稀疏模型
Mistral 以 Apache 2.0 釋出 Leanstral-120B-A6B,第一個專為 Lean 4 設計的開源程式代理。單次推理 18 美元,pass@2 便以約 15 分之一的成本超越 Claude Sonnet。本文解析其架構、FLTEval 成績與三種部署方式。
閱讀文章 ↗OpenAI 推出 Codex Security 研究 preview:找漏洞也幫你修的資安 agent
2026 年 3 月 6 日,OpenAI 把 Codex Security 推進研究 preview:這個應用資安 agent 在程式碼庫中找漏洞並協助修復,鎖定 ChatGPT Enterprise、Business 與 Education 客戶,主打大規模程式碼掃描。
閱讀文章 ↗Cursor agents 大更新:自己測試修改、自己錄下工作過程
2026 年 2 月 24 日,Cursor 宣布 agents 重大更新:coding agent 現在會測試自己的修改,並以影片、log 與截圖記錄工作過程。CNBC 以 AI coding agent 之戰升溫為框架報導。本文分析自我驗證與可審計性為何成為戰場核心。
閱讀文章 ↗Qodo 2.1 推出 Rule System:治好 coding agent 的團隊規範失憶症
2026 年 2 月 17 日,AI 程式碼審查平台 Qodo 發布 2.1 版,推出首個持續學習的 Rule System:自動從過往 PR 萃取團隊規範、由專責代理強制執行並量化違規,VentureBeat 報導精準度提升 11%,為 coding agent 補上治理層。
閱讀文章 ↗16 個 Claude 代理寫出 C 編譯器:兩週、10 萬行、2 萬美元
Anthropic 研究員 Nicholas Carlini 以 16 個平行 Claude Opus 4.6 代理、近 2,000 個工作階段與約 2 萬美元,在兩週內寫出 10 萬行 Rust 的 C 編譯器,可編譯出可開機的 Linux 6.9 核心。本文拆解其極簡編排與限制。
閱讀文章 ↗GPT-5.3-Codex 登場:OpenAI 最強 coding model 參與了自己的開發
2026 年 2 月 5 日,OpenAI 發表 GPT-5.3-Codex,定位為迄今最強的 agentic coding model,媒體報導並指它「幫忙建造了自己」——參與了自身的開發管線。距離 GPT-5.2-Codex 僅三週,本文解析發布節奏與自我開發宣稱的意義。
閱讀文章 ↗Mistral Vibe 2.0 登場:終端機編碼代理轉為付費,進攻企業市場
2026 年 1 月 27 日,Mistral 把終端機編碼代理 Vibe 升級至 2.0:新增子代理、澄清提問、技能指令與代理模式,綁進 Le Chat Pro 與 Team 付費方案,Devstral 2 同步轉為付費 API。
閱讀文章 ↗Ollama 推出 ollama launch:一行指令把 coding agent 接上本機 model
Ollama 於 2026 年 1 月 23 日發表 ollama launch,一行指令自動安裝並設定 Claude Code、OpenCode 與 Codex,讓這些 coding agent 直接對接本機 model。本文解析本地開發工作流在隱私與成本上的實際影響。
閱讀文章 ↗GPT-5.2-Codex 登場:40 萬 token context、原生壓縮與更強工具呼叫
2026 年 1 月 14 日,OpenAI 發布 GPT-5.2-Codex,具備約 40 萬 token 的 context window、原生 compaction 與更佳的 tool calling,同日登上 GitHub Copilot。本文解析對 coding agent 工作流的實際影響。
閱讀文章 ↗Cursor CLI 更新:模型、規則與 MCP 管理走進終端機
2026 年 1 月 8 日,Cursor 發布 CLI 更新:agent 成為主要入口,新增模型、規則、指令與 MCP 管理命令,hooks 改為平行執行、延遲降為十分之一。本文解析這波更新對終端機 coding agent 自動化的意義。
閱讀文章 ↗
2025
2 篇文章Mistral 推出 Mistral Code 進攻企業AI編程
2025年6月4日,法國 Mistral AI 發布企業版 AI 編程助手 Mistral Code,整合 Codestral、Devstral 等四款模型,支援 JetBrains 與 VS Code,可私有化部署並在私有程式碼庫上精調,直接對陣 GitHub Copilot 與 Cursor。
閱讀文章 ↗Windsurf 稱 Anthropic 緊縮 Claude 直接存取
2025年6月3日,AI 編程工具 Windsurf 執行長 Varun Mohan 表示,Anthropic 以不到五天的通知刪減其 Claude 3.x 模型的第一方容量,公司被迫轉向第三方運算,突顯模型供應商與應用層之間的權力失衡。
閱讀文章 ↗
2026
42 ARTICLESFirecrawl Developer Index: A Specialized Retrieval Layer for Coding Agents
Firecrawl launches Developer Index for coding agents, indexing 70M+ artifacts with semantic retrieval and DevDex benchmark.
READ POST ↗Cursor Origin: Code Hosting Redesigned for AI Agents
Cursor launches Origin code hosting in early beta, designed for agent-scale workflows with GitHub sync, app integrations, and faster cloud agents via builds.
READ POST ↗Making Agentic Workflows Visible, Steerable, and Cost-Efficient with GitHub Copilot Canvases
Explore how GitHub Copilot canvases turn agentic workflows into durable, inspectable systems, with real examples and cost insights.
READ POST ↗Grok 4.6: 61 on the Intelligence Index, Frontier Again
SpaceXAI shipped Grok 4.6 on August 12: 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol, with pricing unchanged at $2 and $6 per million tokens.
READ POST ↗DeepSeek Open-Sources Agent Harness in Developer Preview
DeepSeek's agent harness enters an MIT-licensed developer preview: a Cordis plugin kernel, a replayable append-only session log, and four runtime modes, open from day one.
READ POST ↗Docker Sandboxes: microVM Isolation for AI Coding Agents
Docker Sandboxes runs AI coding agents like Claude Code and Codex in disposable microVMs with filesystem, network, and credential controls. The CLI is free; governance is paid.
READ POST ↗Stacked Pull Requests Arrive on GitHub in Public Preview
GitHub's stacked pull requests enter public preview: review each layer independently, merge a stack in one click, plus a gh-stack CLI extension and Copilot coding-agent skill.
READ POST ↗Claude Opus 4.7: A Practical Guide for Product Builders
Claude Opus 4.7 brings better long-task reliability, self-verification, and vision. Learn what changed, how to use it, and key trade-offs.
READ POST ↗Block Open-Sources Buzz: Chat, AI Agents, Git in One Place
Block open-sourced Buzz on July 21, 2026: team chat, AI agents as workspace members, and Git hosting on one signed-event stream, to cut reliance on Slack and GitHub.
READ POST ↗Kimi K3 Arrives: 2.8T Parameters, Million-Token Context, Open Models Go Long-Horizon
Kimi K3: a 2.8T open 3T-class model with native vision and 1M-token context for long-horizon coding. The architecture, kernel and chip cases, pricing, and the gap to Fable 5 and GPT 5.6 Sol.
READ POST ↗Grok Build Open Source: The Harness Is the Part Worth Reading, Not the UI
xAI open-sourced Grok Build's complete harness: agent loop, tools, terminal UI, and extension system — and why the source itself is the definitive reference for extension authors.
READ POST ↗Bun 1.4 Is Rust Now: 64 Claude Agents Ported It in 11 Days
Bun v1.4.0 is the first Rust-written release of the JS runtime: Claude Code workflows plus Claude Fable 5 ported 535k lines of Zig in 11 days with 64 agents, for about $165k.
READ POST ↗Qwen Code Update: Auto Model Fallback and Nested Subagents for Reliable AI Agents
Qwen Code's v0.19.6-0.19.8 releases center on reliability: fallback masks only capacity and rate-limit errors, nested sub-agents get depth caps and two-layer protection, plus session upgrades.
READ POST ↗Devin Security Swarm: Agents Verify Exploits and Ship Fixes
Devin Security Swarm, launched July 1, 2026, uses parallel agents to scan codebases, verify exploitability in sandboxes, and open remediation PRs — 72% recall at $90.23 per run.
READ POST ↗Sourcegraph's Agentic Batch Changes Enters Public Beta
Announced June 30, 2026: Sourcegraph's Batch Changes engine becomes an agent that scopes, executes, and ships migrations across hundreds of repos until every PR is mergeable.
READ POST ↗Leanstral 1.5: Mistral Saturates miniF2F at $4 a Proof
Mistral's Apache-2.0 Leanstral 1.5 saturates miniF2F, solves 587 PutnamBench problems at roughly $4 each, and turns Lean 4 proof engineering into a cheap, repeatable routine.
READ POST ↗Claude Code Late June: Org Model Defaults, MCP Hardening
Between June 22 and 29, Claude Code shipped seven releases: org-level default models and model restrictions, plus fixes closing MCP self-approval and OAuth scope holes.
READ POST ↗Cloudflare Temporary Accounts: Agent Deploys Without Signup
Cloudflare lets AI agents deploy to Workers with no signup: wrangler deploy --temporary provisions a 60-minute account a human can claim, or it is auto-deleted.
READ POST ↗Qwen Code's /fork, /skills, and Cross-Project Memory: Fixing an Agent's Daily Friction
Qwen Code v0.18.0-preview adds /fork background agents that inherit full context, a searchable lockable /skills panel, and user-level cross-project memory in one weekly update.
READ POST ↗From Nextdoor to Codex: How Outcome Engineering Is Redefining the Role of the Engineer
Nextdoor's engineering head explains how Codex shifts engineers from prompting agents to outcome engineering, compressing timelines and moving bottlenecks to strategy.
READ POST ↗AI Coding Startup Cognition Raises $1B at $26B Valuation
Cognition raised over $1B at a $26B post-money valuation, double its September price, on $492M annualized revenue and Devin enterprise usage up 50% month over month.
READ POST ↗Runtime (YC P26): Sandboxed Coding Agents for Whole Teams
YC P26 startup Runtime launched May 21 on Hacker News: sandboxed Claude Code and Codex sessions for whole teams, with env snapshots, secret proxy, and RBAC. Open core.
READ POST ↗From Typing Faster to Thinking Better: How Sea Uses Codex to Reshape Engineering Teams
Sea's David Chen on rolling out Codex across the org: 87% weekly active users, AI agents in CI/CD, and the shift to system orchestration.
READ POST ↗How OpenAI Deploys Codex Safely: Sandboxes, Rules, and Agent-Native Logs
OpenAI shares its internal framework for deploying Codex safely: sandboxing, approval policies, network controls, and agent-native telemetry for auditing and triage.
READ POST ↗antirez's ds4: A Local LLM Inference Engine Built for Metal
Redis creator antirez open-sourced ds4: a C local inference engine tuned for DeepSeek V4 Flash and GLM on Metal, CUDA and ROCm, with a coding agent. MIT licensed.
READ POST ↗Anthropic Postmortem: Why Claude Code Got Worse for a Month
Anthropic's April 23 postmortem ties a month of Claude Code complaints to three changes: lower reasoning effort, a cache bug wiping thinking history, and a verbosity prompt rule.
READ POST ↗Same Model, 2x Gap: Benchmarking Four CLI Coding Agents
Four open-source CLI coding agents running the same model (GLM-4.7) on Terminal-Bench 2.0: Mistral Vibe scored 0.35, Codex 0.15. The scaffolding decides, not the model.
READ POST ↗Codex Update: From Writing Code to Operating Your Computer
OpenAI's Codex update adds background computer use, an in-app browser, memory, automations, and 90+ plugins, making it a more proactive partner across the software development…
READ POST ↗Android CLI: Google's Terminal-First Bet on Coding Agents
Google's Android CLI, open-source Skills, and Knowledge Base let any agent — Gemini CLI, Claude Code, Codex — build Android apps; setup tokens fell over 70%, tasks ran 3x faster.
READ POST ↗Agentic Infrastructure: What Vercel's Three-Layer Model Means for AI Builders
Why coding agents now drive 30% of Vercel deployments, and the three infrastructure shifts that enable machine-driven development.
READ POST ↗GLM-5.1 Arrives for Coding Plan Subscribers: Z.ai Makes 8-Hour Coding Runs a Product Spec
In late March 2026, Zhipu's Z.ai shipped GLM-5.1 to Coding Plan subscribers, with support for autonomous coding runs of up to 8 hours. What it changes for long-horizon coding agents.
READ POST ↗OpenAI Gives Students $100 in Codex Credits
OpenAI launched Codex for Students: verified university students in the U.S. and Canada get $100 of ChatGPT credit usable only in Codex, valid 12 months. GitHub moved students too.
READ POST ↗Mistral Open-Sources Leanstral, a Lean 4 Proof Agent
Mistral released Leanstral-120B-A6B under Apache 2.0 — the first open-source agent purpose-built for Lean 4 proof engineering. One pass costs $18; pass@2 beats Claude Sonnet at roughly 1/15 the cost.
READ POST ↗OpenAI Ships Codex Security in Research Preview: A Security Agent That Finds and Helps Fix
OpenAI moved Codex Security into research preview on March 6, 2026: an app-security agent that finds and helps fix vulnerabilities in codebases, aimed at Enterprise, Business, and Education plans.
READ POST ↗Cursor's Big Agents Update: Coding Agents That Test Their Own Work and Record It
On Feb 24, 2026, Cursor shipped a major agents update: coding agents now test their own changes and record their work via video, logs, and screenshots. Verification is becoming the real battleground.
READ POST ↗Qodo 2.1 Adds a Rule System to Cure Coding Agent Amnesia
Qodo 2.1, out February 17, 2026, adds a continuous-learning Rule System to AI code review: auto-discovered team standards, deterministic enforcement, violation analytics, and an 11% precision boost.
READ POST ↗16 Claude Agents Built a C Compiler: 100K Lines, $20K
Anthropic researcher Nicholas Carlini ran 16 parallel Claude Opus 4.6 agents over ~2,000 Claude Code sessions to write a 100,000-line Rust C compiler that boots Linux 6.9 — for under $20,000.
READ POST ↗GPT-5.3-Codex Arrives: OpenAI's Strongest Coding Model Helped Build Itself
OpenAI introduced GPT-5.3-Codex on February 5, 2026 — its most capable agentic coding model yet, one that helped build itself. Three weeks after GPT-5.2-Codex, the cadence is the story.
READ POST ↗Mistral Vibe 2.0: Terminal Coding Agent Goes Paid
Mistral's terminal coding agent Vibe 2.0 adds subagents, clarifying questions, skills, and modes — now a paid product, with Devstral 2 moving to a paid API at $0.40 in / $2.00 out per million tokens.
READ POST ↗Ollama Ships ollama launch: One Command to Wire Up Local Coding Agents
Ollama January 23, 2026 post introduced ollama launch: one command that installs and configures Claude Code, OpenCode, and Codex against local models. What it changes for coding workflows.
READ POST ↗GPT-5.2-Codex Arrives: 400k-Token Context, Native Compaction, Better Tool Calling
On January 14, 2026, OpenAI released GPT-5.2-Codex: roughly 400k-token context, native compaction, better tool calling — GA in GitHub Copilot the same day. What it changes for coding agents.
READ POST ↗Cursor CLI Update: Model, Rules and MCP Control in Terminal
Cursor's January 8, 2026 CLI release makes agent the primary entrypoint, adds commands for models, rules and MCP management, and runs hooks in parallel for a 10x latency cut.
READ POST ↗
2025
2 ARTICLESMistral Code brings agentic coding to enterprises
On June 4, 2025, Mistral AI launched Mistral Code, an enterprise coding assistant bundling four models for JetBrains and VS Code, with self-hosting and private fine-tuning for Copilot-class rivals.
READ POST ↗Windsurf says Anthropic cut its direct Claude access
On June 3, 2025, Windsurf said Anthropic cut its first-party Claude 3.x capacity with under five days notice, forcing third-party compute — a case study in model providers power over app builders.
READ POST ↗