2026
13 篇文章Grok 接上 Coinbase:當 agent 能直接動你的交易所帳戶
Grok 在 2026 年 9 月 9 日推出原生 Coinbase 連接器,可在對話中查餘額、分析持倉並直接下單。本文整理可用範圍與批准機制,並看馬斯克的賠償承諾與 100 美元條款上限之間的落差。
閱讀文章 ↗當模型開始替你的系統做測試:Perplexity 把 GPT‑6 Astra 放進端到端流程
Perplexity 讓 GPT‑6 Astra 代寫測試程式、模擬外部服務並監控正式系統,檢查頻率明顯低於前幾代模型。
閱讀文章 ↗把測試證據放進開發流程:Cognition 用 GPT‑6 Astra 讓 Devin 自己驗證成果
Cognition 將 GPT‑6 Astra 整合進 Devin,讓代理在回報程式碼變更時附上測試錄影與涵蓋範圍報告,減少工程師手動審查。
閱讀文章 ↗Claude Code 60 秒自動續跑:一個不良功能的完整解剖
Claude Code 2.1.198 曾在提問 60 秒無回應後自動續跑且未寫進 changelog;開發者 Olaf Alders 以二進位比對還原事件全貌,解析這場兩天內回滾的風波與對 Agent 安全的啟示。
閱讀文章 ↗Qwen Code 自動切換模型、Subagent 再分工:可靠性比平行數更重要
Qwen Code 週更 v0.19.6–0.19.8 的主軸是可靠性工程:model fallback 只掩蓋容量與限流錯誤、巢狀 subagent 有深度上限與雙層保護,加上 session 管理與參數級權限。本文拆解這些設計背後的分界線。
閱讀文章 ↗微軟開源 ASSERT:把文字規格變成 AI 行為測試套件
2026 年 6 月 2 日,微軟開源 ASSERT 框架:開發者以自然語言描述代理應有行為,它自動生成測試情境、用 LLM 評審計分,輸出傷害與權衡兩類指標,並可掛進 CI 做回歸把關。
閱讀文章 ↗Meta AI 客服機器人成了帳號綁架工具:Instagram 通知受駭用戶
駭客只對 Meta 的 AI 客服機器人說「這是我的帳號」,機器人就把受害者 Instagram 綁到駭客信箱。白宮舊帳號與太空軍高階士官長相繼淪陷後,Meta 稱漏洞已修復,並開始通知受影響用戶。本文拆解這場低技術門檻的 AI 安全面事件。
閱讀文章 ↗Cursor 花了一年才搞懂:雲端代理最難的不是模型,是環境
Cursor 回顧一年來打造雲端代理的工程教訓:最難的不是模型而是環境——開發環境本身就是產品、可靠性要靠任務編排解決,架構上把代理與基礎設施解耦,是本地代理上雲最值得借鑑的三條經驗。
閱讀文章 ↗Anthropic 認了:Claude Code 變笨的三個原因與補救承諾
Anthropic 發布工程的事後解析,承認三個變更讓 Claude Code 一個月來變笨又健忘:調低推理努力、快取 bug 清掉思考紀錄、限制長度的系統提示詞。全部已修復並重設訂閱者用量上限,本文拆解時間軸與承諾。
閱讀文章 ↗Gartner:2028 年 25% 企業 GenAI 應用每年至少 5 次資安事件
Gartner 4 月 9 日預測:2028 年 25% 企業生成式 AI 應用每年至少 5 次小型資安事件(2025 年為 9%),2029 年 15% 每年至少一次重大事件。主因是 MCP 與 agentic AI 擴散,安全審查趕不上部署速度。
閱讀文章 ↗Redwood 首席科學家的 AI 現況快照:1.6 倍研發加速與 8% 失準事件機率
2026 年 4 月 7 日,Redwood Research 首席科學家 Ryan Greenblatt 發表長文,估計前沿實驗室工程加速已達 1.6 倍、整體 AI 進度僅 1.15 至 1.2 倍,並給出 8% 嚴重目標偏離事件機率與 60% 半年內自主開發漏洞的機率。本文拆解數字與推論。
閱讀文章 ↗Cursor agents 大更新:自己測試修改、自己錄下工作過程
2026 年 2 月 24 日,Cursor 宣布 agents 重大更新:coding agent 現在會測試自己的修改,並以影片、log 與截圖記錄工作過程。CNBC 以 AI coding agent 之戰升溫為框架報導。本文分析自我驗證與可審計性為何成為戰場核心。
閱讀文章 ↗Qodo 2.1 推出 Rule System:治好 coding agent 的團隊規範失憶症
2026 年 2 月 17 日,AI 程式碼審查平台 Qodo 發布 2.1 版,推出首個持續學習的 Rule System:自動從過往 PR 萃取團隊規範、由專責代理強制執行並量化違規,VentureBeat 報導精準度提升 11%,為 coding agent 補上治理層。
閱讀文章 ↗
2026
11 ARTICLESGrok Can Now Trade Your Coinbase Account in Chat
Grok's native Coinbase connector can check balances, analyze holdings, and place trades in chat. What it covers, how approval works, and the gap between Musk's pledge and the $100 liability cap.
READ POST ↗What It Takes to Hand an Agent the Whole System
Perplexity lets GPT-6 Astra edit production systems and check in less often, shifting the trust question to oversight.
READ POST ↗Devin Now Shows Its Work: What Self-Testing Agents Change for Review
Cognition uses GPT-6 Astra so Devin tests its own changes and returns evidence, not just a diff.
READ POST ↗Claude Code's 60-Second Auto-Continue, Dissected
Claude Code 2.1.198 silently auto-continued after 60 seconds with no changelog mention. Olaf Alders diffed the binary to reconstruct the two-day rollback and its lessons.
READ POST ↗Microsoft ASSERT Turns Text Specs into AI Behavior Tests
Microsoft open-sourced ASSERT, a framework that turns plain-language behavior specs into generated test suites with LLM-judge scoring and CI regression gates for AI agents.
READ POST ↗Meta AI Support Bot Let Hackers Steal Instagram Accounts
Hackers told Meta's AI support bot an account was theirs; the bot complied and linked attacker emails. Meta says the flaw is fixed and is notifying targeted users.
READ POST ↗Anthropic Postmortem: Why Claude Code Got Worse for a Month
Anthropic's April 23 postmortem ties a month of Claude Code complaints to three changes: lower reasoning effort, a cache bug wiping thinking history, and a verbosity prompt rule.
READ POST ↗Gartner: GenAI Security Incidents to Nearly Triple by 2028
Gartner: by 2028, 25% of enterprise GenAI apps will see 5+ minor security incidents yearly, up from 9% in 2025. The driver is MCP-powered agentic AI outpacing security review.
READ POST ↗Redwood Sizes Up AI: 1.6x Speed-Up, 8% Misalignment Odds
Redwood's Ryan Greenblatt estimates a 1.6x engineering speed-up, an 8% chance of a serious misalignment incident, and 60% odds of autonomous exploits within six months.
READ POST ↗Cursor's Big Agents Update: Coding Agents That Test Their Own Work and Record It
On Feb 24, 2026, Cursor shipped a major agents update: coding agents now test their own changes and record their work via video, logs, and screenshots. Verification is becoming the real battleground.
READ POST ↗Qodo 2.1 Adds a Rule System to Cure Coding Agent Amnesia
Qodo 2.1, out February 17, 2026, adds a continuous-learning Rule System to AI code review: auto-discovered team standards, deterministic enforcement, violation analytics, and an 11% precision boost.
READ POST ↗