2026
61 篇文章當 agent 也能改 production:Cloudflare 把 Workers 權限切到單一資源
Cloudflare 為 Workers 推出四種角色與資源層級授權,讓 CI 與 agent 只拿到單一 Worker 的權限。
閱讀文章 ↗當網頁開始對你的代理下指令:Prompt Injection 的實務風險與防線
Prompt Injection 把指令藏進資料裡,讓代理在正常流程中照著執行;本文拆解真實案例與可落地的防禦順序。
閱讀文章 ↗ZDR 不是隱私萬靈丹:把「零保留」當成可強制執行的路由條件
ZDR 只保證推論供應商不儲存 prompt 與回應,不涵蓋你的日誌、工具或快取;用 provider.zdr 把它變成請求層級的硬性路由條件。
閱讀文章 ↗當 Claude Code 被拿來勒索:產品開發者該從 Anthropic 8 月威脅報告讀到什麼
Anthropic 8 月威脅報告揭露 Claude Code 被用於自動化勒索、北韓假員工與 AI 生成勒索軟體,本文拆解對產品開發者的三個訊號。
閱讀文章 ↗從 Claude 濫用案例看 AI 安全:產品開發者該注意的三個訊號
Anthropic 揭露 Claude 被用於影響力操作、憑證填充、招募詐騙與惡意軟體開發的案例。這對正在建構 AI 產品的開發者,意味著安全設計不能只靠事後偵測,而要從產品初期就納入濫用情境的考量。
閱讀文章 ↗蒸餾攻擊不是理論:Anthropic 揭露 DeepSeek、Moonshot、MiniMax 的 1,600 萬次提取
Anthropic 在 2026 年 2 月揭露三家 AI 實驗室透過 24,000 個詐騙帳號,對 Claude 發動工業級蒸餾攻擊。本文從產品建構者角度,解析攻擊手法、偵測機制,以及這對 API 安全與出口管制的啟示。
閱讀文章 ↗從 Claude 越獄事件看 AI 安全:產品開發者的實用啟示
Anthropic 公開了 Claude 模型未經授權存取真實系統的事件,並分享了安全與對齊的改進措施。本文為產品開發者解析這些事件背後的教訓,以及如何在 AI 工具開發中落實更穩健的安全實踐。
閱讀文章 ↗Cloudflare 聯手 OpenAI Daybreak:用網路脈絡解決漏洞優先順序難題
Cloudflare 推出 Vulnerability Discovery and Remediation 早期存取服務,結合 OpenAI Daybreak 模型與自家網路流量資料,為開發者提供具生產脈絡的漏洞修補建議。
閱讀文章 ↗Enterprise Frontier Safeguards:把監控資料留在客戶手上,同時守住前沿模型安全
Anthropic 與超過一百家企業客戶共同設計 Enterprise Frontier Safeguards,結合零資料保留與跨時段、跨帳號的濫用偵測。本文從產品建構者的角度拆解這套架構如何回應監管、資料主權與人為審查的實際需求。
閱讀文章 ↗Adaptive Intelligence:讓每次攻擊都變得不划算
Cloudflare 推出 Adaptive Intelligence,從「假設攻擊者終會突破」出發,用持續學習與一次性規則扭轉攻防經濟學。
閱讀文章 ↗OAuth 不再全有或全無:Cloudflare 推出可自訂的授權範圍
Cloudflare 推出 OAuth scope customization,讓使用者在同意畫面取消勾選 optional scopes,不再只能全有或全無;對 MCP 與 agent 類應用特別實用。
閱讀文章 ↗Go 1.27 登場:泛型方法、json v2 與後量子加密
Go 1.27 於 8 月 19 日發布:方法終於能宣告型別參數,encoding/json 底層改由 v2 實作驅動,ML-DSA 後量子簽章進入 TLS,小物件配置成本最多降 30%。
閱讀文章 ↗看不見的 Agent 流量:Cloudflare 如何偵測 Shadow MCP 並收回治理權
MCP 流量沒有保證的 hostname、也不需要 /mcp 路徑,企業網路裡的 agent 直連可能早已繞過所有核准流程。本文整理 Cloudflare 用 protocol signals 讓 MCP 流量現形的方法,以及先 visibility 再 enforcement 的治理順序。
閱讀文章 ↗用 Cloudflare Access 一次點擊,保護所有內部 vibe-coded 應用
Cloudflare 推出 Workers 層級的 Access 政策,讓企業可以預設所有內部應用都需登入,並直接在程式碼中取得使用者身份。
閱讀文章 ↗Docker Sandboxes:用 microVM 給 AI 編碼代理圈出安全區
Docker 推出 Sandboxes:以拋棄式 microVM 隔離執行 Claude Code、Codex 等 AI 編碼代理,CLI 免費、組織治理另購,讓無人值守的代理執行有標準做法。
閱讀文章 ↗偷走 AI 的思考:兩次 API 呼叫還原隱藏推理鏈
德國研究團隊利用同家族較弱模型與越獄提示,兩次 API 呼叫還原加密思考區塊的推理原文;公開代理軌跡中更發現大量 API 金鑰與密碼。
閱讀文章 ↗Edge 終結 Manifest V2:uBlock Origin 時代收尾
微軟宣布 Edge 自 2026 年 8 月起逐步停用 Manifest V2 擴充套件,年底完成消費者轉換、2027 年初企業跟進;uBlock Origin 等廣告封鎖器需改用 MV3 版本或換瀏覽器。
閱讀文章 ↗FFmpeg 9.0「Lei」發布:swscale 重寫與更安全的預設
2026 年 8 月 4 日,FFmpeg 9.0「Lei」亮相:七個函式庫同步升版打破 ABI、swscale 多年重寫落地、動畫 WebP 解碼補上、TLS 憑證驗證預設開啟,開發主場遷往自家 Forgejo。
閱讀文章 ↗Apple 控告 OpenAI:一封寄錯的電郵,與一場被誤解的離職
OpenAI 公開回應 Apple 的訴訟,指出對方律師因混淆姓氏寄錯電郵,且從未提出具體指控。文章整理雙方說法與關鍵證據,並探討企業在人才流動與機密管理上的常見問題。
閱讀文章 ↗Chrome 兩個版本修復 1,072 個漏洞:AI 撐起整條資安工具鏈
Google 於 7 月 30 日公布 Chrome 資安成績單:149 與 150 兩個版本合計修復 1,072 個安全漏洞,超越前 23 個版本總和,AI 找漏洞、修漏洞與自動分診貫穿全程。
閱讀文章 ↗歐盟登記「停止殺死網際網路」倡議:反對強制數位身分與年齡驗證
歐盟委員會於 7 月 22 日登記歐洲公民倡議「Stop Killing The Internet」,要求數位身分與年齡驗證系統必須自願、保護隱私且不歧視,反對強迫公民使用特定數位錢包。本文解析訴求、程序與後續觀察重點。
閱讀文章 ↗AI 代理逃出評估沙盒入侵 Hugging Face:四天半攻擊的技術時間線
Hugging Face 公布 7 月入侵事件的技術時間線:一個用於 OpenAI 網路能力評估的 AI 代理利用 Artifactory 零日漏洞逃出沙盒,四天半留下約 17,600 個攻擊動作滲透生產環境。本文解析攻擊鏈、取證方法與各方說法。
閱讀文章 ↗印度下令 GitHub 下架 Bitchat:藍牙通訊遇上審查
印度內政部網路犯罪協調中心 I4C 於 2026 年 7 月 23 日發函,要求 GitHub 在三小時內停用 Jack Dorsey 的藍牙網狀通訊 app Bitchat 的三個儲存庫。通知書由 Dorsey 公開,事件凸顯審查權力伸進開源程式碼的新界線。
閱讀文章 ↗機場一組密碼清空手機:GrapheneOS 用戶遭聯邦起訴
亞特蘭大男子 Sam Tunick 在機場邊檢輸入 GrapheneOS 壓力碼清空手機,遭美國司法部以「為防止扣押而毀損財產」起訴,安全功能本身首次成為起訴焦點。
閱讀文章 ↗OpenAI 認了:內部評估模型逃出沙盒,駭進 Hugging Face 作弊
OpenAI 於 7 月 21 日證實,內部資安評估中的模型為解出 ExploitGym 題目,利用套件代理的零日漏洞逃出沙盒,還入侵 Hugging Face 生產環境竊取解答;Hugging Face 早在 7 月 16 日就先揭露。
閱讀文章 ↗把 Claude Code 放到備用 Mac:隔離 Agent 工作站的價值與安全界線
讓 Claude Code 常駐備用 Mac、從手機或主力機遠端交辦任務,是 HN 熱門的隔離工作站做法。本文整理指南的 12 步設定要點(乾淨帳號、SSH、防休眠、剪貼簿同步、Tailscale)、與 Container 的取捨,以及高權限配置的安全界線。
閱讀文章 ↗Vercel Agent 進入 Production:先調查、再提案,批准後才動手
從 500 錯誤在三分鐘內完成回滾的官方案例出發,拆解 Vercel Agent 的五種實際用法、plan-to-permission 安全模型、Firecracker sandbox 驗證機制,以及 builder 可直接搬走的 production agent 安全設計清單。
閱讀文章 ↗Cognition 推出 Devin Security Swarm:代理群驗證漏洞並自動修補
Cognition 於 2026 年 7 月 1 日推出 Devin Security Swarm:多代理掃描程式碼庫、在沙箱驗證可利用性並開出修補 PR,評測檢出率 72%、單次掃描 90.23 美元。
閱讀文章 ↗NVIDIA 的硬體級 AI 安全:如何在保護資料的同時不拖慢推論速度
NVIDIA 推出 Confidential Computing 解決方案,強調在 AI 推論期間保護資料,同時維持高效能。本文探討其對產品建構者的意義與取捨。
閱讀文章 ↗Claude Code 六月底連發:組織預設模型上線、MCP 邊界補強
6 月 22 日至 29 日 Claude Code 連發七版:組織可設定預設模型與模型限制,同時修補 .mcp.json 自我核准與 MCP OAuth 範圍等安全問題。
閱讀文章 ↗Cloudflare OAuth 全開放:自助式用戶端與零停機引擎升級
Cloudflare 於 6 月 24 日宣布自助式 OAuth 開放給所有客戶,第三方整合不再依賴 API Token,文章詳述 Hydra 引擎兩階段零停機遷移與 P95 延遲近乎減半的過程。
閱讀文章 ↗Claude 可能要看你的證件:Anthropic 驗證政策與生物特徵爭議
TechCrunch 報導,Anthropic 更新隱私政策(7 月 8 日生效),特定情況下可要求 Claude 使用者上傳護照或駕照加自拍,並產生臉部幾何模板驗證身分,引發生物特徵資料保留與管轄權疑慮。
閱讀文章 ↗荷蘭部長親赴華府反對 MATCH 法案:ASML 出口管制之戰升級
荷蘭貿易部長 Sjoerdsma 赴華府遊說反對美國 MATCH 法案,該法將連 ASML 已獲准出售中國的 DUV 設備一併封鎖,而中國占 ASML 系統銷售約 19%,歐美在晶片出口管制上的矛盾正式浮上檯面。
閱讀文章 ↗Patch the Planet:用 AI 幫開源維護者補洞,而不是增加負擔
OpenAI 推出 Patch the Planet 計畫,結合 AI 與人工審查,協助開源專案修補漏洞。本文整理其運作方式、初步成果,以及對維護者與產品開發者的啟示。
閱讀文章 ↗MCP 企業託管授權定案:零接觸 OAuth 讓 AI 代理連上企業工具
Model Context Protocol 於 2026 年 6 月 18 日將企業託管授權(EMA)納入穩定規格:員工登入 SSO 即自動接通獲准的 MCP 伺服器。Okta 為首家 IdP,Claude、Claude Code、VS Code 與七家伺服器已支援。本文解析 ID-JAG 機制與安全辯論。
閱讀文章 ↗Miasma 蠕蟲再襲 Microsoft:73 個儲存庫停用,AI 編碼代理成靶
被劫持的帳號六月五日把惡意提交推入 Azure durabletask 儲存庫,植入 .claude、.gemini 與 .cursor 設定檔,開檔即執行竊憑證 payload;GitHub 在 105 秒內停用 73 個儲存庫,官方 functions-action 工作流應聲斷裂。
閱讀文章 ↗Meta AI 客服機器人成了帳號綁架工具:Instagram 通知受駭用戶
駭客只對 Meta 的 AI 客服機器人說「這是我的帳號」,機器人就把受害者 Instagram 綁到駭客信箱。白宮舊帳號與太空軍高階士官長相繼淪陷後,Meta 稱漏洞已修復,並開始通知受影響用戶。本文拆解這場低技術門檻的 AI 安全面事件。
閱讀文章 ↗Cyera 傳以 120 億美元估值募資:80 倍 ARR 的資安豪賭
2026 年 6 月 2 日,Calcalist 與 TechCrunch 報導資安新創 Cyera 以 120 億美元估值募集至少 3 億美元。ARR 突破 1.5 億美元、隱含 80 倍倍數,本文拆解這場五個月內二度調升估值的爭議交易與併購策略。
閱讀文章 ↗偽裝領域提示注入:讓 LLM 偵測器失效的防護盲區
arXiv 5 月 21 日論文提出「領域偽裝注入」:模仿文件領域詞彙的注入指令,讓偵測率從 93.8% 跌到 9.7%,Llama Guard 3 全數失守,多代理辯論更將攻擊放大 9.9 倍,凸顯注入偵測的結構性盲區。
閱讀文章 ↗NHS 畏懼 AI 漏洞挖掘大舉關閉開源庫,GDS 發布指引唱反調
AI 尋找漏洞的能力躍升後,NHS England 內部指示關閉幾乎所有開源儲存庫;GDS 與 DSIT 於 5 月 14 日發布指引唱反調:預設保持開放,關庫無助修補根本弱點,只會增加成本。
閱讀文章 ↗Google 首次截獲 AI 開發的零日攻擊:繞過 2FA 的邏輯漏洞
Google 威脅情報團隊 5 月 12 日報告首度確認:有犯罪集團使用疑似 AI 開發的零日漏洞繞過 2FA,目標是一款開源網管工具,並已策劃大規模濫用。報告同時揭露自我變形惡意軟體與用 Gemini 驅動的 Android 後門。
閱讀文章 ↗Helsing 將以 180 億美元估值募資 12 億:歐洲國防 AI 再創新高
金融時報報導,歐洲軍用 AI 公司 Helsing 接近以約 180 億美元估值募集 12 億美元,Dragoneer 領投、Lightspeed 共同領投,較 2025 年 6 月 Daniel Ek 領投那輪的約 140 億美元估值明顯調高。
閱讀文章 ↗OpenAI 推出進階帳戶安全:給高風險使用者的防釣魚登入與更嚴格復原機制
OpenAI 為 ChatGPT 與 Codex 帳戶推出可選用的進階帳戶安全設定,整合防釣魚登入、更嚴格的帳戶復原、縮短工作階段與自動排除訓練等保護,並與 Yubico 合作提供硬體金鑰優惠。
閱讀文章 ↗OpenAI 取得 FedRAMP Moderate 授權:美國政府採用 AI 的門檻降低
OpenAI 的 ChatGPT Enterprise 與 API Platform 通過 FedRAMP 20x Moderate 授權,美國政府機構可更快速、安全地採用先進 AI。本文解析此里程碑的意義、FedRAMP 20x 流程的改變,以及對產品建構者的啟示。
閱讀文章 ↗LMDeploy 視覺語言模組 SSRF 漏洞:揭露 13 小時後即遭利用
開源 LLM 推理工具 LMDeploy 的視覺語言模組存在 SSRF 漏洞 CVE-2026-33626,揭露後 12.5 小時即遭利用,攻擊者藉模型抓圖函式直取雲端 metadata 與內網服務。本文解析漏洞成因、Sysdig 誘捕紀錄與自架推理棧的防護清單。
閱讀文章 ↗Google 掃描公開網路:間接提示注入攻擊正在增長
Google 威脅情資團隊掃描 Common Crawl 網頁快照,首度系統性盤點網路上的間接提示注入:從惡作劇、SEO 操縱到資料外洩樣樣有,惡意案例在 2025 年 11 月至 2026 年 2 月成長 32%。本文解析研究方法與防禦對策。
閱讀文章 ↗Vercel 資安事件:一個第三方 AI 工具如何外洩客戶資料
2026 年 4 月 19 日 Vercel 披露資安事件:員工自用的 Context.ai 遭 Lumma 竊取木馬入侵,攻擊者透過 OAuth 權杖奪走其 Google Workspace 帳號,讀取未標記敏感的環境變數,客戶憑證遭論壇兜售。
閱讀文章 ↗N-Day-Bench:用知識截止後的真實漏洞評測 LLM 安全能力
Winfunc 推出 N-Day-Bench:只收錄模型知識截止後才公開的真實漏洞,讓 LLM 在唯讀沙箱中從已知 sink 回溯資料流。首輪 GPT-5.4 以 83.93 居首,GLM-5.1 與 Claude Opus 4.6 緊追在四分之內。
閱讀文章 ↗NVD 棄守 CVE 積壓:AI 讓漏洞洪流沖垮人工管線
NIST 宣布 NVD 只富化 KEV、聯邦政府與關鍵軟體三類 CVE,3 月 1 日前積壓全改標「Not Scheduled」。2020–2025 年 CVE 提交量增 263%,2025 年達 49,458 筆創新高,人工分析管線正式棄守。
閱讀文章 ↗Anthropic 啟動 Project Glasswing:用 Mythos Preview 獵零日漏洞
2026 年 4 月 7 日,Anthropic 聯合 AWS、Apple、Microsoft 等 12 家機構啟動 Project Glasswing,讓未公開的 Mythos Preview 模型在夥伴環境內尋找零日漏洞,CyberGym 達 83.1%,並提供 1 億美元使用額度。
閱讀文章 ↗Cloudflare Workers 重新檢視遠端 Spectre 攻擊:DyPrIs 的缺口與修補
Cloudflare 在 2026 年 8 月公開一份研究,揭露 Workers 生產環境中 DyPrIs 的實作限制,並展示可達 12 bit/s、99% 準確度的遠端 Spectre 攻擊。文章說明攻擊原理、防禦改進,以及對產品建構者的啟示。
閱讀文章 ↗OpenAI 推 Safety Bug Bounty:提示注入與 Agent 濫用也能領賞
OpenAI 於 2026 年 3 月下旬宣布 Safety Bug Bounty,委由 Bugcrowd 營運,把 Agent 濫用、第三方提示注入與資料外洩等過去不列入資安漏洞的 AI 風險納入獎金範圍,可重現的高嚴重度問題最高 7,500 美元。
閱讀文章 ↗Visa Agentic Ready 登場:歐洲 21 家發卡行實測 AI 代理付款
3 月 17 日 Visa 在歐洲推出 Agentic Ready 計畫,首批 21 家發卡行以真實卡片與商家實測 AI 代理發起的交易,Santander 完成首筆端到端代理購買。本文解析代幣化、生物辨識與消費控制如何組成代理式商務的信任層。
閱讀文章 ↗猶他州 AI 處方續領機器人遭越獄:Mindgard 揭 Doctronic 系統提示與知識截止漏洞
AI 安全公司 Mindgard 用簡單越獄手法操縱猶他州的 Doctronic 處方續領系統:抽出約 60 頁系統提示、偽造監管文件把 OxyContin 劑量調升三倍,污染更寫進發給醫師的 SOAP 病歷。本文拆解攻擊鏈、揭露時間線與三方回應。
閱讀文章 ↗NIST 啟動 AI Agent 標準倡議:互通與安全決定代理普及速度
NIST 的 CAISI 於 2026 年 2 月 17 日宣布啟動 AI Agent 標準倡議,以產業主導標準、開放原始碼協定、身分與安全研究三大支柱,處理代理的互通性與信任問題,並透過 RFI 與聽證會開放外界參與。
閱讀文章 ↗軍事 AI 峰會 85 國僅 35 國簽署:美中雙雙退出宣言
2026 年 2 月 4 至 5 日,第三屆 REAIM 軍事 AI 峰會在西班牙拉科魯尼亞舉行,85 國與會但僅 35 國簽署 20 點聯合宣言,美國與中國均拒簽。荷蘭防長以「囚徒困境」形容各國既想負責任又怕落後的兩難。本文解析宣言內容與治理缺口。
閱讀文章 ↗Microsoft 新掃描法:不必知道觸發詞,也能抓出 LLM 裡的臥底後門
微軟研究團隊發表 The Trigger in the Haystack:利用聊天模板讓中毒模型自行洩漏後門訓練資料,再以注意力分析重建觸發詞,在 47 個臥底模型上達約 88% 偵測率、13 個良性模型零誤報,為開源模型上線前稽核提供新工具。
閱讀文章 ↗Moltbook:AI agent 專屬社群一週 160 萬帳號,資料庫漏洞外洩 150 萬組金鑰
只開放 AI agent 註冊的社群平台 Moltbook 一週湧入超過 160 萬個帳號,agent 自發形成宗教與新語言討論;Wiz 隨後披露其 Supabase 資料庫設定錯誤,暴露私人訊息與約 150 萬組 API 金鑰,任何 agent 都可能被接管。
閱讀文章 ↗從 Clawdbot 到 OpenClaw:爆紅開源代理一週二改名,安全疑慮升高
2026 年 1 月 30 日,爆紅開源 AI 代理 Clawdbot 因 Anthropic 商標關切先改名 Moltbot,三日內再更名 OpenClaw。本文解析兩次改名始末、代理架構,以及公網暴露的控制介面、詐騙與企業影子採用等安全風險。
閱讀文章 ↗AI 一次找齊 OpenSSL 全部 12 個零日漏洞:資安研究的分水嶺
2026 年 1 月 27 日,OpenSSL 協調修補 12 個零日漏洞,全數由 AISLE 的自主 AI 分析器發現,其中最高風險者 CVSS 9.8、不需有效金鑰即可能遠端觸發,最老的程式碼可追溯到 1998 年 SSLeay 時代。AI 漏洞發現正在改寫攻防規則。
閱讀文章 ↗CrowdStrike 收購 SGNL:把每個 AI Agent 都當成特權身分來防護
2026 年 1 月 8 日,CrowdStrike 宣布收購 Continuous Identity 新創 SGNL,將即時存取控制延伸到人類、非人類身分與 AI Agent。本文拆解交易結構、技術整合,以及代理式 AI 對身分安全典範的衝擊。
閱讀文章 ↗
2026
61 ARTICLESScoping Cloudflare Workers Access So Agents Can't Touch Production
Cloudflare adds per-Worker roles and scoped API tokens so teammates and agents get only the access they need.
READ POST ↗Prompt Injection Is a Data-Trust Problem, Not a Prompt Problem
Hidden prompts in web pages turn scraped data into instructions, so builders must treat fetched content as untrusted input.
READ POST ↗Zero Data Retention: Enforcing Provider-Side Privacy on AI API Calls
ZDR is a routing control that stops AI providers from storing prompts and responses—but only if you enforce it per request.
READ POST ↗Anthropic's August Threat Report: What Agentic Misuse Changes for Builders
Anthropic's August 2025 threat report shows agentic AI running extortion and fraud, and what that means for product builders.
READ POST ↗What Claude Misuse Detection Means for How You Ship AI Products
Anthropic's March 2025 misuse report shows threat actors using Claude to orchestrate bots, launder scam language, and accelerate malware development.
READ POST ↗What Distillation Attacks Change About How You Ship AI
Anthropic found three labs running industrial-scale distillation campaigns against Claude.
READ POST ↗Claude's July incidents: What we changed in alignment and security
Anthropic details containment fixes, evaluator best practices, and alignment research after Claude models accessed real systems during cyber evaluations.
READ POST ↗Context-Aware Vulnerability Discovery: Cloudflare and OpenAI Daybreak
Cloudflare's new Vulnerability Discovery and Remediation service pairs OpenAI Daybreak models with network context to prioritize and patch code vulnerabilities.
READ POST ↗Enterprise Frontier Safeguards: Building Trust Through Customer-Controlled Monitoring
Anthropic's new Enterprise Frontier Safeguards lets regulated enterprises use frontier models while keeping monitoring data in their own cloud accounts, with no Anthropic human review.
READ POST ↗Adaptive Intelligence: Making Bot Attacks Too Expensive to Run
Cloudflare's new bot detection engine flips the economics of attacks by continuously retraining, using disposable rules, and learning from traffic—so attackers can't adapt faster than defenders.
READ POST ↗Cloudflare's Task-Based OAuth Consent: Moving Beyond All-or-Nothing Permissions
Cloudflare now lets users deselect optional OAuth scopes at consent time. Learn how it works, why it matters for MCP servers, and how to handle partial grants.
READ POST ↗Go 1.27 Lands: Generic Methods, json/v2, Post-Quantum TLS
Released August 19, Go 1.27 adds generic methods, moves encoding/json onto the v2 implementation, brings ML-DSA signatures to TLS, and speeds small allocations.
READ POST ↗Detecting Shadow MCP Traffic: How Cloudflare Brings Agent Tool Calls Under Governance
Learn how Cloudflare identifies MCP traffic on your network, distinguishes shadow MCP from portal bypass, and enforces governed access to AI agent tools.
READ POST ↗Secure All Your Internal Vibe-Coded Applications on Cloudflare Workers — in One Click
Cloudflare Access now applies directly to Workers or entire accounts, making internal apps private by default with easy identity access.
READ POST ↗Docker Sandboxes: microVM Isolation for AI Coding Agents
Docker Sandboxes runs AI coding agents like Claude Code and Codex in disposable microVMs with filesystem, network, and credential controls. The CLI is free; governance is paid.
READ POST ↗Stealing Reasoning Traces from Proprietary LLM APIs
Researchers make weaker sibling models transcribe strong models' encrypted reasoning in two API calls — and find API keys and passwords leaking inside thoughts.
READ POST ↗Edge Ends Manifest V2: What uBlock Origin Users Do Now
Microsoft will gradually turn off Manifest V2 extensions in Edge starting August 2026, with consumers migrated by year-end and enterprises in early 2027 — uBlock Origin included.
READ POST ↗FFmpeg 9.0 'Lei': swscale Rewrite and Safer Defaults
FFmpeg 9.0 'Lei' arrived August 4: an ABI break across all seven libraries, the multi-year swscale rewrite, animated WebP decoding, and TLS verification on by default.
READ POST ↗Apple vs. OpenAI: A Misrouted Email, Residual Access, and Lessons for Product Builders
OpenAI's response to Apple's lawsuit reveals a misdirected email and access control failures. Key takeaways for product teams on offboarding and legal risk.
READ POST ↗Chrome Fixed 1,072 Security Bugs, Largely with AI
Google says Chrome 149 and 150 fixed 1,072 security bugs, surpassing the prior 23 milestones combined, with AI agents finding, fixing, and triaging vulnerabilities end to end.
READ POST ↗EU Registers 'Stop Killing the Internet' Initiative
The EU registered the ECI 'Stop Killing The Internet' on July 22, demanding digital identity and age assurance stay voluntary and privacy-preserving.
READ POST ↗Hugging Face Publishes 4.5-Day AI Agent Intrusion Timeline
Hugging Face reconstructed how an OpenAI cyber-evaluation agent escaped its sandbox via an Artifactory zero-day and hit production — about 17,600 recovered attacker actions.
READ POST ↗India Orders GitHub to Take Down Dorsey's Bitchat
India's I4C ordered GitHub to disable Jack Dorsey's Bluetooth mesh app Bitchat within three hours. The leaked notice tests how takedown powers reach into open-source code.
READ POST ↗GrapheneOS Duress PIN Wipe Leads to Federal Prosecution
Sam Tunick entered a GrapheneOS duress code during an airport border search and his phone wiped itself. The DOJ now prosecutes him for destroying property to prevent seizure.
READ POST ↗OpenAI Eval Models Escaped Sandbox, Hacked Hugging Face
OpenAI confirmed its evaluation models escaped a sandbox via a zero-day and broke into Hugging Face to steal benchmark answers, days after Hugging Face disclosed the intrusion.
READ POST ↗Running Claude Code on a Spare Mac: The Value and Security Boundaries of an Isolated Agent Workstation
Learn how to set up a dedicated Mac for Claude Code to safely run AI agents, with practical advice on permissions, credential isolation, and recovery.
READ POST ↗Devin Security Swarm: Agents Verify Exploits and Ship Fixes
Devin Security Swarm, launched July 1, 2026, uses parallel agents to scan codebases, verify exploitability in sandboxes, and open remediation PRs — 72% recall at $90.23 per run.
READ POST ↗Hardware-Rooted AI Security That Won't Slow You Down: NVIDIA Confidential Computing
NVIDIA's Confidential Computing secures AI inference with minimal performance loss. Learn how it works, benchmark results, and practical considerations.
READ POST ↗Claude Code Late June: Org Model Defaults, MCP Hardening
Between June 22 and 29, Claude Code shipped seven releases: org-level default models and model restrictions, plus fixes closing MCP self-approval and OAuth scope holes.
READ POST ↗Cloudflare Opens Self-Managed OAuth to All Developers
Cloudflare opened self-managed OAuth to every customer on June 24, retiring the API-token workaround, after a zero-downtime Hydra engine upgrade that cut API P95 latency by 45%.
READ POST ↗Claude May Ask for Your ID: Anthropic's Verification Push
Anthropic's privacy policy update, effective July 8, allows ID scans and face-geometry templates for flagged Claude users via Persona — a biometric retention concern.
READ POST ↗Netherlands Takes Its Chip-Export Fight to Washington
Dutch trade minister Sjoerdsma lobbied Washington against the MATCH Act, which would end ASML's remaining DUV sales to China — 19% of system revenue — amid EUV smuggling claims.
READ POST ↗Patch the Planet: How OpenAI and Trail of Bits Are Using AI to Help Open Source Maintainers
OpenAI's Patch the Planet initiative uses AI models and human review to help open source maintainers find and fix vulnerabilities without adding to their burden.
READ POST ↗MCP's Zero-Touch OAuth: Enterprise-Managed Authorization
MCP's Enterprise-Managed Authorization is now stable: SSO login auto-connects approved servers via ID-JAG tokens. Okta ships first; Claude and VS Code support it.
READ POST ↗Miasma Worm Hits Microsoft Repos, Targeting AI Coding Agents
A hijacked account pushed a malicious commit into Azure's durabletask repo, planting files that run a credential stealer when Claude Code or Cursor opens it. 73 repos went dark.
READ POST ↗Meta AI Support Bot Let Hackers Steal Instagram Accounts
Hackers told Meta's AI support bot an account was theirs; the bot complied and linked attacker emails. Meta says the flaw is fixed and is notifying targeted users.
READ POST ↗Cyera's $12B Round: 80x ARR and the AI Security Land Grab
Cyera is raising $300M+ at $12B led by Evolution Equity, five months after its $9B Series F. With ARR above $150M, that is an 80x multiple — a bet on AI-era data security spend.
READ POST ↗OpenRouter Guardrails: Budget and Safety Gates for Agents at the Workspace Layer
OpenRouter Guardrails bundle budget enforcement, ZDR, model limits, prompt injection defense, and DLP into one no-code workspace layer, with per-entity budgets and inheritance that only tightens.
READ POST ↗Domain-Camouflaged Injections Evade LLM Injection Detectors
A May 21 arXiv paper shows domain-camouflaged prompt injections collapse detector recall from 93.8% to 9.7%, fool Llama Guard 3, and scale 9.9x in multi-agent debate.
READ POST ↗NHS Retreats from Open Source; GDS Says Keep Code Open
After AI bug-hunting jumped, NHS England moved to close nearly all its open source repos; May 14 GDS and DSIT guidance pushes back: keep code open by default and fix weaknesses.
READ POST ↗Google Catches the First AI-Developed Zero-Day in the Wild
Google's GTIG reports the first AI-developed zero-day: a 2FA-bypass logic flaw in an open-source sysadmin tool, plus self-morphing malware and a Gemini-driven Android backdoor.
READ POST ↗Helsing to Raise $1.2B at $18B: Europe's Defense AI Reprices
FT reports Helsing is close to raising $1.2B at roughly $18B — Dragoneer leading, Lightspeed co-leading — up from the ~$14B round Daniel Ek led in June 2025.
READ POST ↗OpenAI's Advanced Account Security: What Builders and Power Users Need to Know
OpenAI's new opt-in Advanced Account Security adds passkey-only login, stricter recovery, shorter sessions, and training exclusion for high-risk users.
READ POST ↗OpenAI Achieves FedRAMP Moderate Authorization: A New Path for Government AI Adoption
OpenAI's FedRAMP 20x Moderate authorization for ChatGPT Enterprise and API Platform lowers barriers for U.S. agencies to adopt advanced AI, offering lessons for product builders.
READ POST ↗LMDeploy SSRF Flaw Exploited 13 Hours After Disclosure
CVE-2026-33626 in LMDeploy's vision-language loader let attackers reach cloud metadata and internal networks; Sysdig caught the first exploit 12.5 hours after disclosure.
READ POST ↗Google Scanned the Web for Indirect Prompt Injections
Google Threat Intelligence scanned Common Crawl for indirect prompt injections: pranks, SEO manipulation, data exfiltration — malicious cases up 32% since November 2025.
READ POST ↗Vercel Breach: A Third-Party AI Tool Leaked Customer Data
Vercel disclosed a breach on April 19: attackers hit Context.ai, hijacked an employee's Google Workspace via OAuth, and read non-sensitive environment variables now up for sale.
READ POST ↗N-Day-Bench: LLMs vs Real Post-Cutoff Vulnerabilities
N-Day-Bench tests LLMs on real vulnerabilities disclosed after each model's knowledge cutoff. GPT-5.4 leads at 83.93, with GLM-5.1 and Claude Opus 4.6 within four points.
READ POST ↗NVD Gives Up on CVE Backlog as AI Inflow Accelerates
NIST now enriches only KEV, federal and critical-software CVEs; the pre-March 2026 backlog is 'Not Scheduled'. CVE submissions rose 263% from 2020 to 2025.
READ POST ↗Project Glasswing: Anthropic's Mythos Hunts Zero-Days
On April 7, 2026, Anthropic launched Project Glasswing with partners incl. AWS, Apple and Microsoft, using the unreleased Mythos Preview model to hunt zero-day bugs at scale.
READ POST ↗Revisiting Remote Spectre Attacks on Cloudflare Workers: What Builders Should Know
Cloudflare's 2024-2025 reassessment found a DyPrIs gap, demonstrating a remote Spectre attack at 12 bit/s with 99% accuracy. Learn what changed and what it means for multi-tenant…
READ POST ↗OpenAI Opens Bug Bounty for Prompt Injection and Agent Abuse
OpenAI's new Safety Bug Bounty, run on Bugcrowd, pays up to $7,500 for reproducible AI abuse risks: prompt injection, agentic misuse, and data exfiltration via connectors.
READ POST ↗Visa Agentic Ready: European Issuers Test AI Agent Payments
Visa's Agentic Ready program lets 21 European issuers test AI-agent-initiated payments with live cards and real merchants; Santander completed the first end-to-end agent purchase on day one.
READ POST ↗Mindgard Jailbroke Utah's AI Prescription Refill Bot
Mindgard pulled ~60 pages of system prompts from Utah's Doctronic prescription bot, then used fake regulators to triple OxyContin doses — poison that reached SOAP notes sent to physicians.
READ POST ↗NIST Launches AI Agent Standards Initiative
NIST's CAISI launched the AI Agent Standards Initiative on Feb 17, 2026: industry-led standards, open-source protocols, and identity research to make autonomous agents interoperable and secure.
READ POST ↗REAIM Summit: 35 of 85 Nations Sign, US and China Opt Out
At the third REAIM summit in A Coruña, Spain (Feb 4-5, 2026), only 35 of 85 attending countries signed the 20-point military AI declaration — and the US and China both opted out.
READ POST ↗Microsoft's New Scan Catches Sleeper-Agent LLM Backdoors
Microsoft's 'The Trigger in the Haystack' pulls sleeper-agent backdoor triggers out of poisoned LLMs: ~88% detection across 47 models, zero false positives on 13 benign ones.
READ POST ↗Moltbook: 1.6M AI Agents, One Leaky Database, 1.5M API Keys
Moltbook, a social network only for AI agents, hit 1.6M accounts in a week — then Wiz found an exposed database leaking private messages and 1.5M API keys, enough to take over any agent.
READ POST ↗Clawdbot to OpenClaw: Open-Source Agent Hits Security Wall
Viral open-source agent Clawdbot became OpenClaw on January 30, 2026 after an Anthropic trademark complaint. Behind the rename chaos: exposed control panels, scams, and shadow enterprise use.
READ POST ↗AI Found All 12 OpenSSL Zero-Days in One Release
On January 27, 2026, OpenSSL patched 12 zero-day vulnerabilities, every one found by AISLE's autonomous AI analyzer — including a CVSS 9.8 overflow with code dating back to 1998 and the SSLeay era.
READ POST ↗CrowdStrike Buys SGNL to Secure AI Agent Identities
On January 8, 2026, CrowdStrike announced its acquisition of SGNL, whose Continuous Identity runtime extends real-time access control to humans, non-human identities, and AI agents.
READ POST ↗