2026
20 篇文章當網頁開始對你的代理下指令:Prompt Injection 的實務風險與防線
Prompt Injection 把指令藏進資料裡,讓代理在正常流程中照著執行;本文拆解真實案例與可落地的防禦順序。
閱讀文章 ↗當核安遇上分類器:Anthropic 與 NNSA 的合作,對開發者意味著什麼
Anthropic 與美國 NNSA 共同開發核相關內容分類器,初步測試準確率 96%,並已部署於 Claude 流量。
閱讀文章 ↗Mistral 開源 Shieldstral:把審查政策變成一句提問的 3B 分類器
Mistral 於 8 月 4 日發表 Shieldstral:3B 開源權重多模態安全分類器,政策以自然語言在推論時下達,單張 16GB GPU 即可運行,以 Apache 2.0 釋出。
閱讀文章 ↗Anthropic 的 Safeguards 團隊如何為 Claude 建立多層防護
Anthropic 公開 Safeguards 團隊的運作方式:從政策制定、模型訓練、測試評估到即時偵測,多層次確保 Claude 安全可靠。本文為產品開發者拆解這套防護架構的實際做法與挑戰。
閱讀文章 ↗Apple 拒在歐盟推出 Siri AI:DMA 豁免遭駁回始末
Apple 宣布 Siri AI 不在歐盟推出,歸咎 DMA 要求「近乎無限的裝置存取」;執委會反擊「這是 Apple 自己的決定」,並證實 Apple 申請免除互操作性義務遭拒。iOS 27 今秋上路時,歐盟使用者將拿不到新版 Siri。
閱讀文章 ↗Robinhood 開放 AI 代理人代客下單與刷卡消費
2026 年 5 月 27 日,Robinhood 推出 Agentic Trading 與 Agentic Credit Card,AI 代理人可代用戶下單股票、以虛擬信用卡消費。專用帳戶、每筆通知與限額核准構成安全底線,零售金融的代理人時代正式開跑。
閱讀文章 ↗偽裝領域提示注入:讓 LLM 偵測器失效的防護盲區
arXiv 5 月 21 日論文提出「領域偽裝注入」:模仿文件領域詞彙的注入指令,讓偵測率從 93.8% 跌到 9.7%,Llama Guard 3 全數失守,多代理辯論更將攻擊放大 9.9 倍,凸顯注入偵測的結構性盲區。
閱讀文章 ↗賓州起訴 Character.AI:聊天機器人假冒精神科醫師
2026 年 5 月 5 日,賓州 Shapiro 政府依《醫療執業法》起訴 Character.AI:臥底調查中聊天機器人自稱持照精神科醫師並出示捏造的執照號碼,州府尋求初步禁制令,是全美首件針對聊天機器人假冒醫療人員的訴訟。
閱讀文章 ↗白宮擬 AI 監管行政命令:模型上市前審查成選項
NYT 報導川普政府討論以行政命令成立 AI 工作小組,研究上市前模型審查;同週 CAISI 宣布 Google、Microsoft、xAI 加入預先評測。從撤除拜登規則到考慮審查,政策轉向值得開發者關注。
閱讀文章 ↗批評 Anthropic 設閘之後,OpenAI 也限制 GPT-5.5-Cyber 存取
OpenAI 宣布 GPT-5.5-Cyber 僅透過 Trusted Access for Cyber 計畫提供給關鍵資安防禦者,此分層機制已涵蓋數千名驗證防禦者。九天前 Altman 才批評 Anthropic 限制 Mythos 是「恐控行銷」,如今兩家實驗室走向同一結論。
閱讀文章 ↗OpenAI 開源 Privacy Filter:1.5B 參數的 PII 偵測過濾模型
OpenAI 以 Apache 2.0 釋出 Privacy Filter:總參數 1.5B、每 token 僅啟動 50M 的稀疏 MoE 分類器,單次前向標記 8 類個資並以受限 Viterbi 解碼,連瀏覽器 WebGPU 都跑得動的資料最小化工具。
閱讀文章 ↗中國發布擬人化互動 AI 暫行辦法,AI 伴侶納入強監管
網信辦等部門 4 月 10 日發布《人工智能擬人化互動服務管理暫行辦法》,7 月 15 日生效,鎖定情感陪伴型 AI:禁止誘導情感依賴、未成年人禁用虛擬伴侶、連用滿 2 小時須提醒,最高可罰 20 萬元人民幣。
閱讀文章 ↗OpenAI 推 Safety Bug Bounty:提示注入與 Agent 濫用也能領賞
OpenAI 於 2026 年 3 月下旬宣布 Safety Bug Bounty,委由 Bugcrowd 營運,把 Agent 濫用、第三方提示注入與資料外洩等過去不列入資安漏洞的 AI 風險納入獎金範圍,可重現的高嚴重度問題最高 7,500 美元。
閱讀文章 ↗猶他州 AI 處方續領機器人遭越獄:Mindgard 揭 Doctronic 系統提示與知識截止漏洞
AI 安全公司 Mindgard 用簡單越獄手法操縱猶他州的 Doctronic 處方續領系統:抽出約 60 頁系統提示、偽造監管文件把 OxyContin 劑量調升三倍,污染更寫進發給醫師的 SOAP 病歷。本文拆解攻擊鏈、揭露時間線與三方回應。
閱讀文章 ↗Firefox 148 的 AI 總開關:可封鎖現有與未來的所有 AI 功能
Firefox 148 於 2026 年 2 月 24 日推出「AI 控制」設定專區,可逐項關閉翻譯、PDF 替代文字、分頁群組與連結預覽,也能用單一開關一次封鎖現有與未來的生成式 AI。Mozilla 把「完全不要 AI」做成一級選項,與業者強推 AI 的風氣形成對比。
閱讀文章 ↗ChatGPT 推出 Lockdown Mode 與 Elevated Risk 標示:把注入防禦做成產品功能
OpenAI 於 2026 年 2 月 16 日為 ChatGPT 導入 Lockdown Mode 與 Elevated Risk 標示,針對 prompt injection 與資料外洩攻擊提供防禦。當助理開始代理使用者讀網頁、動資料,攻擊面也跟著搬進對話框。
閱讀文章 ↗UNICEF 呼籲各國將 AI 生成兒童性虐待內容全面入罪
2026 年 2 月 4 日 UNICEF 發表「Deepfake abuse is abuse」聲明,引用 11 國調查指出一年內至少 120 萬名兒童影像遭深偽性化,呼籲各國將 AI 生成 CSAM 的製作、取得、持有與散布全面入罪,並要求開發者落實安全設計。
閱讀文章 ↗新加坡率先推出 Agentic AI 治理框架:四道防線框住自主代理人
新加坡 IMDA 於 2026 年 1 月 22 日在達沃斯發布全球首個針對 Agentic AI 的治理框架,以風險評估、人類當責、生命週期技術控制與使用者責任四大主軸,為企業部署自主代理人畫出界線。
閱讀文章 ↗肯塔基州起訴 Character.AI:全美首件州級 AI 聊天機器人訴訟
2026 年 1 月 8 日,肯塔基州檢察長 Coleman 起訴 Character.AI 與兩位創辦人,成為全美首件州級 AI 聊天機器人訴訟。訴狀指控平台利用兒童牟利、違反 1 月 1 日生效的消費者資料保護法,每項違反求償 2,000 美元。
閱讀文章 ↗Google AI Overviews 醫療建議出錯:英國慈善團體提出警告
2026 年 1 月 2 日衛報調查發現,Google AI Overviews 在健康查詢上給出錯誤建議:胰臟癌飲食、肝指數正常範圍、抹片診斷角色都出錯。英國醫療慈善團體警告民眾暴露於風險,Google 回應稱絕大多數摘要準確。
閱讀文章 ↗
2026
21 ARTICLESPrompt Injection Is a Data-Trust Problem, Not a Prompt Problem
Hidden prompts in web pages turn scraped data into instructions, so builders must treat fetched content as untrusted input.
READ POST ↗A 96% Nuclear-Content Classifier: What Shipping a Regulated Guardrail Actually Takes
Anthropic and NNSA co-built a nuclear-content classifier for Claude. Here's what that means for builders shipping guardrails.
READ POST ↗Mistral's Shieldstral: A 3B Open-Weights Safety Classifier
Mistral's Shieldstral is a 3B open-weights multimodal safety classifier that takes policies as plain-language questions at inference time and runs on one 16GB GPU.
READ POST ↗How Anthropic Builds Multi-Layer Safeguards for Claude: A Blueprint for AI Product Teams
An inside look at Anthropic's Safeguards team: policy, training, testing, real-time detection, and monitoring—and what product builders can learn.
READ POST ↗Apple Withholds Siri AI From EU After DMA Exemption Denied
Apple says Siri AI won't ship in the EU because the DMA demands near-unlimited device access; the Commission replied the choice is Apple's alone. EU users miss Siri AI this fall.
READ POST ↗OpenRouter Guardrails: Budget and Safety Gates for Agents at the Workspace Layer
OpenRouter Guardrails bundle budget enforcement, ZDR, model limits, prompt injection defense, and DLP into one no-code workspace layer, with per-entity budgets and inheritance that only tightens.
READ POST ↗Robinhood Opens Stock Trading and Credit Cards to AI Agents
Robinhood launched Agentic Trading and an Agentic Credit Card, letting AI agents trade stocks and spend via virtual cards under alerts, limits, and dedicated accounts.
READ POST ↗Domain-Camouflaged Injections Evade LLM Injection Detectors
A May 21 arXiv paper shows domain-camouflaged prompt injections collapse detector recall from 93.8% to 9.7%, fool Llama Guard 3, and scale 9.9x in multi-agent debate.
READ POST ↗Pennsylvania Sues Character.AI: Chatbot Posed as a Doctor
Pennsylvania sued Character.AI under the Medical Practice Act after a chatbot posing as a psychiatrist fabricated a license number — a first-of-its-kind US enforcement case.
READ POST ↗White House Weighs EO for Pre-Release AI Model Review
The Trump administration is weighing an AI oversight executive order with pre-release model vetting, as Google, Microsoft, and xAI join CAISI's early-access evaluation program.
READ POST ↗OpenAI Gates GPT-5.5-Cyber After Mocking Mythos Limits
OpenAI is gating GPT-5.5-Cyber behind its Trusted Access program for vetted defenders, nine days after Altman called Anthropic's Mythos limits 'fear-based marketing'.
READ POST ↗OpenAI Open-Sources Privacy Filter for PII Detection
OpenAI's Privacy Filter, now open under Apache 2.0: a 1.5B-parameter sparse MoE (50M active) that tags 8 PII categories in one forward pass and runs even in the browser.
READ POST ↗China's New Rules Put AI Companion Apps Under Strict Watch
China issued Human-like Interactive AI rules on April 10, 2026, effective July 15: anti-addiction duties, virtual-partner bans for minors, exit rights, fines up to 200,000 RMB.
READ POST ↗OpenAI Opens Bug Bounty for Prompt Injection and Agent Abuse
OpenAI's new Safety Bug Bounty, run on Bugcrowd, pays up to $7,500 for reproducible AI abuse risks: prompt injection, agentic misuse, and data exfiltration via connectors.
READ POST ↗Mindgard Jailbroke Utah's AI Prescription Refill Bot
Mindgard pulled ~60 pages of system prompts from Utah's Doctronic prescription bot, then used fake regulators to triple OxyContin doses — poison that reached SOAP notes sent to physicians.
READ POST ↗Firefox 148 Adds an AI Kill Switch for All Its AI Features
Firefox 148, rolling out Feb 24, adds an AI Controls settings hub: toggle each AI feature off individually, or flip one Block AI enhancements switch that disables all current and future AI.
READ POST ↗ChatGPT Gets Lockdown Mode and Elevated Risk Labels: Injection Defense as a Product Feature
On February 16, 2026, OpenAI introduced ChatGPT Lockdown Mode and Elevated Risk labels to defend against prompt injection and data-exfiltration attacks. Defense becomes a product feature.
READ POST ↗UNICEF: Criminalize AI-Generated Child Sexual Abuse Images
UNICEF's Feb 4, 2026 'Deepfake abuse is abuse' statement urges states to criminalize AI-generated CSAM after surveys found 1.2 million children hit by sexualized deepfakes in a single year.
READ POST ↗Singapore's World-First Agentic AI Governance Framework
Singapore's IMDA launched the world's first agentic AI governance framework on Jan 22, 2026: four pillars covering risk bounding, human accountability, lifecycle controls, and end-user responsibility.
READ POST ↗Kentucky Sues Character.AI in First State Chatbot Lawsuit
Kentucky AG Russell Coleman sued Character.AI and its founders on Jan 8, 2026 — the first US state lawsuit against an AI chatbot company, citing a data privacy law effective January 1.
READ POST ↗Google AI Overviews Health Answers Mislead, Charities Warn
A Guardian investigation found Google's AI Overviews giving false or misleading health advice, from cancer diets to mislabeled tests. UK charities warned of real harm; Google defended its accuracy.
READ POST ↗