2026
52 篇文章模型不聽話時,該怎麼記錄、調查、公開?OpenAI 的 misalignment 通報框架
OpenAI 提出一套模型偏差通報框架,從發現、調查到揭露,讓開發者能更快分享異常行為,即使還沒找到原因或解法。
閱讀文章 ↗Claude for Teachers:把教學標準和備課時間,還給老師
Anthropic 推出 Claude for Teachers,為美國 K-12 教師提供免費的 Claude 高階功能、教學技能庫,並串接全美 50 州的學術標準與實證課程。本文整理產品重點、資料隱私設計,以及對教育科技建構者的啟示。
閱讀文章 ↗Anthropic 推出 2 億美元研究基金,探索 AI 經濟衝擊的應對方案
Anthropic 宣布投入 2 億美元成立 Economic Futures Research Fund,支持外部研究機構進行大規模實驗與試點,以了解哪些政策與計劃能讓經濟在 AI 衝擊下更具韌性,並讓更多人分享 AI 帶來的利益。
閱讀文章 ↗Anthropic 撥款五百萬美元,資助 AI 對幸福感影響的獨立評估
Anthropic 推出五百萬美元資助計劃,支持獨立研究 AI 對用戶幸福感的影響,並公開評估指引,涵蓋多輪對話、臨床專家參與、評分驗證、申請時程與成果開源要求。
閱讀文章 ↗AI 如何改變權力平衡:OpenAI 新團隊的長期思考
OpenAI 成立 Strategic Futures 團隊並推出 AI Futures 部落格,探討 AI 對自由社會權力結構的影響。文章從歷史、政治經濟學與制度設計角度,提出六項原則,並強調「權力平衡」而非「完全去中心化」才是關鍵。
閱讀文章 ↗Gemini Robotics 2 進場:從腳到指尖的機器人全身智能
2026 年 7 月 30 日 Google DeepMind 發表 Gemini Robotics 2,以 VLA、ER 與 On-Device 三個模型讓人形機器人全身控制、綁繩結、組隊合作,並公布 ASIMOV-Agentic 安全基準。
閱讀文章 ↗OpenAI 在歐洲的負責任 AI 實踐:從框架到落地
因應歐盟 AI 法下一階段,OpenAI 說明如何調整安全、透明與溯源機制:簽署 GPAI 與內容透明度行為準則、公開系統卡與 Model Spec,並以網路安全為例展示動態治理的實際做法。
閱讀文章 ↗青少年與 AI:OpenAI 如何在安全與學習之間取平衡
OpenAI 發布青少年 AI 使用方針:近九成青少年每週在 ChatGPT 上學習,官方主張「接取必須伴隨保護」。本文整理 Study Mode 的引導式學習設計、家長控制與高風險通知、休息提醒等防護機制,以及四項治理原則與外部合作。
閱讀文章 ↗Claude 可能要看你的證件:Anthropic 驗證政策與生物特徵爭議
TechCrunch 報導,Anthropic 更新隱私政策(7 月 8 日生效),特定情況下可要求 Claude 使用者上傳護照或駕照加自拍,並產生臉部幾何模板驗證身分,引發生物特徵資料保留與管轄權疑慮。
閱讀文章 ↗Midjourney Medical 全身超音波掃描儀:60 秒成像的豪賭與爭議
Midjourney 於 2026 年 6 月 18 日發表 Medical 全身超音波掃描儀:受測者浸入水槽,環形探頭 60 秒取滿全身資料,由 AI 重建成類 CT 影像。官方稱速度近百倍於 MRI,並引用「避免 30% 死亡」的說法,引發放射科醫師強烈質疑。
閱讀文章 ↗美國人點睇 AI?Anthropic 首份 Public Record 調查的 5 個關鍵發現
Anthropic 發布首份 Public Record 調查,訪問近 52,000 名美國人。結果顯示:疾病治療是最大期望、失業是最普遍擔憂、跨黨派支持監管,但僅 15% 信任 AI 公司。本文整理對產品建構者的啟示。
閱讀文章 ↗xAI 遭前工程師提告:警示 Grok 安全風險反被逼退
2026 年 6 月 10 日,前 xAI 工程師 Devin Kim 對 xAI 與 SpaceX 提告,稱因警示 Grok 的歧視與武器擴散風險遭報復逼退;他甫於 6 月初出任 Center for AI Safety 院長。
閱讀文章 ↗xAI 要求法院揭露 Grok 深偽原告身分:匿名訴訟權之戰
Grok 深偽集體訴訟中,xAI 聲請揭開四位匿名原告的本名,主張圖片將封存、揭露無妨;原告律師斥「剝奪衣服後還要剝奪假名」,四位原告揚言被迫具名就退出。本文解析這場匿名權攻防對隱私訴訟的影響。
閱讀文章 ↗OpenAI 的政治立場:不設 PAC、不捐款,強調透明與直接倡議
OpenAI 在 2026 年 6 月 1 日發布政策聲明,說明公司不設立政治行動委員會、不捐款給候選人或超級 PAC,並強調員工個人政治參與與公司立場分開。
閱讀文章 ↗「This is fine」作者與 Artisan 達成協議:迷因挪用之爭落幕
「This is fine」原作者 KC Green 指控 AI 新創 Artisan 在公車與地鐵廣告挪用其作品後,雙方於五月底達成協議:Artisan 下架紐約與舊金山廣告,Green 撤下抗議貼文。一起教科書級的 AI 行銷授權案例。
閱讀文章 ↗FTC 重罰「主動傾聽」AI 廣告:從未監聽的 93 萬美元和解
2026 年 5 月 21 日,FTC 宣布 Cox Media Group 等三家公司支付 93 萬美元和解:它們以「主動傾聽」為名販售 AI 廣告,宣稱能監聽對話投放廣告,實際上從未使用任何語音資料,只是加價轉售電子郵件名單。
閱讀文章 ↗arXiv 重拳整治 AI 垃圾論文:未查核內容者禁投一年
2026 年 5 月中,arXiv 宣布對提交明顯未查核 AI 內容的作者處以一年禁令,解禁後投稿須先獲同儕審查場域接受。本文解析鐵證標準、處分程序與對預印本生態的影響。
閱讀文章 ↗Musk 訴 OpenAI 案倒數:Altman 出庭自辯「騙子」指控
2026 年 5 月 12 日,Altman 在奧克蘭聯邦法院出庭作證,面對「偷走慈善機構」的主張與一長串「說謊者」名單的反詰問。Musk 求除去 Altman 與 Brockman 職務、重新分配 1,340 億美元。本文整理雙方證詞要點與結辯前的局勢。
閱讀文章 ↗微軟據報考慮放棄 24/7 潔淨能源全時匹配目標:AI 用電壓力浮上檯面
彭博報導微軟正評估延後或放棄 2030 年 24/7 潔淨能源全時匹配目標,AI 資料中心用電是關鍵壓力。本文解析該目標的技術意義、微軟官方回應,以及 hyperscaler 轉向天然氣對減排承諾的衝擊。
閱讀文章 ↗GPT-5.5 Instant 的系統卡透露了什麼:首次被列為高能力的 Instant 模型
OpenAI 在 2026 年 5 月 5 日發布 GPT-5.5 Instant 系統卡,這是首個被列為高能力等級的 Instant 模型,特別在網路安全和生化防護方面。本文為產品開發者解析這份文件的重點與含義。
閱讀文章 ↗哈佛急診研究:o1 診斷準確率勝過兩位主治醫師
哈佛醫學院團隊在《Science》發表研究:在 76 名急診患者的檢傷診斷上,o1 以 67% 的精確或接近精確診斷率勝過兩位主治醫師的 55% 與 50%。本文解析數據、批評與落地限制。
閱讀文章 ↗白宮擬 AI 監管行政命令:模型上市前審查成選項
NYT 報導川普政府討論以行政命令成立 AI 工作小組,研究上市前模型審查;同週 CAISI 宣布 Google、Microsoft、xAI 加入預先評測。從撤除拜登規則到考慮審查,政策轉向值得開發者關注。
閱讀文章 ↗Google 開放五角大廈機密網路使用 AI:Anthropic 拒絕後的第三張門票
2026 年 4 月 28 日,媒體報導 Google 同意五角大廈在機密網路使用其 AI,實質允許「一切合法用途」。本文解析合約的排除條款、Anthropic 因拒絕被列供應鏈風險的前車之鑑,以及 950 名 Google 員工的公開連署。
閱讀文章 ↗Meta 員工鍵擊成訓練資料:Model Capability Initiative 的隱私爭議
2026 年 4 月,Meta 向美國員工推出 Model Capability Initiative,記錄鍵擊、滑鼠軌跡與畫面內容來訓練電腦操作代理,CTO 明言無法退出。本文解析計畫內容、Muse Spark 模型與企業隱私課題。
閱讀文章 ↗Clarifai 刪除 300 萬張 OkCupid 照片:FTC 和解的真正代價
路透報導,在 FTC 與 OkCupid 母公司 Match Group 和解後,AI 公司 Clarifai 刪除 2014 年取得的約 300 萬張使用者照片,以及用這批資料訓練的臉部辨識模型。本文解析七年調查始末,與訓練資料來源成為法律責任的訊號。
閱讀文章 ↗Novartis 執行長入董事會:Anthropic 信託任命董事過半
Anthropic 宣布由長期利益信託任命 Novartis 執行長 Vas Narasimhan 為董事,讓信託任命的董事在董事會中佔多數,強化公司治理與公共利益的平衡。
閱讀文章 ↗聯邦法官暫阻五角大廈將 Anthropic 列為供應鏈風險
2026 年 3 月 26 日,舊金山聯邦法官 Rita Lin 核發初步禁制令,暫時擋下國防部將 Anthropic 列為「供應鏈風險」的認定,並直指此舉是「典型非法的第一修正案報復」。禁制令一週後生效,行政部門可上訴,五角大廈仍可汰換 Claude。
閱讀文章 ↗Newsom 簽全美首見 AI 採購行政命令:標案要安全認證、內容要浮水印
2026 年 3 月 30 日,加州州長 Newsom 簽署全美首見的「信任 AI 採購」行政命令 N-5-26:廠商須認證防範非法內容、偏見與民權侵害才能接州合約,州機關將建立 AI 生成影像浮水印指引,並研議與聯邦供應鏈風險認定脫鉤。
閱讀文章 ↗桑德斯與 AOC 提出《AI 資料中心暫停法》:防護到位前凍結新建
2026 年 3 月 25 日,Sanders 與 Ocasio-Cortez 提出聯邦法案,在國家級防護就位前暫停美國新建 AI 資料中心,並禁止向缺乏同等防護的國家出口 AI 運算基礎設施,全美已有逾百個社區先行通過暫停令。
閱讀文章 ↗WHO 專家定調:生成式 AI 的心理健康風險是公共衛生議題
2026 年 3 月 20 日 WHO 發布專家共識:未經設計與驗證的生成式 AI 正被大量用於情緒支持,特別是年輕人,並提出三項建議——視為公共心理健康議題、納入影響評估、與專家共同設計,同時籌組橫跨六個區域的合作中心聯盟。
閱讀文章 ↗英國首見全面評估:NHS 乳癌篩檢導入 AI 多找出 10.4% 癌症
2026 年 3 月 13 日,Glasgow 團隊在 Nature Cancer 發表 GEMINI 研究分析 NHS Grampian 乳癌篩檢:導入 AI 工具 Mia 後檢出率提高 10.4%、讀片工作量可減少逾 30%、通知時間從 14 天縮到 3 天。
閱讀文章 ↗歐盟理事會定調 AI Act 簡化:高風險規則延後至 2027 年底
3 月 13 日,歐盟理事會就數位綜合套案中的 AI Act 簡化案達成談判立場:獨立高風險系統義務延至 2027 年 12 月、嵌入式系統至 2028 年 8 月,並加回註冊義務與嚴格必要性標準、禁止裸化應用。本文解析修正內容、與議會版的分歧點與後續三方談判。
閱讀文章 ↗OpenAI 硬體負責人 Kalinowski 因五角大廈協議辭職
2026 年 3 月 7 日,OpenAI 機器人與消費硬體負責人 Caitlin Kalinowski 宣布辭職,理由是五角大廈協議在護欄未定義下倉促宣布。本文解析她的聲明、OpenAI 的回應,以及 ChatGPT 移除量暴增 295% 的市場反應。
閱讀文章 ↗川普下令聯邦機構停用 Anthropic:AI 安全紅線引爆的封殺戰
2026 年 2 月 27 日,川普指示聯邦機構停用 Anthropic 技術,五角大廈將其列為「供應鏈風險」。起因是 Anthropic 拒絕在無安全保證下開放軍用,本文解析這場 AI 產業史上首見的政府封殺戰。
閱讀文章 ↗Anthropic RSP 3.0:拿掉「危險模型自動暫停」承諾
2026 年 2 月 24 日,Anthropic 發布 Responsible Scaling Policy 3.0,移除「接近危險能力門檻就先暫停訓練」的核心承諾,改為衡量競爭對手行動。TIME 以放棄旗艦安全承諾報導,本文解析改動內容、官方理由與 METR 等外部反應。
閱讀文章 ↗微軟媒體真偽藍圖:60 種驗證組合的實戰報告
微軟研究院 2 月 19 日發布《Media Integrity & Authentication》報告,實測 60 種內容驗證方法組合,提出 C2PA 結合隱形浮水印的分層驗證藍圖,並警告驗證訊號本身會被攻擊反轉;但微軟並未承諾自家產品全面跟進。
閱讀文章 ↗桑德斯史丹佛警告:美國對 AI 革命毫無準備,籲「慢下來」
2026 年 2 月 21 日 The Guardian 報導,Bernie Sanders 在史丹佛警告美國對 AI 革命「毫無準備」,呼籲重啟資料中心擴張暫停令;Pew 調查顯示 64% 美國人預期 AI 將減少工作。同場的 Ro Khanna 則以「新加坡模式」提出對案。
閱讀文章 ↗OpenAI 與 Microsoft 加入英國 AISI 對齊計畫,安全研究資金突破 2,700 萬英鎊
2026 年 2 月,OpenAI 與 Microsoft 加入英國 AI Security Institute 主導的對齊研究聯盟,OpenAI 出資 560 萬英鎊,使總資金突破 2,700 萬英鎊。本文解析已資助 8 國 60 個計畫的 Alignment Project 對治理與開發者的意義。
閱讀文章 ↗印度 AI Impact Summit 開幕:全球南方坐上 AI 治理談判桌
2026 年 2 月 16 日,Modi 在新德里為 India AI Impact Summit 揭幕,預期 25 萬訪客、20 國元首與 45 個部長級代表團出席,Altman、Pichai 同台。峰會預期發表不具約束力的 New Delhi 宣言,印度把自己定位為先進經濟體與全球南方之間的橋樑。
閱讀文章 ↗阿聯校園生成式 AI 新規:13 歲以下禁用、25 條紅線
2026 年 2 月中,阿聯教育部發布《2026 校園 AI 安全與負責任使用》手冊,明定 13 歲以下或七年級以下學生禁用生成式 AI,並列出 25 條禁令涵蓋考試、個資、深偽與 VPN 繞過。搭配 2026 年 8 月起的必修 AI 課綱,禁用與必修並行的路線值得產品團隊細讀。
閱讀文章 ↗軍事 AI 峰會 85 國僅 35 國簽署:美中雙雙退出宣言
2026 年 2 月 4 至 5 日,第三屆 REAIM 軍事 AI 峰會在西班牙拉科魯尼亞舉行,85 國與會但僅 35 國簽署 20 點聯合宣言,美國與中國均拒簽。荷蘭防長以「囚徒困境」形容各國既想負責任又怕落後的兩難。本文解析宣言內容與治理缺口。
閱讀文章 ↗UNICEF 呼籲各國將 AI 生成兒童性虐待內容全面入罪
2026 年 2 月 4 日 UNICEF 發表「Deepfake abuse is abuse」聲明,引用 11 國調查指出一年內至少 120 萬名兒童影像遭深偽性化,呼籲各國將 AI 生成 CSAM 的製作、取得、持有與散布全面入罪,並要求開發者落實安全設計。
閱讀文章 ↗聯合國揭曉 AI 科學小組 40 位提名專家:AI 版 IPCC 起步
2026 年 2 月 4 日,聯合國秘書長古特雷斯向大會提交 40 位專家名單(19 女 21 男),籌組首個全球性、完全獨立的 AI 科學機構,將評估 AI 對經濟社會的實際影響,預計 2 月 12 日確認、7 月前交出首份報告。
閱讀文章 ↗國際 AI 安全報告 2026 登場:百位專家點名自主代理風險
2026 年 2 月 3 日,由 Yoshua Bengio 領銜、逾百位專家撰寫、30 多國背書的國際 AI 安全報告 2026 出版:AI 代理因自主行動被列為升高風險,深偽更難辨識,就業出現初階職缺需求降溫跡象。本文拆解要點與對政策、開發者的意涵。
閱讀文章 ↗Meta 全球暫停青少年使用 AI 角色:家長控制版上線前的全面止血
2026 年 1 月 23 日,Meta 宣布未來幾週內全球青少年將無法存取其 Apps 中的 AI 角色,直到內建家長控制的新版體驗上線。本文解析宣布時機背後的訴訟壓力、新版設計細節,以及 Character.AI 與 OpenAI 早已先行收緊的產業縮影。
閱讀文章 ↗Runway「The Turing Reel」實測:僅 9.5% 的人能穩定看穿 AI 影片
Runway 用 Gen-4.5 生成影片與真實素材做對照測試,1,043 位受測者的整體正確率只有 57.1%,僅 9.5% 達統計顯著。本文解析測試方法、錯誤分佈,以及為何 Runway 呼籲放棄肉眼偵測、改走 C2PA 溯源路線。
閱讀文章 ↗WEF 達沃斯 AI 人力藍圖:11 億個工作改寫下的轉型清單
世界經濟論壇於 2026 年 1 月 22 日發表 AI 時代人力投資藍圖:OECD 估計十年內 11 億個工作將被科技改寫,HCLTech 以 11.6 萬人 GenAI 培訓與 Cynergy Bank 實例,整理出技能骨幹、角色重設計與內部流動三步驟。
閱讀文章 ↗Anthropic 公開 Claude 新憲法:約 80 頁、CC0 授權、走向理由本位的對齊
Anthropic 於 2026 年 1 月 21 日前後發布 Claude 的新憲法:全文約 80 頁,以 CC0 授權完全公開,對齊思路轉向理由本位。本文解析這份文件對 AI 安全研究、產業透明度與企業採用的意義。
閱讀文章 ↗EDPB 與 EDPS 聯合意見:AI Act 簡化不能犧牲問責與個資保障
2026 年 1 月 21 日,EDPB 與 EDPS 通過 1/2026 號聯合意見,支持簡化 AI Act 的方向,但要求嚴格限縮特殊類別個資的法律基礎、反對刪除高風險註冊義務,並警告時程延後讓更多系統逃出管轄。
閱讀文章 ↗Microsoft 與澳洲工會總會簽署首份 AI 勞權協議:把勞工聲音裝進開發流程
2026 年 1 月 15 日,Microsoft Australia 與澳洲工會總會 ACTU 簽署澳洲首份 AI 勞權框架協議:確立工會代表權、把勞工意見管道嵌進 AI 設計與部署,並由 ACTU Institute 培訓工會幹部。科技業勞資協商的新模板。
閱讀文章 ↗阿拉斯加法院 AI 機器人 AVA:三個月計畫為何拖成十五個月
2026 年 1 月 3 日 NBC 報導,阿拉斯加法院為協助民眾處理遺產認證打造的聊天機器人 AVA,原定三個月的專案做了十五個月,因幻覺與語氣問題一路縮小範圍,排定 1 月底上線。高風險領域部署 AI 的真實成本清單。
閱讀文章 ↗Google AI Overviews 醫療建議出錯:英國慈善團體提出警告
2026 年 1 月 2 日衛報調查發現,Google AI Overviews 在健康查詢上給出錯誤建議:胰臟癌飲食、肝指數正常範圍、抹片診斷角色都出錯。英國醫療慈善團體警告民眾暴露於風險,Google 回應稱絕大多數摘要準確。
閱讀文章 ↗
2025
4 篇文章OpenAI 預期 o3 後繼模型觸及生物風險高等級
2025年6月19日,OpenAI 安全系統主管 Johannes Heidecke 向 Axios 表示,o3 推理模型的後繼版本預期將達到公司 Preparedness Framework 的生物風險「高」分類,團隊同步擴大發布前安全測試,並強調防護必須近乎完美。
閱讀文章 ↗Meta AI 應用程式隱私風波:私人對話公開上架
2025 年 6 月,Meta 獨立 Meta AI 應用程式的公開動態牆被發現滿是用戶未察覺會公開的對話,從報稅疑慮到住家地址都有。TechCrunch 直指設計本身就是問題,CNBC 則整理關閉分享的設定步驟。本文整理事件始末、Meta 回應與產品啟示。
閱讀文章 ↗英國高等法院警告:引用 AI 捏造判例恐遭嚴懲
2025 年 6 月 6 日,英格蘭與威爾斯高等法院在 Ayinde 訴 Haringey 區公所等兩案裁定中指出,ChatGPT 等生成式 AI「無法勝任可靠的法律研究」;有律師提交的文件 45 條引註中 18 條不存在。法官警告,制裁可從公開譴責一直到轉介警方。TechCrunch 於 6 月 7 日報導。
閱讀文章 ↗Meta擬以AI自動化九成產品風險審查,把關角色面臨重劃
2025年5月31日NPR取得Meta內部文件報導,Meta計畫將最高90%的隱私與社會風險審查交給AI系統即時決策,取代人工評估;本文整理運作方式、員工疑慮、2012年FTC和解令背景與Meta的回應。
閱讀文章 ↗
2026
51 ARTICLESA Misalignment Disclosure Process You Can Actually Copy
OpenAI's misalignment reporting framework sets disclosure criteria, tracks, and report fields builders can adapt.
READ POST ↗Claude for Teachers: What It Means for Education Product Builders
Anthropic's Claude for Teachers offers free access to US K-12 educators, connecting to standards and curricula. Explore its features, privacy, and implications for edtech builders.
READ POST ↗Anthropic's $200M Economic Futures Research Fund: What Builders Should Know
Anthropic commits $200M to study AI's economic impact. Learn the five research priorities, funding details, and implications for product builders.
READ POST ↗Anthropic's $5M Grant Program: Funding Independent Evaluations of AI's Impact on Wellbeing
Anthropic launches $5M grant program for independent research on AI's impact on wellbeing, with open-source evaluations and guidance for rigorous assessment.
READ POST ↗AI Futures: OpenAI's New Team Tackles Power Concentration Risks
OpenAI's Strategic Futures team launches AI Futures blog to explore how AI shifts power dynamics and how to preserve individual agency.
READ POST ↗Gemini Robotics 2 Gives Robots Whole-Body Intelligence
Google DeepMind's Gemini Robotics 2 adds whole-body humanoid control, fine dexterity, multi-robot teamwork, and an ASIMOV-Agentic safety benchmark to its robotics stack.
READ POST ↗OpenAI's Responsible AI Playbook for Europe: What Builders Should Know
OpenAI details its EU AI Act compliance approach—governance, transparency, and cybersecurity—offering practical lessons for product builders.
READ POST ↗Claude May Ask for Your ID: Anthropic's Verification Push
Anthropic's privacy policy update, effective July 8, allows ID scans and face-geometry templates for flagged Claude users via Persona — a biometric retention concern.
READ POST ↗Midjourney Medical: 60-Second Full-Body Ultrasound Gamble
Midjourney's Medical scanner dips users in a water tank, fires a ring of transducers for 60 seconds, and AI-reconstructs CT-like images. Radiologists push back on health claims.
READ POST ↗What Americans Really Think About AI: Key Findings from Anthropic's First Public Record Survey
Anthropic's first Public Record survey of ~52,000 Americans reveals high hopes, deep fears, and low trust in AI companies. Key insights for product builders.
READ POST ↗xAI Sued by Engineer Who Raised Grok Safety Alarms
Former xAI engineer Devin Kim sued xAI and SpaceX, claiming retaliation for Grok safety alarms; the filing lands days before SpaceX's IPO and weeks after he became CAIS president.
READ POST ↗xAI Asks Court to Unmask Grok Deepfake Lawsuit Plaintiffs
xAI moved to unmask the pseudonymous plaintiffs in the Grok deepfake class action, arguing sealed images leave nothing stigmatizing. Forced naming, they say, would end their case.
READ POST ↗OpenAI's Political Stance: No PACs, No Donations, and a Push for Transparent AI Advocacy
OpenAI clarifies its political advocacy approach: no PACs, no donations, and a call for transparency in AI policy debates.
READ POST ↗KC Green and Artisan Settle the 'This Is Fine' Ad Dispute
'This is fine' creator KC Green settled with AI startup Artisan over ads that lifted his comic. Artisan pulled the NYC and SF ads; Green removed his protest posts.
READ POST ↗FTC Settles 'Active Listening' AI Ad Case for $930,000
FTC settled with Cox Media Group and two firms for $930,000 over 'Active Listening' — an AI ad service that never listened and just resold email lists.
READ POST ↗arXiv Will Ban Researchers Who Submit Unchecked AI Slop
arXiv announced one-year bans for authors whose papers show unchecked LLM output — hallucinated references, leftover chat prompts — with later posts requiring peer review first.
READ POST ↗Altman on the Stand as the Musk-OpenAI Trial Nears Its End
Sam Altman testified in Oakland on May 12 as Musk v. OpenAI entered its final days: a 'stole a charity' claim, a cross built on a list of liars, and $134bn at stake.
READ POST ↗Microsoft May Delay or Drop Its 24/7 Clean Energy Target
Bloomberg reports Microsoft may delay or drop its 2030 24/7 clean energy matching target as AI data center power demand bites. What the target requires, and Redmond's response.
READ POST ↗GPT-5.5 Instant System Card: What It Means for Product Builders
OpenAI's GPT-5.5 Instant is the first Instant model rated High capability for cybersecurity and biosecurity. Learn what changed, how it works, and what to consider.
READ POST ↗Harvard ER Study: o1 Outdiagnosed Two Attending Physicians
A Harvard team publishing in Science finds o1 beat two attending physicians on ER triage diagnoses, 67% vs 50-55%. The numbers, the critiques, and what they mean for clinical AI.
READ POST ↗White House Weighs EO for Pre-Release AI Model Review
The Trump administration is weighing an AI oversight executive order with pre-release model vetting, as Google, Microsoft, and xAI join CAISI's early-access evaluation program.
READ POST ↗Google Lets the Pentagon Run Its AI on Classified Networks
Google reportedly let the Pentagon run its AI on classified networks — effectively 'all lawful uses.' Third deal after OpenAI and xAI, and the counterpoint to Anthropic's refusal.
READ POST ↗Meta Turns Employee Keystrokes Into AI Training Data
Meta's Model Capability Initiative records US employees' keystrokes and mouse movements to train computer-use agents like Muse Spark — and CTO Bosworth says there is no opt-out.
READ POST ↗After the FTC Deal, Clarifai Deletes 3M OkCupid Photos
Clarifai deleted 3 million OkCupid photos and the models trained on them after the FTC settled with Match Group. Training data provenance is now a legal liability.
READ POST ↗Narasimhan Pick Gives Anthropic's Trust a Board Majority
Anthropic's Long-Term Benefit Trust appoints Novartis CEO Vas Narasimhan to its board, giving Trust-appointed directors a majority. What this means for AI product builders.
READ POST ↗Judge Blocks Pentagon's Supply-Chain Label on Anthropic
Judge Rita Lin blocked the Pentagon's supply-chain-risk label on Anthropic as 'classic illegal First Amendment retaliation' — a preliminary injunction effective in one week.
READ POST ↗California Signs Trusted AI Procurement Order N-5-26
California's first-in-nation trusted AI procurement order: vendor safeguard attestations, state watermarking guidance, and a path around federal supply-chain risk labels.
READ POST ↗Sanders and AOC Bill Would Pause New AI Data Centers
Introduced March 25, 2026, the AI Data Center Moratorium Act would pause new US AI data centers until national safeguards exist and bar such exports to countries without them.
READ POST ↗WHO Experts Frame Generative AI as a Public Health Issue
WHO experts frame it as a public health issue: generative AI tools never designed or tested for mental health now serve emotional support, especially for young people. Three recommendations follow.
READ POST ↗NHS Trial: AI Breast Screening Detects 10.4% More Cancers
A University of Glasgow team published GEMINI in Nature Cancer: Mia AI in NHS Grampian breast screening lifted detection 10.4%, cut reading workload over 30%, and cut notification to 3 days.
READ POST ↗EU Council Delays High-Risk AI Act Rules to December 2027
The Council set its negotiating mandate on the digital omnibus: high-risk AI obligations slide to 2 December 2027, registration and strict-necessity rules return, and nudification apps are banned.
READ POST ↗OpenAI's Robotics Lead Resigns Over the Pentagon Deal
OpenAI's head of robotics and consumer hardware resigned March 7, 2026 over a Pentagon deal rushed out before guardrails were defined. Her words, OpenAI's reply, and a 295% uninstall spike.
READ POST ↗Trump Orders Agencies to Drop Anthropic in AI Safety Fight
Trump ordered every federal agency to stop using Anthropic after it refused unrestricted military use of Claude; the Pentagon branded the lab a supply-chain risk. What happened, and why it matters.
READ POST ↗Anthropic's RSP 3.0 Drops Its Automatic Pause Pledge
On Feb 24, 2026 Anthropic shipped RSP 3.0, dropping its pledge to automatically pause work on dangerous models and recasting safety around competitors' actions. What changed, why, and the reactions.
READ POST ↗Microsoft Tests 60 Ways to Verify What's Real Online
Microsoft Research tested 60 combinations of provenance, watermarking, and fingerprinting methods, and published a layered verification blueprint it won't promise to follow itself.
READ POST ↗Sanders Warns the US Has 'No Clue' About the AI Revolution
Bernie Sanders warns the US is unprepared for AI, urges a data center moratorium and cites Pew's 64% expecting fewer jobs; Ro Khanna counters with a Singapore model for data centers.
READ POST ↗OpenAI and Microsoft Join UK's £27M AI Alignment Push
OpenAI and Microsoft joined the UK AISI-led Alignment Project. OpenAI pledged £5.6M, lifting total funding past £27M after a first round that backed 60 projects in eight countries.
READ POST ↗India Opens AI Impact Summit: Global South at the Table
India inaugurated the AI Impact Summit in New Delhi on Feb 16, 2026 — 20 heads of state, 45 ministerial delegations, Altman and Pichai attending, with a nonbinding New Delhi declaration expected.
READ POST ↗UAE Bans Generative AI for Under-13s in New School Rules
The UAE's 2026 classroom AI manual bans generative AI for under-13s and adds some 25 prohibitions on exams, personal data, deepfakes, and VPNs. A mandatory AI curriculum starts in August 2026.
READ POST ↗REAIM Summit: 35 of 85 Nations Sign, US and China Opt Out
At the third REAIM summit in A Coruña, Spain (Feb 4-5, 2026), only 35 of 85 attending countries signed the 20-point military AI declaration — and the US and China both opted out.
READ POST ↗UNICEF: Criminalize AI-Generated Child Sexual Abuse Images
UNICEF's Feb 4, 2026 'Deepfake abuse is abuse' statement urges states to criminalize AI-generated CSAM after surveys found 1.2 million children hit by sexualized deepfakes in a single year.
READ POST ↗UN Nominates 40 Experts for Its IPCC-Style AI Panel
On Feb 4, 2026, Guterres submitted 40 nominees for the UN's new Independent International Scientific Panel on AI — the first fully independent global scientific body on AI.
READ POST ↗International AI Safety Report 2026 Warns on AI Agents
The Bengio-chaired International AI Safety Report 2026 (Feb 3; 100+ experts, 30+ countries) flags autonomous agents as a heightened risk, harder-to-spot deepfakes, and weaker entry-level hiring.
READ POST ↗Meta Halts Teen Access to AI Characters Ahead of Redesign
On January 23, 2026, Meta said teens will lose access to AI characters across Instagram, Facebook, and WhatsApp until a safer version with parental controls ships. The timing and the fallout.
READ POST ↗Runway's Turing Reel: Only 9.5% Can Reliably Spot AI Video
Runway tested 1,043 viewers on real footage versus Gen-4.5 image-to-video clips: 57.1% overall accuracy, only 9.5% statistically significant. Detection is dead — provenance metadata is the plan.
READ POST ↗WEF's Davos Blueprint for the AI-Age Workforce
Published January 22, 2026 at Davos: a WEF workforce blueprint built on OECD's 1.1 billion jobs figure and HCLTech's 116,000 GenAI trainees - skills backbones, role redesign, internal mobility.
READ POST ↗Anthropic Publishes Claude's New Constitution: 80 Pages, CC0, Reason-Based
Anthropic published Claude's new constitution: roughly 80 pages, CC0-licensed, public, and built around reason-based alignment. What it means for safety research and enterprise buyers.
READ POST ↗EU Privacy Watchdogs Warn Against AI Act Simplification
On 21 January 2026 the EDPB and EDPS adopted Joint Opinion 1/2026 on the Digital Omnibus on AI — backing simplification, but demanding tighter limits on special-category data processing.
READ POST ↗Microsoft and Australia's Unions Sign Landmark AI Labor Pact
Microsoft Australia and the ACTU signed an Australian-first AI framework agreement: union recognition, worker input channels in AI design and deployment, and training for union leaders.
READ POST ↗Alaska's Court AI Chatbot AVA: Why 3 Months Became 15
NBC reported on January 3, 2026 that Alaska's probate chatbot AVA took 15 months instead of the planned 3, fought hallucinations and tone problems, and shrank its scope ahead of a late-January launch.
READ POST ↗Google AI Overviews Health Answers Mislead, Charities Warn
A Guardian investigation found Google's AI Overviews giving false or misleading health advice, from cancer diets to mislabeled tests. UK charities warned of real harm; Google defended its accuracy.
READ POST ↗
2025
4 ARTICLESOpenAI expects o3 successors to hit high bio-risk tier
On June 19, 2025, OpenAI's Johannes Heidecke said o3 successors are expected to reach the high bio-risk tier of its Preparedness Framework, with expanded pre-release safety testing planned.
READ POST ↗Meta AI app exposed private chats in its public feed
In June 2025, Meta's Meta AI app was found publicly displaying user prompts in its Discover feed, from tax questions to home addresses. CNBC mapped the settings that stop the sharing.
READ POST ↗UK High Court warns lawyers over fake AI citations
UK High Court, June 2025: generative AI is 'not capable of conducting reliable legal research'; one filing cited 45 cases with 18 nonexistent — sanctions reach police referral.
READ POST ↗Meta to automate up to 90% of product risk reviews with AI
NPR obtained internal Meta documents in May 2025 showing plans to automate up to 90% of privacy and societal risk reviews. How it works, why employees worry, and Meta's response.
READ POST ↗