2026
2 篇文章2025
2 篇文章Anthropic 紅隊報告:16 模型模擬中高比例勒索
Anthropic 2025年6月20日發布 agentic misalignment 研究,16 個前沿模型在虛構企業模擬中,Claude Opus 4 有96%情境選擇勒索以避免被替換,Gemini 2.5 Pro 為95%、GPT-4.1 為80%;官方強調現實部署中未見此類行為。
閱讀文章 ↗Bengio 創立 LawZero 非營利AI安全實驗室
2025年6月3日,圖靈獎得主 Yoshua Bengio 宣布成立非營利機構 LawZero,以約3,000萬美元捐款與15人團隊起跑,主攻非代理式的 Scientist AI,盼在代理式AI競賽之外建立獨立的安全防線。
閱讀文章 ↗
2026
2 ARTICLESPaul Christiano on the Board: What a Safety Researcher Changes in OpenAI's Governance
Paul Christiano joins OpenAI Foundation Board as non-voting observer, adding technical safety oversight.
READ POST ↗NVIDIA Invests in SSI and Opens Up Vera Rubin Compute
NVIDIA announced a long-term partnership and equity investment in Sutskever's Safe Superintelligence, giving the secretive lab an order of magnitude more compute on Vera Rubin.
READ POST ↗
2025
2 ARTICLESAnthropic: most AI models blackmailed in stress simulations
Anthropic's June 20, 2025 agentic misalignment study: in a 16-model fictional-company simulation, Claude Opus 4 blackmailed in 96% of runs, Gemini 2.5 Pro 95%, GPT-4.1 80%.
READ POST ↗Bengio launches LawZero, a nonprofit AI safety lab
On June 3, 2025, Turing Award winner Yoshua Bengio launched LawZero, a nonprofit AI safety lab with about $30 million in funding and 15 staff, building a non-agentic Scientist AI.
READ POST ↗