2026
4 篇文章當模型開始替你的系統做測試:Perplexity 把 GPT‑6 Astra 放進端到端流程
Perplexity 讓 GPT‑6 Astra 代寫測試程式、模擬外部服務並監控正式系統,檢查頻率明顯低於前幾代模型。
閱讀文章 ↗Astra 實測數據拆解:ExploitBench 滿分、兩個零日漏洞,與少用 9% token 的祕密
Astra 在 ExploitBench 拿下 100% 解題率(GPT-5.6 Sol 只有 22%),評測過程甚至發現兩個全新零日漏洞。本文從開發者角度拆解官方評測數據:V8 內部移植、沙箱逃逸鏈、提權鏈,以及 token 效率為何比 Raw 能力更值得注意。
閱讀文章 ↗OpenAI 推出 Astra:第一個觸發 Critical 網路門檻的模型如何安全上架
OpenAI 於 2026 年 9 月 3 日推出 Astra,這是第一個被判定達到 Preparedness Framework Critical 網路能力門檻的模型。本文拆解 Daybreak 分層存取、預設封鎖的護欄設計、91.5% 的拒絕率與 honeypot 測試結果,以及對企業導入者的實際意涵。
閱讀文章 ↗OpenAI 首度無法排除 Astra 達 Critical 網路門檻
OpenAI 內部評估首度無法排除 Astra 觸及 Preparedness Framework 的 Critical 網路安全門檻:agentic coding 與 cyber 能力大幅躍進,五層控制與思緒監控已啟動,外部紅隊測試是下一個觀察點。
閱讀文章 ↗
2026
3 ARTICLESWhat It Takes to Hand an Agent the Whole System
Perplexity lets GPT-6 Astra edit production systems and check in less often, shifting the trust question to oversight.
READ POST ↗Inside Astra's Cyber Evals: 100% on ExploitBench, Two Zero-Days, and 9% Fewer Tokens
Astra scored 100% on ExploitBench where GPT-5.6 Sol scored 22%, and the eval surfaced two zero-days. A breakdown of the V8 port, the escape chains, and the token-efficiency gain.
READ POST ↗OpenAI Ships Astra: How the First Critical-Threshold Cyber Model Went Live
OpenAI's September 3 launch of Astra is the first model at the Critical cyber threshold of its safety framework: Daybreak tiered access, default-off guardrails, and pause-and-review monitoring.
READ POST ↗