2026
6 篇文章OpenRouter Fusion 登場:多模型融合輸出在 DRACO 勝過最強單一模型
OpenRouter 於 6 月 12 日推出 Fusion:單一 API 呼叫把 prompt 平行送進多個模型,再由裁判模型整合共識與矛盾後交回原模型作答。DRACO 深度研究基準上,融合面板拿下 69.0%,超越所有單一模型,平價組合以約一半成本逼近 Fable 5。
閱讀文章 ↗Deep Max:為何 Exa 把代理搜尋的速度推上新高?
Exa 推出 Deep Max 代理搜尋端點,結合前沿 LLM 與平行搜尋,在主要 agentic 搜尋基準達到 SOTA 準確率、速度領先對手達 20 倍,但定價未公開。本文拆解其設計原理與產品端的實際取捨。
閱讀文章 ↗DeepSearchQA 實測:Ultra 準度勝 GPT-5.4,成本僅 43%
沙盒直譯器留住中間資料:20 步研究從 128K 壓到 30K tokens;Ultra 以 $300/千次拿到 70%,勝過 GPT-5.4。
閱讀文章 ↗Web Search 與 Deep Research:2026 年 Agent 的資料層已經變成基礎設施
2026 年,AI agent 的 web search 與 deep research 已從實驗性功能變成生產級基礎設施。本文整理 Firecrawl 部落格的分析,說明兩者差異、市場變化、實際應用案例,以及如何整合進 agentic stack。
閱讀文章 ↗Deep Research 接上任何 MCP:ChatGPT 研究代理轉向可信來源
2026 年 2 月 10 日 OpenAI 改版 ChatGPT Deep Research:可連接任何 MCP 或應用、把網路搜尋限縮在指定網站、研究計畫可先編輯、中途可即時調整方向,底層換上 GPT-5.2。研究代理正從開放網路轉向受控的資料層。
閱讀文章 ↗Perplexity Model Council:三個前線模型同時作答,分歧也是答案
Perplexity 在 2 月 5 日推出 Model Council:同一個問題交給三個前線模型平行作答,再由另一個模型綜整共識與分歧。搭配升級版 Deep Research 與記憶改善,多模型交叉驗證正式走進搜尋產品。
閱讀文章 ↗
2026
5 ARTICLESOpenRouter Fusion: Multi-Model Answers Beat Single Models
OpenRouter Fusion sends one prompt to a model panel, then a judge fuses the answers. On DRACO a fused panel scored 69.0%, above every solo model, at a fraction of the cost.
READ POST ↗Parallel Beats GPT-5.4 on DeepSearchQA at 43% of the Cost
Sandboxed Python keeps a 20-step research task under 30K tokens instead of 128K; Parallel's Ultra scores 70% on DeepSearchQA at $300 per 1K requests.
READ POST ↗Web Search and Deep Research for AI Agents: From Experiment to Infrastructure
How web search and deep research became production infrastructure for AI agents in 2026, with architecture, use cases, and integration steps.
READ POST ↗Deep Research Now Connects to Any MCP and Trusted Sites
OpenAI's Feb 10, 2026 Deep Research overhaul: connect any MCP app, restrict searches to trusted sites, edit the plan, steer mid-run, now on GPT-5.2. Research agents move to a controlled data layer.
READ POST ↗Perplexity's Model Council: Three Models Answer as One
Perplexity's Model Council runs one query through three frontier models in parallel, then synthesizes agreements and disagreements into one answer — ensemble verification shipped in a search product.
READ POST ↗