2026
11 篇文章Fable 5.1 的省錢關鍵不是模型,而是你的快取讀取比例
Firecrawl 實測 57 次 API 呼叫後發現,Fable 5.1 只有在快取讀取密集的長代理任務才便宜,其他情境反而更貴。
閱讀文章 ↗把多模型辯論變成一個 API 呼叫:OpenRouter Fusion 的取捨與使用時機
OpenRouter Fusion 用平行面板加裁判來提升研究品質,但代價是 4-5 倍成本與 2-3 倍延遲。
閱讀文章 ↗把翻譯模型放進產品前,先看吞吐量與長文件的真實落差
North Small Translate 以 25B 活躍參數在 WMT26 拿下 83.6 分,並在長文件與吞吐量上拉開差距,改變了翻譯功能的建置取捨。
閱讀文章 ↗別再只看每百萬 token 報價:Amazon Bedrock 上的 OpenAI 模型該怎麼挑
AWS 的開源評測把「答對一題要多少錢」和「代理跑幾輪才答對」攤開來算,讓模型選擇回到工作負載本身。
閱讀文章 ↗細模型的大優勢:企業如何用小語言模型省錢又高效
當企業面對眾多 AI 模型時,選擇哪一個才能兼顧效能與成本,成了關鍵課題。大型語言模型(LLM)常佔據新聞版面,但許多組織發現,較小、專用的小語言模型(SLM)反而能提供顯著優勢:更低的運算需求、更少的訓練資料、更省能源,以及更具成本效益的解決方案。
閱讀文章 ↗選影像模型,終於不用只看範例圖了
OpenRouter 推出 Visual Image Benchmarks,用七類挑戰性 Prompt 比較 39 個影像模型,補上文字基準之外的空白。產品工作者可以先用它篩選,再拿自己的內容驗證。
閱讀文章 ↗OpenRouter 完成 1.13 億美元 B 輪:模型路由層估值翻倍至 13 億美元
2026 年 5 月,OpenRouter 宣布 1.13 億美元 B 輪融資,由 CapitalG 領投,估值翻倍至約 13 億美元。每月處理 100 兆 token、支援 400+ 模型,本文解析模型路由層為何成為 Agent 時代的關鍵基礎設施。
閱讀文章 ↗Interfaze 混合架構登場:自報基準碾壓 flash 級對手
JigsawStack 團隊推出混合架構模型 Interfaze,把專用神經網路併進 omni-transformer,主打 OCR、語音轉文字等確定性任務:OCRBench V2 70.7% 對 Gemini-3-Flash 55.8%。本文檢視九項自報基準與切換風險。
閱讀文章 ↗OpenAI 的「哥布林」之謎:獎勵機制如何悄悄塑造模型行為
OpenAI 公開調查 GPT-5.1 以來模型頻繁提及「哥布林」等生物的現象,發現源於「Nerdy」個性訓練的獎勵訊號,並因回饋迴圈擴散。本文拆解根因、傳播機制與教訓,對產品開發者與 AI 學習者深具啟發。
閱讀文章 ↗GPT-5.4 前夕:從發佈節奏看 OpenAI 的版本加速與 API 對策
從官方 release notes 看:GPT-5.2 於 2025 年 12 月 11 日發佈,GPT-5.2-Codex 於 2026 年 1 月 14 日上線,GPT-5.3-Codex 於 2 月 5 日登場,不到兩個月三個版本。本文分析加速的命名節奏與 API 開發者的因應對策。
閱讀文章 ↗GPT-4o 從 ChatGPT 退役:0.1% 使用率背後的模型整併
OpenAI 宣布 2 月 13 日將 GPT-4o、GPT-4.1 系列與 o4-mini 從 ChatGPT 退役,僅 0.1% 使用者每天還在用 GPT-4o,API 不受影響。本文整理時程、決策理由,與 2025 年 8 月風波後「充分預告」承諾的兌現。
閱讀文章 ↗
2026
12 ARTICLESFable 5.1's Cache Discount: Where the Bill Actually Moves
Fable 5.1 cut cache reads 75%, but 57 billed runs show the saving depends on how often your agent re-reads context.
READ POST ↗OpenRouter Fusion: What a Compound Model Changes for Your Escalation Path
Fusion adds a multi-model deliberation loop to one API call, trading latency and tokens for better research answers.
READ POST ↗North Small Translate: What 16k Context and 1.4x Throughput Change for Translation Pipelines
Cohere's open-weight translation model pairs 16k context with 1.4x throughput, shifting how builders handle long documents.
READ POST ↗Cost Per Correct Answer: Picking an OpenAI Model on Bedrock
AWS benchmark data shows turn count and retry rate, not token price, decide what a correct answer actually costs.
READ POST ↗Small Language Models: The Enterprise Case for Right-Sizing AI
Cohere's guide to SLMs shows why smaller models can cut costs, run locally, and even beat larger ones on specific tasks. Learn how to build a model portfolio that matches size to job.
READ POST ↗OpenRouter's Image Benchmarks: A Practical Guide for Product Builders
OpenRouter launches visual image benchmarks to help developers compare 39 image models across 7 challenge categories, with practical advice for product builders.
READ POST ↗OpenRouter MCP: Live Model Data for Smarter Agent Decisions
OpenRouter's MCP server gives coding agents live model data, benchmarks, pricing, docs, and test inference. Eleven tools, only one billed; the OAuth key has a seven-day expiry and a ten-dollar cap.
READ POST ↗OpenRouter's $113M Series B Values Model Routing at $1.3B
OpenRouter raised a $113M Series B led by CapitalG at a reported $1.3B valuation, processing 100 trillion tokens monthly across 400+ models — why routing now matters.
READ POST ↗Interfaze's Hybrid Architecture Takes On Flash-Tier Models
JigsawStack's Interfaze hybrid model pairs specialized DNN/CNN blocks with an omni-transformer: 70.7% on OCRBench V2 vs 55.8% for Gemini-3-Flash, $1.50/M input, 1M context.
READ POST ↗Where the Goblins Came From: How Reward Signals Quietly Shape Model Behavior
OpenAI traced a 175% spike in 'goblin' mentions to a reward signal for the Nerdy personality. Learn how RL feedback loops spread quirks and what product builders can do.
READ POST ↗Before GPT-5.4: Reading OpenAI's Accelerating Release Cadence
Per the official release notes: GPT-5.2 shipped Dec 11, 2025; GPT-5.2-Codex on Jan 14, 2026; GPT-5.3-Codex on Feb 5 — three releases in under two months. What the faster pace means for API builders.
READ POST ↗OpenAI Retires GPT-4o From ChatGPT: The 0.1% Problem
OpenAI announced January 29, 2026: GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini leave ChatGPT on February 13. Only 0.1% of users still pick GPT-4o daily, most run GPT-5.2, and the API is unaffected.
READ POST ↗