2026
15 篇文章用 Amazon Bedrock prompt caching 把重複的 context 成本壓低 90%
Amazon Bedrock 的 prompt caching 讓重複輸入的 token 成本最多降 90%,同時縮短首字延遲,適合多輪問答與 agent 工作流。
閱讀文章 ↗Fable 5.1 的省錢關鍵不是模型,而是你的快取讀取比例
Firecrawl 實測 57 次 API 呼叫後發現,Fable 5.1 只有在快取讀取密集的長代理任務才便宜,其他情境反而更貴。
閱讀文章 ↗細模型的大優勢:企業如何用小語言模型省錢又高效
當企業面對眾多 AI 模型時,選擇哪一個才能兼顧效能與成本,成了關鍵課題。大型語言模型(LLM)常佔據新聞版面,但許多組織發現,較小、專用的小語言模型(SLM)反而能提供顯著優勢:更低的運算需求、更少的訓練資料、更省能源,以及更具成本效益的解決方案。
閱讀文章 ↗用 Zstandard 與 Pingora 節省 PB 級快取儲存:Cloudflare 的 Cache Transcoding 原型
Cloudflare 工程實習生打造 Cache Transcoding 原型,在 Pingora 快取內以 Zstandard 壓縮文字資產,用少量 CPU 換取可觀的儲存與跨資料中心頻寬節省。本文解析其取捨、設計與測試結果。
閱讀文章 ↗Copilot canvases:把 agentic workflow 從聊天捲動搬上可審視的畫布
GitHub 在 Copilot app 推出 canvases,讓開發者與 agent 在持久、共享的表面上協作。本文拆解兩個實戰 canvas 的設計細節、2,000 到 3,000 AI credits 的成本帳本,以及 builder 可以直接帶走的四步落地藍圖。
閱讀文章 ↗OpenAI 大幅降價 GPT-5.6 Luna 與 Terra,推出 Fast mode 加速 Sol
OpenAI 於 2026 年 7 月 30 日宣布 GPT-5.6 Luna 與 Terra 的 API 價格大幅調降,並推出 Fast mode 讓 Sol 的處理速度提升至 2.5 倍。本文整理價格變動、效率提升的技術背景,以及對產品開發者的實務意涵。
閱讀文章 ↗OpenAI 提出「每美元有用智慧」框架:四個維度衡量 AI 投資回報
OpenAI 提出「每美元有用智慧」衡量框架:從工作產出、任務成本、可靠性與規模效益四個維度取代採用率指標,說明為何每 token 成本會誤導選型,幫產品建構者與企業衡量 AI 投資的真實回報。
閱讀文章 ↗Baseten 傳以 130 億美元估值募 15 億美元:五個月估值跳 160%
WSJ 報導 Baseten 接近完成 15 億美元輪次,估值最高 130 億美元,距 1 月以 50 億美元估值完成的 3 億美元 Series E 僅五個月。本文拆解雙軌定價設計、開源推理路由生意,與推理淘金熱背後的毛利風險。
閱讀文章 ↗DeepSeek 把 V4 Pro 七五折降價常態化:前沿模型價格戰的底牌
2026 年 5 月 22 日,DeepSeek 宣布旗艦 V4 Pro 的 75% 折扣轉為常態價:輸出每百萬 token 0.87 美元,僅為原價四分之一,Flash 更只要 0.28 美元。本文解析降價常態化的成本邏輯與對 API 市場的連鎖影響。
閱讀文章 ↗資料中心用電狂潮推升燃氣電廠造價:兩年暴漲 66% 的代價
BNEF 報告指出,新建複循環燃氣電廠成本兩年內上漲 66%、每 kW 達 2,157 美元,施工期延長 23%,主因是資料中心需求。本文解析燃氣輪機短缺與燃氣、再生能源兩條路線之爭。
閱讀文章 ↗GPT-5.6 Luna 讓 Replit Free Mode 成真:模型價格效能如何打開軟體創作的大門
Replit 與 OpenAI 合作,以 GPT-5.6 Luna 驅動 Free Mode,讓數百萬用戶免費打造軟體。本文探討模型經濟學的轉變、產品設計的連續性,以及這對產品開發者的啟示。
閱讀文章 ↗GPT-5.4 mini 與 nano 登場:小模型接手免費層的主力位置
OpenAI 於 3 月 17 日推出 GPT-5.4 mini 與 nano,mini 自 3 月 18 日起進入 Free 與 Go 方案。距離 GPT-5.4 旗艦發表僅十二天,本文看小型版本在產品階層與成本結構中的角色,以及開發者的選型建議。
閱讀文章 ↗Meta 四款自研 MTIA 晶片登場:兩年四代的矽節奏
2026 年 3 月 11 日,Meta 發表 MTIA 300、400、450、500 四款自研 AI 晶片,數十萬顆已在生產環境運轉,並宣稱約每六個月推出新一代。本文解析逐代規格跳幅、chiplet 模組化設計,與 NVIDIA、Broadcom 之間的分層算力佈局。
閱讀文章 ↗Gemini 3.1 Flash Lite 預覽登場:Gemini 3 家族最快最便宜的模型
Google 於 3 月 3 日釋出 Gemini 3.1 Flash Lite 預覽版:輸入每百萬 token 0.25 美元、輸出 1.50 美元,是 Gemini 3 家族最快最便宜的選項。本文解析定價、基準與選型建議。
閱讀文章 ↗Microsoft Maia 200 登場:FP4 破 10 petaFLOPS 的自研推理晶片
2026 年 1 月 26 日,Microsoft 發表第二代自研 AI 加速器 Maia 200:FP4 算力超過 10 petaFLOPS、1,400 億顆以上電晶體、750W 封裝,宣稱 FP4 效能為 Amazon Trainium3 的三倍。本文解析規格對比與自研矽的成本帳。
閱讀文章 ↗
2026
14 ARTICLESPrompt Caching on Bedrock: Where the 90% Input Savings Actually Come From
Amazon Bedrock prompt caching cuts repeated-context input costs up to 90% and lowers TTFT, but only if you place cache points and TTLs deliberately.
READ POST ↗Fable 5.1's Cache Discount: Where the Bill Actually Moves
Fable 5.1 cut cache reads 75%, but 57 billed runs show the saving depends on how often your agent re-reads context.
READ POST ↗Small Language Models: The Enterprise Case for Right-Sizing AI
Cohere's guide to SLMs shows why smaller models can cut costs, run locally, and even beat larger ones on specific tasks. Learn how to build a model portfolio that matches size to job.
READ POST ↗How Cloudflare Could Save Petabytes of Cache Storage with Zstandard and Pingora
Cloudflare's Cache Transcoding prototype compresses cache entries with Zstandard inside Pingora, trading a small CPU increase for significant storage and bandwidth savings.
READ POST ↗Making Agentic Workflows Visible, Steerable, and Cost-Efficient with GitHub Copilot Canvases
Explore how GitHub Copilot canvases turn agentic workflows into durable, inspectable systems, with real examples and cost insights.
READ POST ↗OpenAI Slashes GPT-5.6 Luna and Terra Prices, Adds Fast Mode for Sol
OpenAI cuts GPT-5.6 Luna and Terra API prices, introduces Fast mode for Sol, and shares efficiency gains. Learn what it means for builders.
READ POST ↗Baseten's $1.5B Round Bets Big on Open-Source Inference
Baseten is reportedly raising $1.5B at up to a $13B valuation, five months after a $300M Series E at $5B. Inside the split-priced round and the open-source inference bet.
READ POST ↗DeepSeek Makes Its 75% V4 Pro Price Cut Permanent
DeepSeek confirmed May 22, 2026 its 75% V4 Pro discount is permanent: $0.87 per million output tokens, a quarter of list price; Flash is $0.28. What it says about API economics.
READ POST ↗Data Center Boom Pushes Gas Plant Costs Up 66% in Two Years
New combined-cycle gas plant costs jumped 66% in two years to $2,157/kW, per BNEF, with build times up 23%. Data center demand drives it, and turbine waitlists stretch years out.
READ POST ↗How GPT-5.6 Luna Powers Replit's Free Mode: Model Economics in Action
OpenAI's GPT-5.6 Luna now powers Replit Free Mode. Learn how price-performance shifts enable free tiers and what product builders can learn.
READ POST ↗GPT-5.4 mini and nano Arrive: Small Models Take Over the Free Tier
OpenAI released GPT-5.4 mini and nano on March 17; mini reaches Free and Go tiers from March 18, twelve days after the flagship. On small models in the product ladder and the cost equation.
READ POST ↗Meta Unveils Four MTIA Chips: A Six-Month Silicon Cadence
Meta announced four MTIA chips on March 11, 2026 — hundreds of thousands already in production, 25x FLOPS gains across generations, and a chiplet platform built for a roughly six-month cadence.
READ POST ↗Gemini 3.1 Flash Lite: Fastest, Cheapest Gemini 3 Model
Google shipped Gemini 3.1 Flash Lite in preview on March 3: $0.25 per million input tokens, up to 363 tokens per second, built for high-volume workloads. Pricing, benchmarks, and guidance.
READ POST ↗Microsoft's Maia 200: A 10-PetaFLOPS Bet on AI Inference
Microsoft's Maia 200, announced January 26, 2026, delivers 10+ petaFLOPS at FP4 in a 750W package and claims 3x Trainium3 FP4 throughput. Inside the specs and the custom-silicon cost math.
READ POST ↗