2026
18 篇文章把 30 秒長鏡頭寫進 API:Seedance 2.5 的規格取捨與計費邏輯
Seedance 2.5 支援 30 秒單次生成與影片參考輸入,但解析度上限只有 720p,計費依影片 token 而非秒數。
閱讀文章 ↗MiniMax H3 開源:一次生成 2K 影片與立體聲的全模態模型
MiniMax 於 7 月 31 日發表 Hailuo 後繼者 H3,數日後釋出權重:33B 全模態 Transformer 同時生成最長 15 秒、768p 起步 2K 的影片與原生立體聲,ComfyUI Day-0 支援,量化後 RTX 3060 也能本地運行。
閱讀文章 ↗Seedance 2.5 登場:30 秒一鏡到底,一次參考 30 張圖與 10 段音訊
ByteDance Seed 團隊 7 月 31 日發表影片生成模型 Seedance 2.5:單次生成 30 秒、可多輪延伸成多分鐘成片,一次可引用 30 張圖、10 段影片與 10 段音訊,已上架即夢與豆包。
閱讀文章 ↗FLUX 3 登場:影片、圖像、聲音一個骨幹全包
Black Forest Labs 發表多模態基礎模型 FLUX 3:最長 20 秒含原生音訊的影片生成、多語言文字渲染,並推出以 FLUX 3 為骨幹的機器人動作模型 FLUX-mimic。
閱讀文章 ↗DeepMind GenCeption:影片生成模型就是通用視覺學習器
Google DeepMind 提出 GenCeption,把文字轉影片擴散模型改造成可用文字指令驅動的通用前饋感知模型,在深度、分割、相機姿態等任務達到 SOTA,所需訓練數據最多減少 500 倍。
閱讀文章 ↗Character.AI 進軍微短劇:c.ai Series 原創劇集登場
Character.AI 推出 c.ai Series 自製微短劇格式,首波三齣劇集登場,18 歲以上觀眾能直接與劇中角色對話、改寫劇情,把核心聊天產品接進影集,成為微短劇市場最新的互動敘事玩家。
閱讀文章 ↗Nano Banana 2 Lite 登場:4 秒生圖、單張 3.4 美分的高速影像路線
2026 年 6 月 30 日,Google 發表 Nano Banana 2 Lite:約 4 秒完成文字生圖、每張 0.034 美元,同場推出 Gemini Omni Flash 影片模型,用「快出圖、再動圖」的串接工作流把低成本高速生成推向主流。
閱讀文章 ↗Adobe 收購 Topaz Labs:裝置端 AI 影像增強補強 Firefly
Adobe 於 2026 年 6 月 25 日宣布收購影像增強公司 Topaz Labs,交易預計下半年完成。Topaz 的 Astra、Wonder 模型與 Neurostream 本地運算技術將整合進 Firefly 與創作套件,把大型視訊模型搬進消費級 GPU。
閱讀文章 ↗Decart Oasis 3 世界模型開放 API:即時生成三鏡頭駕駛世界
2026 年 6 月 10 日,Decart 推出首個以 API 開放的互動世界模型 Oasis 3:22 FPS 三鏡頭駕駛場景、每秒 0.02 美元,背後是 3 億美元新輪募資;實測仍見主題漂移與物理失真。
閱讀文章 ↗OpenRouter 影片生成上線:一個 API 路由所有影片模型
OpenRouter 把影片生成納入統一路由層:Seedance、Veo 3.1、Wan、Sora 2 Pro 走同一個 schema 與計費,非同步 job 模型加上 /api/v1/videos/models 能力探索端點。本文整理四大正規化設計、參數差異的地雷,以及 LLM prompt 接影片的多模態工作流。
閱讀文章 ↗Sora 1 美國下架:OpenAI 把影片生成收進 ChatGPT
2026 年 3 月 13 日,Sora 1 在美國正式下架,未匯出的舊內容一併刪除;同週 The Information 報導 OpenAI 計畫把 Sora 影片生成搬進 ChatGPT。本文解析這波產品整併的商業邏輯,以及創作者該記住的資料風險教訓。
閱讀文章 ↗PixVerse 籌得 3 億美元 C 輪,成為亞洲 AI 影片獨角獸
2026 年 3 月 12 日,阿里巴巴支持的 AI 影片生成平台 PixVerse 完成 3 億美元 C 輪融資,估值突破 10 億美元,由 CDH Investments 領投,是亞洲 AI 影片類別最大一輪。本文解析其產品線、1.6 億月活背後的數據與新加坡全球佈局。
閱讀文章 ↗ByteDance Seedance 2.0 登場:音視訊統一生成,15 秒多鏡頭立體聲
2026 年 2 月 12 日,ByteDance Seed 團隊發布 Seedance 2.0,以統一多模態架構一次生成 15 秒多鏡頭影片與同步立體聲音軌,支援文字、圖片、音訊、影片四種輸入混合,先在中國上線,再經第三方平台走向海外。
閱讀文章 ↗Google Project Genie 上線:世界模型即時生成可玩世界,遊戲股應聲重挫
Google 於 1 月 29 日向 AI Ultra 訂閱者開放 Project Genie,首個基於 Genie 3 世界模型的公開產品,可從文字與圖片即時生成可互動 3D 世界。次日 Unity 一度暴跌近 24%,遊戲股全面下挫。
閱讀文章 ↗xAI 開放 Grok Imagine API:影片生成原生帶音訊
xAI 於 2026 年 1 月 28 日推出 Grok Imagine API,提供原生音訊的影片生成,並宣稱在品質、成本與延遲上達到 state-of-the-art。本文解析影片生成 API 市場的競爭格局與開發者的評估重點。
閱讀文章 ↗Runway「The Turing Reel」實測:僅 9.5% 的人能穩定看穿 AI 影片
Runway 用 Gen-4.5 生成影片與真實素材做對照測試,1,043 位受測者的整體正確率只有 57.1%,僅 9.5% 達統計顯著。本文解析測試方法、錯誤分佈,以及為何 Runway 呼籲放棄肉眼偵測、改走 C2PA 溯源路線。
閱讀文章 ↗YouTube 執行長年度信:打擊 AI slop 與深偽列為 2026 首要任務
2026 年 1 月 21 日,YouTube 執行長 Neal Mohan 發表年度信,把打擊 AI slop 與深偽偵測列為年度優先,同時預告 AI 分身 Shorts 與文字生遊戲。一手緊縮、一手擴張的平台 AI 策略解析。
閱讀文章 ↗Overworld 開源 Waypoint-1:鍵盤滑鼠即時操控的擴散世界模型
2026 年 1 月 20 日,Overworld 在 Hugging Face 發布 Waypoint-1:以 10,000 小時遊戲影片訓練的即時互動世界模型,2.3B 開源,RTX 5090 上 30 FPS。文中解析 rectified flow 架構、WorldEngine 推理優化與世界模型離實用的距離。
閱讀文章 ↗
2025
2 篇文章Midjourney V1 影片模型問世:圖生影片最長約 21 秒
2025 年 6 月 18 日,Midjourney 推出首款影片生成模型 V1,可將靜態圖片轉為 5 秒短片,最長延長至約 21 秒,定價約為圖片生成的 8 倍。本文整理功能、用量限制、定價,以及它在 Veo 3 與 Sora 之間的位置。
閱讀文章 ↗Veo 3開放Gemini Pro訂閱戶試用,AI Ultra推進73國
2025年5月30日Google擴大Veo 3供應範圍:Gemini應用程式的AI Pro訂閱戶可試用這款含音訊的影片生成模型,AI Ultra方案亦推進至73國;DeepMind執行長Hassabis稱上線數日已生成數百萬支影片。
閱讀文章 ↗
2026
19 ARTICLESSeedance 2.5: Long Takes, Reference Inputs, and the Per-Second Bill
Seedance 2.5 trades 4K for 30-second takes and cheaper video-reference billing — what that changes for clip pipelines.
READ POST ↗MiniMax H3 Goes Open With 2K Video and Native Stereo Audio
MiniMax announced H3 on July 31 and released the weights days later: a 33B omni-modal Transformer that generates up to 15 seconds of 2K video with native stereo audio.
READ POST ↗Seedance 2.5 Pushes AI Video to 30-Second One-Takes
ByteDance's Seed team launched Seedance 2.5: 30-second single-pass video, multi-round extension into minutes, and up to 30 images, 10 videos, and 10 audio clips as references per run.
READ POST ↗FLUX 3: One Backbone for Video, Images, Audio — and Robots
Black Forest Labs launched FLUX 3, a multimodal model with 20-second native-audio video, plus FLUX-mimic, a robot action model built on the same backbone.
READ POST ↗DeepMind GenCeption: Video Generation as Vision Pretraining
DeepMind's GenCeption turns a text-to-video diffusion backbone into a text-steered perception model, hitting SOTA on depth, segmentation, and pose with up to 500x less data.
READ POST ↗Character.AI Enters Microdrama With c.ai Series Shows
Character.AI launches c.ai Series, three microdramas where viewers 18+ can chat with the characters and roleplay alternate storylines, wiring its core chat product into shows.
READ POST ↗Google's Nano Banana 2 Lite: 4-Second, 3.4-Cent Images
On June 30, 2026, Google launched Nano Banana 2 Lite — images in about 4 seconds at $0.034 each — plus the Gemini Omni Flash video model, pairing fast stills with cheap video.
READ POST ↗Adobe Acquires Topaz Labs to Bring On-Device AI to Firefly
Adobe is acquiring Topaz Labs, whose Astra and Wonder models and Neurostream on-device tech will join Firefly and the creative apps; the deal closes in H2 2026.
READ POST ↗Runway Agent 2.0: From AI Video Generator to Marketing Operations Partner
Runway's Agent 2.0 bridges the gap between analyzing ad performance and creating new assets—in one conversation. Built for marketers, not just creators.
READ POST ↗Decart Oasis 3: An API-Served World Model for AV Training
Decart ships Oasis 3, the first world model callable by API: 22 FPS three-camera driving scenes at $0.02/sec, a $300M raise behind it, and real caveats in hands-on testing.
READ POST ↗Video Generation Is Live on OpenRouter: One API to Route Every Video Model
OpenRouter brings video into its unified routing layer: Seedance, Veo 3.1, Wan, and Sora 2 Pro behind one schema and bill, with async jobs and a capability endpoint built for coding agents.
READ POST ↗Sora 1 Goes Dark in the US as Video Moves Into ChatGPT
On March 13, 2026, OpenAI retired Sora 1 for US users and deleted unexported content; reports say Sora video is heading into ChatGPT. The consolidation logic and the creator data lesson.
READ POST ↗PixVerse Raises $300M Series C as an AI Video Unicorn
On March 12, 2026, Alibaba-backed PixVerse closed a $300M Series C led by CDH Investments at a $1B+ valuation — Asia's largest AI video round — and opened a global office in Singapore.
READ POST ↗Seedance 2.0: ByteDance's Unified Audio-Video Model Ships
ByteDance's Seed team launched Seedance 2.0 on Feb 12, 2026: one architecture turns text, images, audio, and video references into 15-second multi-shot video with synced stereo sound. China first.
READ POST ↗Project Genie: Google's World Model Rattles Gaming Stocks
Google opened Project Genie to AI Ultra users on Jan. 29 — the first public product built on Genie 3, turning prompts into playable 3D worlds in real time. Unity plunged nearly 24% the next day.
READ POST ↗xAI Opens the Grok Imagine API: Video Generation with Native Audio
xAI launched the Grok Imagine API on January 28, 2026: video generation with native audio, claiming state-of-the-art quality, cost, and latency. What matters for generative video teams.
READ POST ↗Runway's Turing Reel: Only 9.5% Can Reliably Spot AI Video
Runway tested 1,043 viewers on real footage versus Gen-4.5 image-to-video clips: 57.1% overall accuracy, only 9.5% statistically significant. Detection is dead — provenance metadata is the plan.
READ POST ↗YouTube CEO Makes AI Slop and Deepfakes a 2026 Priority
Neal Mohan's 2026 letter names AI slop reduction and deepfake detection as YouTube's top priorities — while teasing likeness Shorts and AI games. Tightening and expanding AI at once.
READ POST ↗Overworld Open-Sources Waypoint-1, a Real-Time World Model
Overworld open-sourced Waypoint-1, a real-time interactive video diffusion world model trained on 10,000 hours of gameplay, running at 30 FPS on an RTX 5090 via its WorldEngine stack.
READ POST ↗
2025
2 ARTICLESMidjourney launches V1, its first AI video model
Midjourney shipped V1, its first video model, on June 18, 2025: image-to-video clips up to about 21 seconds, priced at roughly eight times image generations. How it works, and what it costs.
READ POST ↗Veo 3 opens to Gemini Pro users, AI Ultra in 73 countries
On May 30, 2025, Google opened Veo 3 trials to Gemini AI Pro subscribers and took the AI Ultra plan to 73 countries, days after users generated millions of videos with the audio-capable model.
READ POST ↗