2026
24 篇文章Muse Spark 把「想久一點」變成可調參數:多代理推理與思考時間懲罰的取捨
Meta 發表 Muse Spark,以思考時間懲罰與多代理並行推理控制延遲,並開放 meta.ai 與私有 API 預覽。
閱讀文章 ↗把視覺模型塞進輪椅之後:RAMMP 專案揭露的邊緣部署取捨
Meta 的 DINO 與 SAM 被用於匹茲堡大學 RAMMP 輔助行動平台,重點不是模型多強,而是邊緣裝置上的精度與即時性取捨。
閱讀文章 ↗AMIE 的下一步:AI 醫療從文字走向即時視訊問診
Google Research 與 Google DeepMind 在 2026 年 8 月發表 AMIE 的視訊問診研究,展示 AI 如何透過多代理架構理解視覺與聽覺線索。本文從產品建造者角度,拆解這項技術對醫療 AI 產品設計的啟示與限制。
閱讀文章 ↗MiniMax H3 開源:一次生成 2K 影片與立體聲的全模態模型
MiniMax 於 7 月 31 日發表 Hailuo 後繼者 H3,數日後釋出權重:33B 全模態 Transformer 同時生成最長 15 秒、768p 起步 2K 的影片與原生立體聲,ComfyUI Day-0 支援,量化後 RTX 3060 也能本地運行。
閱讀文章 ↗Seedance 2.5 登場:30 秒一鏡到底,一次參考 30 張圖與 10 段音訊
ByteDance Seed 團隊 7 月 31 日發表影片生成模型 Seedance 2.5:單次生成 30 秒、可多輪延伸成多分鐘成片,一次可引用 30 張圖、10 段影片與 10 段音訊,已上架即夢與豆包。
閱讀文章 ↗FLUX 3 登場:影片、圖像、聲音一個骨幹全包
Black Forest Labs 發表多模態基礎模型 FLUX 3:最長 20 秒含原生音訊的影片生成、多語言文字渲染,並推出以 FLUX 3 為骨幹的機器人動作模型 FLUX-mimic。
閱讀文章 ↗Kimi K3 登場:2.8T 參數、百萬 Token Context,開源模型走向長程 Agent
Moonshot 發布 Kimi K3:2.8T 參數的開源 3T 級模型,原生視覺、1M token context,主打長程 coding 與知識工作。本文整理 Delta Attention 架構、kernel 編譯器與晶片設計等案例、API 定價與 64 卡部署建議,以及官方自認仍落後 Fable 5 與 GPT 5.6 Sol 的誠實定位。
閱讀文章 ↗DeepMind GenCeption:影片生成模型就是通用視覺學習器
Google DeepMind 提出 GenCeption,把文字轉影片擴散模型改造成可用文字指令驅動的通用前饋感知模型,在深度、分割、相機姿態等任務達到 SOTA,所需訓練數據最多減少 500 倍。
閱讀文章 ↗Meta Brain2Qwerty v2:腦波轉文字達 61%
Meta 發表 Brain2Qwerty v2:以腦磁圖搭配端到端深度學習與微調語言模型,將非侵入式腦波解碼文字的字準確率從過去的 8% 推升至 61%,並開源 v1 與 v2 的訓練程式碼。
閱讀文章 ↗OpenRouter Unified Image API:統一請求之前,先讓程式看得懂模型差異
OpenRouter 推出專屬 Image API:單一請求格式接 30 多個模型,能力、參數與計價全部變成可查詢的 schema。本文整理 capability descriptors、三種計費單位、streaming 支援,以及官方的 migration 建議。
閱讀文章 ↗Snap Specs 開賣:2,195 美元的 on-device AR 眼鏡豪賭
2026 年 6 月 16 日,Snap 正式發表 Specs AR 眼鏡:2,195 美元、雙 Snapdragon 處理器、51 度視野、運算全在眼鏡上完成。本文整理規格與情境式 AI 功能,並看發表後股價下跌逾 5% 的市場反應。
閱讀文章 ↗WWDC26:Siri AI 改版、新一代 Apple Intelligence 與 iOS 27 一次揭曉
2026 年 6 月 8 日,Apple 在 WWDC26 主題演講發表 Siri AI 改版、新一代 Apple Intelligence 與 iOS 27。本文看蘋果 AI 的更新節奏、助理路線的產品邏輯,與對使用者的實際意義。
閱讀文章 ↗Thinking Machines 互動模型:把即時對話訓練進模型本體
Thinking Machines Lab 發表「互動模型」研究預覽:276B 參數 MoE、12B 啟動,以 200 毫秒微回合實現全雙工語音視訊互動,FD-bench 77.8 領先 GPT-realtime-2.0 與 Gemini Live,並由背景模型補足推理與工具呼叫。
閱讀文章 ↗Google Health Coach 5 月 19 日登場:Gemini 教練與 Fitbit 品牌重生
Google 宣布 Gemini 驅動的 Health Coach 於 2026 年 5 月 19 日全球上線,Fitbit 應用同步更名 Google Health,Health Premium 每月 9.99 美元,AI Pro 與 Ultra 訂戶免費使用,無螢幕手環 Fitbit Air 同日開賣。
閱讀文章 ↗微軟 MAI 三模型上線 Foundry:轉錄、語音與影像
2026 年 4 月 2 日,微軟把自研的 MAI-Transcribe-1、MAI-Voice-1、MAI-Image-2 送上 Microsoft Foundry 公開預覽:轉錄 GPU 成本砍半、一秒生成一分鐘語音、影像模型首發衝上 Arena.ai 第三名。
閱讀文章 ↗Qwen3.5-Omni 登場:原生全模態、聽 113 種語言、說 36 種
阿里巴巴 Qwen 團隊推出原生全模態模型 Qwen3.5-Omni,單一管線處理文字、影像、音訊與視訊並即時說話,256K 上下文、可聽逾 10 小時音訊,Plus 版宣稱在一般音訊任務超越 Gemini 3.1 Pro。
閱讀文章 ↗Gemini Embedding 2 公開預覽:文字、影像、音訊、影片共用一個向量空間
Google DeepMind 於 2026 年 3 月 10 日推出 Gemini Embedding 2 公開預覽:第一個原生多模態嵌入模型,把文字、影像、影片、音訊與 PDF 映射到單一 3072 維向量空間,支援 Matryoshka 截斷。本文解析輸入限制、成本槓桿與對 RAG 管線的實際影響。
閱讀文章 ↗World Labs 募資 10 億美元:李飛飛的空間智慧賭注升級
2026 年 2 月 18 日,李飛飛創辦的 World Labs 宣布募資 10 億美元,NVIDIA、AMD、Autodesk、Fidelity 等參投,Reuters 報導估值約 50 億美元。資金將投入世界模型與 Marble 3D 生成產品,把「空間智慧」從研究概念推向可出貨的產品線。
閱讀文章 ↗ByteDance Seedance 2.0 登場:音視訊統一生成,15 秒多鏡頭立體聲
2026 年 2 月 12 日,ByteDance Seed 團隊發布 Seedance 2.0,以統一多模態架構一次生成 15 秒多鏡頭影片與同步立體聲音軌,支援文字、圖片、音訊、影片四種輸入混合,先在中國上線,再經第三方平台走向海外。
閱讀文章 ↗Google Project Genie 上線:世界模型即時生成可玩世界,遊戲股應聲重挫
Google 於 1 月 29 日向 AI Ultra 訂閱者開放 Project Genie,首個基於 Genie 3 世界模型的公開產品,可從文字與圖片即時生成可互動 3D 世界。次日 Unity 一度暴跌近 24%,遊戲股全面下挫。
閱讀文章 ↗Meta 超級智能實驗室半年交出首批模型:Bosworth 的達沃斯進度報告
2026 年 1 月 21 日,Meta CTO Bosworth 在達沃斯證實,重組後的超級智能實驗室已於本月內部交付首批模型,評價「非常好」,但後訓練仍有大量工作。本文整理記者會要點、眼鏡與神經腕帶的載具策略,以及對 2026–2027 消費 AI 定型期的判斷。
閱讀文章 ↗Deepgram 募得 1.3 億美元、估值 13 億,收購 OfOne 進軍語音點餐
2026 年 1 月 13 日,語音 AI 公司 Deepgram 以 13 億美元估值完成 1.3 億美元 C 輪,AVP 領投,並收購 YC 培育的 OfOne。現金流已轉正的公司為何募資?語音 API 市場的競局又將如何變化?
閱讀文章 ↗Boston Dynamics 量產版 Atlas 現身 CES:全電動人形機器人開始出貨
CES 2026 上 Boston Dynamics 發表全電動量產版 Atlas:56 自由度、可舉 50 公斤、自主換電池,2026 年部署名額全數給了現代汽車與 Google DeepMind,2027 年初才開放外部客戶。人形機器人從示範影片走進工廠。
閱讀文章 ↗CES 2026 人形機器人日:Atlas 首度公開亮相與量產時間表
CES 2026 展場首日,Boston Dynamics 首度公開展示 Atlas,Hyundai 規劃數萬台工廠部署,Qualcomm 推出 Dragonwing IQ10,Mobileye 以 9 億美元收購 Mentee。人形機器人從展示走向部署,本文整理關鍵事實與現實檢查。
閱讀文章 ↗
2026
23 ARTICLESMuse Spark's Real Bet: Cheaper Pre-Training, Parallel Thinking at Inference
Meta's Muse Spark claims order-of-magnitude pre-training efficiency and a Contemplating mode that scales agents, not latency.
READ POST ↗Running Vision Models On-Device: What RAMMP Changes for Assistive Robotics Builders
Meta's DINO and SAM models move onto battery-powered assistive robots, trading precision for real-time reliability.
READ POST ↗AMIE: Google's Medical AI Moves from Text to Real-Time Video Consultations
Google Research and DeepMind advance AMIE to real-time video consultations, using multi-agent architecture to interpret visual and auditory cues in simulated clinical settings.
READ POST ↗MiniMax H3 Goes Open With 2K Video and Native Stereo Audio
MiniMax announced H3 on July 31 and released the weights days later: a 33B omni-modal Transformer that generates up to 15 seconds of 2K video with native stereo audio.
READ POST ↗Seedance 2.5 Pushes AI Video to 30-Second One-Takes
ByteDance's Seed team launched Seedance 2.5: 30-second single-pass video, multi-round extension into minutes, and up to 30 images, 10 videos, and 10 audio clips as references per run.
READ POST ↗FLUX 3: One Backbone for Video, Images, Audio — and Robots
Black Forest Labs launched FLUX 3, a multimodal model with 20-second native-audio video, plus FLUX-mimic, a robot action model built on the same backbone.
READ POST ↗Kimi K3 Arrives: 2.8T Parameters, Million-Token Context, Open Models Go Long-Horizon
Kimi K3: a 2.8T open 3T-class model with native vision and 1M-token context for long-horizon coding. The architecture, kernel and chip cases, pricing, and the gap to Fable 5 and GPT 5.6 Sol.
READ POST ↗DeepMind GenCeption: Video Generation as Vision Pretraining
DeepMind's GenCeption turns a text-to-video diffusion backbone into a text-steered perception model, hitting SOTA on depth, segmentation, and pose with up to 500x less data.
READ POST ↗Meta Brain2Qwerty v2: 61% Word Accuracy Without Surgery
Meta's Brain2Qwerty v2 decodes sentences from non-invasive MEG recordings at 61% word accuracy, up from 8% for prior methods, using end-to-end deep learning and fine-tuned LLMs.
READ POST ↗Snap Ships Specs: A $2,195 On-Device AR Glasses Gamble
Snap's Specs AR glasses open for preorder at $2,195 with dual Snapdragon chips, a 51-degree FOV, and fully on-device compute. Shares fell more than 5% after the Long Beach debut.
READ POST ↗WWDC26: Apple Unveils the Siri AI Revamp, Next-Gen Apple Intelligence, and iOS 27
At its WWDC26 keynote on June 8, 2026, Apple unveiled the Siri AI revamp, the next generation of Apple Intelligence, and iOS 27. A look at Apple's AI cadence and its assistant-first bet.
READ POST ↗Thinking Machines' Interaction Models Go Real-Time
Thinking Machines Lab previews interaction models: a 276B MoE, 12B active, trained for full-duplex speech and video via 200ms micro-turns. FD-bench 77.8, 0.40s turn latency.
READ POST ↗Google Health Coach: Gemini Coaching and the End of Fitbit
Google's Gemini-powered Health Coach goes global May 19 at $9.99/month, free for AI Pro and Ultra subs, as the Fitbit app becomes Google Health and Fitbit Air ships same day.
READ POST ↗Microsoft Opens Its MAI Audio and Image Models in Foundry
On April 2, 2026, Microsoft put three first-party MAI models into Foundry preview: half-cost transcription, a minute of speech in under a second, and a top-three image model.
READ POST ↗Qwen3.5-Omni: Alibaba's Omnimodal Model Speaks 36 Languages
Alibaba's Qwen3.5-Omni handles text, image, audio and video in one native pipeline with real-time speech out: 256K context, 113 recognition languages, three variants led by Plus.
READ POST ↗Gemini Embedding 2: One Vector Space for Every Modality
Google's Gemini Embedding 2 hits Public Preview on March 10, 2026 — the first natively multimodal embedding model, mapping text, images, video, audio, and PDFs into one 3072-dim vector space.
READ POST ↗World Labs Raises $1 Billion to Scale Spatial Intelligence
World Labs, Fei-Fei Li's spatial intelligence startup, raised $1B on Feb 18, 2026 from NVIDIA, AMD, Autodesk and Fidelity at a reported $5B valuation, to scale world models and its Marble 3D product.
READ POST ↗Seedance 2.0: ByteDance's Unified Audio-Video Model Ships
ByteDance's Seed team launched Seedance 2.0 on Feb 12, 2026: one architecture turns text, images, audio, and video references into 15-second multi-shot video with synced stereo sound. China first.
READ POST ↗Project Genie: Google's World Model Rattles Gaming Stocks
Google opened Project Genie to AI Ultra users on Jan. 29 — the first public product built on Genie 3, turning prompts into playable 3D worlds in real time. Unity plunged nearly 24% the next day.
READ POST ↗Meta's Superintelligence Labs Ship First Models From Davos
At Davos, Meta CTO Bosworth said the Superintelligence Labs delivered first key models internally this month, six months in. Heavy post-training remains; wearables are the vehicle.
READ POST ↗Deepgram Raises $130M at $1.3B, Buys OfOne
Deepgram raised a $130M Series C at a $1.3B valuation led by AVP and bought YC-backed OfOne. A cashflow-positive voice AI platform with 1,300+ customers, in a voice market headed to $14-20B by 2030.
READ POST ↗Boston Dynamics Electric Atlas Goes Into Production at CES
At CES 2026, Boston Dynamics unveiled the production all-electric Atlas: 56 DoF, 50 kg payload, self-swapping batteries — every 2026 deployment already committed to Hyundai and Google DeepMind.
READ POST ↗Humanoids Take CES 2026: Atlas Goes Public With Ship Dates
On CES 2026's first show day, Boston Dynamics unveiled Atlas, Hyundai lined up tens of thousands of factory robots, and Mobileye paid $900M for Mentee. The facts — and a reality check.
READ POST ↗