2026
23 篇文章當實驗數據多到看不完:SAM 3 與 DINOv3 在 Genesis Mission 裡的角色
美國國家實驗室把 Meta 開源視覺模型微調後部署到 300 張 A100,將一個月的標註工作壓到 15 分鐘。
閱讀文章 ↗用 Amazon Bedrock AgentCore 打造跨 WhatsApp 的點餐助理:文字、語音、通話一條龍
AWS 官方部落格分享如何用 Amazon Bedrock AgentCore 與 Amazon Nova 2 打造多模態 WhatsApp 點餐助理,整合文字、語音訊息與語音通話,並共用單一後端與跨頻道記憶。本文拆解架構決策與部署重點。
閱讀文章 ↗Ox Alpha 就是 GLM-5.3-Flash:MIT 開源,OpenRouter 週流量佔 31%
智譜揭曉匿名冠軍 Ox Alpha 就是 GLM-5.3-Flash 並以 MIT 授權開源:預覽期在十萬顆國產晶片上處理 62 兆 tokens,OpenRouter 週流量佔 31%、登頂程式碼模型榜,消息公布當日港股大漲逾 12%。
閱讀文章 ↗Qwen3.8-27B 開源釋出:體積小、看得懂圖,只是預設想太多
阿里巴巴 Qwen 兌現一週前承諾,釋出 Apache 2.0 的 Qwen3.8-27B 開放權重:262K 上下文、原生視覺、單機可跑;Simon Willison 實測稱讚能力,也點出 xhigh 推理預設讓簡單任務慢上數倍。
閱讀文章 ↗Mistral 開源 Shieldstral:把審查政策變成一句提問的 3B 分類器
Mistral 於 8 月 4 日發表 Shieldstral:3B 開源權重多模態安全分類器,政策以自然語言在推論時下達,單張 16GB GPU 即可運行,以 Apache 2.0 釋出。
閱讀文章 ↗Gemini Robotics 2 進場:從腳到指尖的機器人全身智能
2026 年 7 月 30 日 Google DeepMind 發表 Gemini Robotics 2,以 VLA、ER 與 On-Device 三個模型讓人形機器人全身控制、綁繩結、組隊合作,並公布 ASIMOV-Agentic 安全基準。
閱讀文章 ↗Interfaze 混合架構登場:自報基準碾壓 flash 級對手
JigsawStack 團隊推出混合架構模型 Interfaze,把專用神經網路併進 omni-transformer,主打 OCR、語音轉文字等確定性任務:OCRBench V2 70.7% 對 Gemini-3-Flash 55.8%。本文檢視九項自報基準與切換風險。
閱讀文章 ↗ChatGPT Images 2.0 的實用觀察:從精準控制到多語言排版
OpenAI 在 2026 年 4 月 21 日推出 ChatGPT Images 2.0,主打更精準的圖像控制、多語言文字排版與更真實的人物呈現。本文從產品建構者的角度整理重點與限制。
閱讀文章 ↗OpenRouter 影片生成上線:一個 API 路由所有影片模型
OpenRouter 把影片生成納入統一路由層:Seedance、Veo 3.1、Wan、Sora 2 Pro 走同一個 schema 與計費,非同步 job 模型加上 /api/v1/videos/models 能力探索端點。本文整理四大正規化設計、參數差異的地雷,以及 LLM prompt 接影片的多模態工作流。
閱讀文章 ↗Liquid AI 推出 LFM2.5-VL-450M:跑得進手機的 450M 視覺語言模型
Liquid AI 於 4 月 8 日發布 LFM2.5-VL-450M,預訓練從 10T 擴到 28T tokens,新增物件偵測與函式呼叫能力,量化後在 Jetson Orin 上每幀僅 233 毫秒,瞄準邊緣裝置上的多模態應用。
閱讀文章 ↗Gemma 4 開源發布:Apache 2.0、MoE 與 256K 上下文
2026 年 4 月 2 日 Google DeepMind 發布 Gemma 4:全系列改用 Apache 2.0 授權,四種尺寸從 2B 端側到 31B Dense,支援 140+ 語言、256K 上下文與影像語音輸入,31B 躍上 Arena 開源模型第三名。
閱讀文章 ↗Meta 開源 TRIBE v2:預測大腦如何回應影像、聲音與語言
Meta FAIR 釋出 TRIBE v2,以 700 多名受試者的 fMRI 資料訓練,能 zero-shot 預測新受試者、新語言與新任務的大腦反應,解析度較同類模型高約 70 倍,模型、程式碼與展示皆以 CC BY-NC 4.0 釋出。
閱讀文章 ↗Google Lyria 3 Pro 登場:AI 音樂生成從 30 秒片段走向三分鐘完整曲目
3 月 25 日 Google DeepMind 推出 Lyria 3 Pro,可生成長達 3 分鐘、含完整段落結構的曲目,較 Lyria 3 的 30 秒大幅躍進。本文看它進駐 Gemini、Vids、Vertex AI 與 ProducerAI 的布局、SynthID 浮水印與訓練資料立場。
閱讀文章 ↗小米 MiMo-V2 登場:兆級參數、百萬上下文與 87 億美元豪賭
小米 3 月 18 日發表 MiMo-V2-Pro(總參數逾 1 兆、啟動 42B、100 萬 token 上下文)、全模態 MiMo-V2-Omni 與 TTS 模型,隔日雷軍宣布三年至少投入 87 億美元。模型曾以 Hunter Alpha 匿名稱霸 OpenRouter。本文解析規格、開源策略與對開發者的影響。
閱讀文章 ↗MIT Wave-Former:用 Wi-Fi 訊號與生成式 AI 看穿牆壁重建 3D 物體
MIT Media Lab 團隊發表 Wave-Former,結合毫米波無線訊號與生成式補形,重建完全被遮蔽的日常物體 3D 形狀,準確度較現有最佳方法提升近 20%,論文將於 CVPR 2026 發表。本文解析其原理、訓練資料策略與機器人應用。
閱讀文章 ↗PixVerse 籌得 3 億美元 C 輪,成為亞洲 AI 影片獨角獸
2026 年 3 月 12 日,阿里巴巴支持的 AI 影片生成平台 PixVerse 完成 3 億美元 C 輪融資,估值突破 10 億美元,由 CDH Investments 領投,是亞洲 AI 影片類別最大一輪。本文解析其產品線、1.6 億月活背後的數據與新加坡全球佈局。
閱讀文章 ↗Rhoda AI 走出隱身:4.5 億美元 A 輪押注機器人基礎模型
2026 年 3 月 10 日,Rhoda AI 結束 18 個月隱身,宣布 4.5 億美元 A 輪與 17 億美元估值,並發表以數億支網路影片預訓練的機器人智慧模型 FutureVision,採 Direct Video Action 架構,鎖定工廠產線等非結構化環境。
閱讀文章 ↗Qwen3.5 小模型補齊戰線:9B 在多項基準超越 gpt-oss-120b
阿里巴巴 2 月底補上 Qwen3.5 的 0.8B–9B 四款小模型:Apache-2.0 開源、原生 262K 情境、支援影像影片輸入。9B 以 9.65B 密集參數在 MMLU-Pro、GPQA Diamond 超越 gpt-oss-120b,目標是手機與筆電上的本地多模態。
閱讀文章 ↗Gemini 3.1 Flash Lite 預覽登場:Gemini 3 家族最快最便宜的模型
Google 於 3 月 3 日釋出 Gemini 3.1 Flash Lite 預覽版:輸入每百萬 token 0.25 美元、輸出 1.50 美元,是 Gemini 3 家族最快最便宜的選項。本文解析定價、基準與選型建議。
閱讀文章 ↗阿里巴巴開源 Qwen3.5:397B 參數、201 種語言的 Agent 時代模型
2026 年 2 月 16 日除夕,阿里巴巴開源 Qwen3.5:首發 Qwen3.5-Plus 為 397B-A17B 稀疏 MoE 原生多模態模型,支援 201 種語言、內建可操作手機與電腦的視覺 agent,長情境解碼吞吐量達 Qwen3-Max 的 8.6 倍。本文解析規格、開源策略與中國 Agent 競賽。
閱讀文章 ↗Waymo World Model 登場:用 Genie 3 生成超擬真自駕模擬世界
2026 年 2 月 6 日,Waymo 發表建構在 Genie 3 之上的 World Model,同時生成相機影像與光達點雲,支援路線重模擬與罕見情境測試。本文解析三種控制機制、預訓練的槓桿,與對自駕安全驗證的意義。
閱讀文章 ↗Kimi K2.5 開源釋出:原生視覺加上 Agent Swarm 多代理協作
Moonshot AI 於 2026 年 1 月 27 日發表 Kimi K2.5:MIT 授權的開源 MoE model,具備原生視覺與 Agent Swarm 多代理協作。本文解析它對自架多模態 agent 堆疊與多代理系統生態的意義。
閱讀文章 ↗Samsung The First Look 2026:娛樂、家庭、照護三種 AI 同伴
CES 2026 開展前,Samsung 在 The First Look 發表「Your Companion to AI Living」:電視內建 Vision AI Companion、SmartThings 用戶突破 4.3 億、Family Hub 結合 Gemini,並把 AI 推向健康照護與保險。
閱讀文章 ↗
2026
23 ARTICLESSAM 3 and DINOv3 Cut Beamline Segmentation From a Month to 15 Minutes
Meta's open vision models let Berkeley Lab's SYNAPS-I label 3D beamline volumes in about 15 minutes instead of a month.
READ POST ↗Build a WhatsApp Ordering Assistant That Remembers You Across Text, Voice Notes, and Calls
AWS shows how to deploy a multimodal WhatsApp ordering assistant with Amazon Bedrock AgentCore, Amazon Nova 2, and MCP.
READ POST ↗Ox Alpha Was GLM-5.3-Flash: 31% of OpenRouter Traffic
The anonymous Ox Alpha that topped OpenRouter is GLM-5.3-Flash: 62T tokens on 100,000 Chinese chips, 31% of weekly traffic, MIT-licensed.
READ POST ↗Qwen3.8-27B: Great Open Weights That Overthink by Default
Alibaba's Qwen shipped the Apache 2.0 Qwen3.8-27B a week after promising it: 262K context, native vision, laptop-class local AI — behind a costly xhigh reasoning default.
READ POST ↗Mistral's Shieldstral: A 3B Open-Weights Safety Classifier
Mistral's Shieldstral is a 3B open-weights multimodal safety classifier that takes policies as plain-language questions at inference time and runs on one 16GB GPU.
READ POST ↗Gemini Robotics 2 Gives Robots Whole-Body Intelligence
Google DeepMind's Gemini Robotics 2 adds whole-body humanoid control, fine dexterity, multi-robot teamwork, and an ASIMOV-Agentic safety benchmark to its robotics stack.
READ POST ↗Interfaze's Hybrid Architecture Takes On Flash-Tier Models
JigsawStack's Interfaze hybrid model pairs specialized DNN/CNN blocks with an omni-transformer: 70.7% on OCRBench V2 vs 55.8% for Gemini-3-Flash, $1.50/M input, 1M context.
READ POST ↗ChatGPT Images 2.0: What Builders Should Know About Text, Multilingual Layouts, and Realism
OpenAI's ChatGPT Images 2.0 brings better text rendering, multilingual support, and realistic styles. Here's what product builders need to know.
READ POST ↗Video Generation Is Live on OpenRouter: One API to Route Every Video Model
OpenRouter brings video into its unified routing layer: Seedance, Veo 3.1, Wan, and Sora 2 Pro behind one schema and bill, with async jobs and a capability endpoint built for coding agents.
READ POST ↗Liquid AI Ships LFM2.5-VL-450M, a 450M Edge VLM
Liquid AI released LFM2.5-VL-450M on April 8: 28T-token pretraining, new detection and function-calling skills, and 233 ms per frame on a Jetson Orin once quantized.
READ POST ↗Gemma 4 Ships Under Apache 2.0: Google's Open Model Reset
Google DeepMind released Gemma 4 on April 2, 2026 under Apache 2.0 — four sizes from 2B edge to 31B dense, 140+ languages, 256K context, Arena top-3. What it changes for builders.
READ POST ↗Meta's TRIBE v2 Predicts Brain Responses Like a Digital Twin
Meta FAIR's TRIBE v2, trained on fMRI from 700+ volunteers, zero-shot predicts brain responses for new subjects and languages at 70x the resolution of similar models.
READ POST ↗Google Lyria 3 Pro: AI Music Moves From Clips to Full Tracks
Google DeepMind's Lyria 3 Pro generates full three-minute songs with verse-chorus structure, rolling into Gemini, Vids, Vertex AI, and ProducerAI with SynthID watermarking.
READ POST ↗Xiaomi MiMo-V2: Trillion-Parameter Models, $8.7B AI Push
Xiaomi launched MiMo-V2-Pro (1T+ params, 42B active, 1M context) plus Omni and TTS models, and pledged $8.7B over three years. It had topped OpenRouter as anonymous 'Hunter Alpha'.
READ POST ↗MIT Wave-Former: Seeing Through Walls With Wi-Fi and AI
MIT's Wave-Former pairs millimeter-wave signals with generative shape completion to rebuild fully occluded 3D objects at nearly 20% higher accuracy. CVPR 2026 paper; companion RISE extends to rooms.
READ POST ↗PixVerse Raises $300M Series C as an AI Video Unicorn
On March 12, 2026, Alibaba-backed PixVerse closed a $300M Series C led by CDH Investments at a $1B+ valuation — Asia's largest AI video round — and opened a global office in Singapore.
READ POST ↗Rhoda AI Exits Stealth With $450M for Robot Intelligence
Rhoda AI left 18 months of stealth on March 10, 2026 with a $450M Series A at a $1.7B valuation and FutureVision, a robot foundation model pre-trained on hundreds of millions of internet videos.
READ POST ↗Qwen3.5 Small Models: 9B Rivals gpt-oss-120b at the Edge
Alibaba adds 0.8B-9B small models to Qwen3.5 under Apache-2.0: native 262K context, image and video input. The 9B beats gpt-oss-120b on MMLU-Pro and GPQA Diamond, targeting phones and laptops.
READ POST ↗Gemini 3.1 Flash Lite: Fastest, Cheapest Gemini 3 Model
Google shipped Gemini 3.1 Flash Lite in preview on March 3: $0.25 per million input tokens, up to 363 tokens per second, built for high-volume workloads. Pricing, benchmarks, and guidance.
READ POST ↗Alibaba Open-Sources Qwen3.5 for the Agentic AI Era
Alibaba open-sourced Qwen3.5 on Lunar New Year's Eve: a multimodal 397B-A17B MoE with 201 languages, 8.6x Qwen3-Max decoding throughput, and visual agents that operate phones and desktops.
READ POST ↗Waymo's World Model: Genie 3 Powers Driving Simulation
Waymo's World Model, built on Google DeepMind's Genie 3, generates camera and lidar data for hyper-realistic driving simulation — re-simulating routes and testing rare events at scale.
READ POST ↗Kimi K2.5 Goes Open Source: Native Vision and Agent Swarm Coordination
Moonshot AI released Kimi K2.5 on January 27, 2026: an MIT-licensed open-source MoE model with native vision and Agent Swarm coordination. What it means for self-hosted agent stacks.
READ POST ↗Samsung Unveils AI Companions for TV, Home, Health at CES
Before CES 2026, Samsung's The First Look unveiled 'Your Companion to AI Living': Vision AI Companion on TVs, SmartThings past 430M users, Gemini-powered Family Hub, and a push into health.
READ POST ↗