2026
10 篇文章把商品標籤交給小模型:SageMaker serverless 客製化的取捨
AWS 示範用 SageMaker serverless 客製化 Qwen3-8B 做商品標籤,重點在把訓練與推論的責任拆開。
閱讀文章 ↗把電子郵件拆成三十個小模型:Fyxer 如何讓 AI 助理值得信賴
Fyxer 用 30–50 個專門模型拆分郵件工作,並以使用者編輯回饋做 DPO 訓練,讓 AI 草稿接受率達 53%。
閱讀文章 ↗細模型的大優勢:企業如何用小語言模型省錢又高效
當企業面對眾多 AI 模型時,選擇哪一個才能兼顧效能與成本,成了關鍵課題。大型語言模型(LLM)常佔據新聞版面,但許多組織發現,較小、專用的小語言模型(SLM)反而能提供顯著優勢:更低的運算需求、更少的訓練資料、更省能源,以及更具成本效益的解決方案。
閱讀文章 ↗Mistral Forge:不微調、不 RAG,從零訓練企業專屬模型的另類路線
Mistral 在 NVIDIA GTC 2026 推出 Forge 平台,讓企業與政府用自有資料從零訓練客製模型,隨平台推出 Mistral Small 4。本文解析這條有別於 OpenAI/Anthropic 微調路線的策略、前進部署工程師模式,以及誰真的需要自訓模型。
閱讀文章 ↗NeMo Automodel 接上 Diffusers:微調影像與影片模型不必先改格式
六個 recipe 涵蓋 FLUX 與 Wan:1.3B 單張 40GB A100 起跳、最大 32B;平行化改 YAML、checkpoint 免轉換回 Hub。
閱讀文章 ↗DharmaOCR 對上更新模型:OCR 專用訓練為何仍有價值
Dharma-AI 用巴西葡萄牙文基準比較 DharmaOCR、Mistral OCR4 與 Unlimited-OCR:兩階段訓練、逐 token 漂移機制、Chico Buarque 誤轉案例,以及 0.925 對 0.798 的廠商自評分數背後,產品團隊該看的四個訊號。
閱讀文章 ↗Adaption 推出 AutoScientist:把模型訓練的研究迴圈全自動化
2026 年 5 月 13 日,Sara Hooker 創辦的 Adaption 推出 AutoScientist,同時最佳化資料與訓練配方,聲稱勝率從 48% 升到 64%,免費開放 30 天。本文解析這套自動化訓練系統的能力與驗證難題。
閱讀文章 ↗Config 募 2,700 萬美元種子輪:韓國製造業押注機器人資料
2026 年 5 月,Config 以逾 2 億美元估值完成超額認購的 2,700 萬美元種子輪,三星創投領投,現代、LG、SK 跟投。它不造機器人,而是幫所有人訓練機器人模型,已累積逾 10 萬小時人類動作資料,自比機器人資料的台積電。
閱讀文章 ↗Sakana AI 發布 Doc-to-LoRA 與 Text-to-LoRA:一次前向傳遞生成 LoRA 配接器
Sakana AI 開源 Doc-to-LoRA 與 Text-to-LoRA 兩個超級網路研究:前者把整份文件壓進不到 50 MB 的 LoRA 配接器,後者用一句任務描述即時產生配接器。本文解析架構、實測數字與限制。
閱讀文章 ↗Thinking Machines 大地震:共同創辦人全數出走,500 億募資卡關
NYT 長篇揭露 Thinking Machines 內鬥:三位早期核心逼宮失敗,CTO Barret Zoph 遭 Murati 解職後 OpenAI 立即收編;Meta 收購探詢破局,500 億美元估值的新輪募資陷入苦戰,AI 人才戰爭的代價浮上檯面。
閱讀文章 ↗
2026
9 ARTICLESHow Serverless Fine-Tuning Changes Product Tagging Economics
SageMaker serverless model customization lets you fine-tune Qwen3-8B for structured product tagging without managing training instances, shifting the cost and ops trade-off for catalog enrichment.
READ POST ↗What Fyxer's 53% Draft Acceptance Rate Changes for How You Build Trustworthy AI Assistants
Fyxer's specialized-model email system shows how fine-tuning on real assistant workflows and user edits builds AI trust.
READ POST ↗Small Language Models: The Enterprise Case for Right-Sizing AI
Cohere's guide to SLMs shows why smaller models can cut costs, run locally, and even beat larger ones on specific tasks. Learn how to build a model portfolio that matches size to job.
READ POST ↗Mistral Forge: No Fine-Tuning, No RAG — Training Enterprise Models From Scratch
Mistral's Forge trains enterprise models from scratch on proprietary data — no fine-tuning, no RAG. Embedded engineers, Mistral Small 4, and who actually needs a custom model.
READ POST ↗NeMo Automodel Meets Diffusers: 6 Recipes, Zero Conversion
Six recipes cover FLUX and Wan (1.3B on one 40GB A100 up to 32B); parallelism is a YAML edit, checkpoints round-trip the Hub with no conversion.
READ POST ↗AutoScientist: Adaption Automates Model Training Research
Adaption, Sara Hooker's lab, launched AutoScientist on May 13, 2026: it automates the model-training research loop, claims 35% average gains and win rates rising from 48% to 64%.
READ POST ↗Config Raises $27M to Be the TSMC of Robot Training Data
Config closed an oversubscribed $27M seed at a $200M+ valuation, led by Samsung with Hyundai, LG, and SK joining — a neutral data layer for robot foundation models.
READ POST ↗Sakana AI's Doc-to-LoRA: Documents Become LoRA in One Pass
Sakana AI open-sources Doc-to-LoRA and Text-to-LoRA, hypernetworks that generate LoRA adapters in a single sub-second forward pass — long documents in under 50 MB, task adapters from one sentence.
READ POST ↗Thinking Machines Turmoil: Founders Gone, $50B Round Stalls
NYT reconstructs the Thinking Machines shake-up: Murati fired CTO Barret Zoph, OpenAI rehired the defectors within the hour, Meta's takeover talks died, and a $50B round is stuck.
READ POST ↗