2026
2 篇文章2026
2 ARTICLESWhat Deploying Qwen3.8-2.4T-A95B on HyperPod Changes for Self-Hosting Frontier Models
A practical walkthrough for serving a 2.4T open-weights MoE on a single 8-GPU node with vLLM and SageMaker HyperPod.
READ POST ↗Nari Labs Gets Qwen3-TTS Talking in Under 50 ms
Nari Labs open-sourced a Qwen3-TTS 1.7B serving stack that hits 10 requests per second with sub-50 ms p95 time-to-first-audio on a single H100.
READ POST ↗