Qwen

Qwen3.8-Max Goes GA: 1M Context and Open Weights

Alibaba's Qwen3.8-Max went GA on August 3: 1M-token context, multimodal input, $2/$6 pricing, and the first open weights in the Max series, built for coding and cowork.

Qwen3.8-Max Goes GA: 1M Context and Open Weights — article cover
On this page6 SECTIONS
  1. Specs and Pricing: A Million Tokens of Context
  2. The Fine Print: Slow and Verbose
  3. Benchmarks and the Cowork Frame
  4. What Open Weights Change
  5. What It Means for Developers
  6. Sources

On August 3, 2026, Alibaba’s Qwen team released Qwen3.8-Max out of preview. The announcement’s title states the ambition plainly: “A New Bar for Coding and Cowork.” This is the general-availability version of a model that surfaced in mid-July as a preview locked to subscription plans with no published weights — and, per launch-day announcements, Qwen3.8-Max now becomes the first model in the Max series with open weights, with the smaller Qwen3.8-27B set to follow within a week.

The community noticed. The Hacker News launch discussion reached 1,124 points and more than 600 comments. In a summer already saturated with flagship releases, that level of attention makes a point: open weights at the frontier remain the thing developers care about most.

Specs and Pricing: A Million Tokens of Context

Combining the launch information with the specs catalogued by Artificial Analysis, Qwen3.8-Max is a reasoning model with extended thinking. It accepts text, image, and video input, and carries a 1-million-token context window — roughly 1,500 A4 pages. Pricing sits in the current flagship band: $2 per million input tokens and $6 per million output tokens, with an 88% cache discount that brings the blended rate to about $1.18 per million tokens. For read-heavy workloads — RAG, repeated questioning over long documents, agents that re-read context — that cache discount matters more than the headline price. The model is already listed on OpenRouter, so third-party applications can wire it in without friction.

The Fine Print: Slow and Verbose

Artificial Analysis scores it 40 on its Intelligence Index, ranking 30th out of 200 models and well above the median of 24. Its evaluation spans benchmarks including Terminal-Bench v4.0, SciCode, and Humanity’s Last Exam, and community discussion broadly places its coding performance alongside Claude and the rest of the top tier.

The same data carries two caveats. First, speed: 37.9 output tokens per second, far below the 65.8 median. Second, verbosity: the full evaluation run generated roughly 180 million tokens, about double the median. The good news is responsiveness — 2.55 seconds to first token, better than the 3.67-second median. Read together, these numbers describe a model built for sustained work rather than real-time conversation: a fit for background tasks and batched agent workflows, a poor fit for experiences where the user watches every token arrive.

Benchmarks and the Cowork Frame

The benchmark mix itself is the signal. Terminal-Bench measures long-horizon operation inside a terminal, SciCode measures scientific computing code, and the suite leans agentic throughout — designed for models that work inside harnesses, not chat leaderboards. Putting “Cowork” in the title amounts to a product statement: the intended customers are agent systems and coding tools, and long-horizon human-machine collaboration is the first scenario. That matches the observation running through the launch discussion — many benchmarks now put it level with closed rivals like Claude, and the open weights are the differentiator.

What Open Weights Change

July’s preview read to much of the community as “a flagship locked behind a subscription wall.” Opening the weights a month later restores Qwen’s familiar two-track strategy: monetize the API, compete for the ecosystem with open weights. For teams that self-host, this is the first time Max-class weights can be downloaded, fine-tuned, and deployed privately. For Qwen, it is a direct answer to open-weight rivals such as Kimi K3. Two things to watch next: the exact license terms, and whether the 27B build sustains the local-deployment reputation Qwen3.6 earned.

What It Means for Developers

Three takeaways. First, there is one more frontier-class open-weight option, with pricing and a cache discount that favor read-heavy applications — architectures that reuse long context instead of resending it will benefit most. Second, slow output plus high verbosity means prompt and agent design now has to manage token budgets seriously, because both latency and cost scale with it. Third, “coding and cowork” in the announcement title confirms where evaluation is heading: from single-turn question answering toward sustained, agentic work. When choosing models, long-horizon stability is worth more than a one-shot leaderboard score.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL