Anthropic

Claude Opus 4.8: Same Price, 41 Days Later, Hundreds of Subagents

Anthropic ships Claude Opus 4.8 41 days after 4.7 at unchanged pricing, adding Dynamic Workflows with hundreds of subagents, effort control, and sharper uncertainty flagging.

Claude Opus 4.8: Same Price, 41 Days Later, Hundreds of Subagents — article cover
On this page6 SECTIONS
  1. A 41-Day Cadence
  2. What the Benchmarks Show
  3. Dynamic Workflows: Hundreds of Subagents in One Session
  4. Effort Control and the Messages API
  5. Honesty Gains and the Mythos Timeline
  6. Sources

On May 28, 2026, Anthropic released Claude Opus 4.8 — 41 days after Opus 4.7 — available everywhere the same day at unchanged pricing: $5 per million input tokens, $25 per million output tokens. The company’s framing is unusually blunt. This is its strongest coding model yet, and it is “more likely to flag uncertainties about its work and less likely to make unsupported claims.”

For developers, the story is less about headline capability and more about three things that change daily work: Dynamic Workflows, which coordinate hundreds of parallel subagents in a single session; effort control, which lets you dial reasoning spend per task; and a Messages API that finally accepts system entries mid-conversation without breaking the prompt cache.

A 41-Day Cadence

TechCrunch highlights the interval: 41 days is far shorter than previous Opus generations. Two forces are behind it. Opus 4.7 landed to a lukewarm reception, and competitive pressure from OpenAI’s Codex push and Google’s Gemini 3.5 Flash (released May 19) made standing still expensive. On pricing, the standard tier holds at 4.7 levels, while a new Fast mode runs at $10 per million input and $50 per million output — roughly 2.5x faster and 3x cheaper than fast modes on earlier models. Speed is now a separately priced option, not a vague subscription promise.

What the Benchmarks Show

  • SWE-bench Pro rises from 64.3% on 4.7 to 69.2%
  • Online-Mind2web hits 84%, ahead of both Opus 4.7 and GPT-5.5
  • Artificial Analysis measures 15% fewer turns per task and 35% fewer output tokens than 4.7 — though still about 30% more turns than GPT-5.5
  • GDPval-AA posts 1,890, a 137-point jump over 4.7 and a new lead
  • The model is roughly 4x less likely than 4.7 to let flaws in its own code pass unremarked

One caveat belongs next to those numbers: GPT-5.5 scored 83.4% on Terminal-Bench 2.1 using the Codex CLI harness. Cross-model comparisons depend heavily on the harness and evaluation conditions. Before committing, look for reproductions in an environment that matches yours, not a single leaderboard entry.

Dynamic Workflows: Hundreds of Subagents in One Session

Dynamic Workflows ships as a research preview inside Claude Code for Enterprise, Team, and Max plans. The pattern: Claude plans the work upfront, fans out hundreds of parallel subagents within one session, then verifies the results before handing anything back. Anthropic’s example is concrete — codebase-scale migrations across hundreds of thousands of lines, from kickoff to merge, with the existing test suite as the pass bar. If your team has been hand-rolling an orchestrator for exactly this kind of job, the message is that the orchestration layer is being absorbed into the model itself. What frameworks used to bolt on is becoming native model behavior.

Effort Control and the Messages API

Effort control is live on claude.ai and Cowork for all plans, defaulting to “high,” with “extra” (“xhigh” in Claude Code) and “max” for genuinely hard problems. This is a direct cost lever: most routine tasks do not need maximum reasoning effort, and now you can stop paying for it. The Messages API change is smaller in surface area but larger in practice: system entries can now appear inside the messages array, so instructions can be updated mid-task without invalidating the prompt cache. For long-running agents, cache reuse is where the token budget actually lives, so this converts a common workaround into supported behavior.

Honesty Gains and the Mythos Timeline

Calibration is the improvement early testers call out most. Bridgewater highlighted the model’s tendency to proactively flag issues with the inputs and outputs of an analysis. Scott Wu, CEO of Devin (Cognition), says 4.7’s comment verbosity and tool-calling problems are fixed. Databricks measured roughly 61% lower token costs versus 4.7 in its Genie environment, and one tester reported it as the first model to complete every case end-to-end on a “Super-Agent” benchmark. On alignment, Anthropic says misaligned behavior rates are substantially lower than 4.7 and similar to Claude Mythos Preview, with “new highs” on its measures of prosocial traits. Mythos itself remains restricted after its preview raised cybersecurity concerns; the official line is that it reaches all customers “in the coming weeks” once cyber safeguards are strengthened.

This tracks what we argued in our 2026 opening outlook: generation gaps are compressing, and differentiation is shifting from single-turn capability to reliability and cost control over long tasks. Opus 4.8 is the most complete specimen of that trend so far.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL