Three weeks ago, Anthropic's Mythos-class architecture lived in exactly one place: Claude Fable 5, priced at $10 per million input tokens and $50 per million output for the hardest agentic work. The open question was whether that architecture would ever come down in price, or stay a flagship-only luxury. On July 1, Anthropic answered it. Claude Sonnet 5 runs on the same Mythos-class architecture, launches at $2 per million input tokens and $10 per million output — a rate Anthropic has since made standard instead of moving to $3/$15 — not a new "Sonnet 4.7," but a second model built on Fable 5's foundation at the Sonnet price tier.
One architecture, two tiers
Sonnet 5 launches below Sonnet 4.6 during the introductory window: $2 per million input tokens and $10 per million output through August 31, 2026 — and on August 31 Anthropic confirmed $2/$10 as the standard price, cancelling the September 1 move to $3/$15. That keeps it below Sonnet 4.6's $3/$15 and Opus 4.8's $5/$25. Cached input is $0.20 per million, and Batch cuts both directions in half to $1/$5. Context stays at 1M tokens, matching Sonnet 4.6 and Opus 4.8, but max output jumps to 128,000 tokens, double Sonnet 4.6's 64,000 and matching Opus 4.8's 128,000. The API id is claude-sonnet-5, and Anthropic's retirement commitment is not sooner than June 30, 2027.
The architecture brings two more inherited traits. Adaptive thinking is always on, with no extended-thinking toggle to flip — the same behavior Fable 5 introduced. The same safety classifiers also apply: classified offensive-cyber, biology/chemistry, or distillation requests return an explicit refusal. Another-model retry occurs only when configured by the application or product.
What the benchmarks say
The headline number is SWE-bench Verified: Sonnet 5 scores 89.4%, ahead of Claude Opus 4.8's 88.6% — a mid-tier model beating last month's flagship on a closely watched metric. SWE-bench Pro, the harder agentic-coding test, comes in at 71.8%. Terminal-Bench 2.1 lands at 85.6%. Reasoning tells a different story: GPQA Diamond is 92.0%, behind Opus 4.8's 93.6%, so Opus 4.8 keeps the edge on graduate-level science reasoning. On ARC-AGI-2, Sonnet 5 scores 20.0, ahead of Sonnet 4.6's 15.0. Humanity's Last Exam without tools comes in at 42.5%. The rest of the sheet: LMSYS Arena 1435, MMLU 93.8%, HumanEval 96.0%, MATH 93.5%.
Read the pattern honestly: Sonnet 5 closes almost all of the coding gap to Opus 4.8, and actually passes it on SWE-bench Verified, while giving up ground on the hardest reasoning benchmark. At $2 / $10, now the standard rate, that is a sharp trade: 40% of Opus 4.8's input and output rates.
Sonnet 5 vs Sonnet 4.6 vs Opus 4.8
| Spec | Claude Sonnet 5 | Claude Sonnet 4.6 | Claude Opus 4.8 |
|---|---|---|---|
| Price (in/out per 1M) | $2 / $10 (standard since Aug. 31) | $3 / $15 | $5 / $25 |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Max output | 128K | 64K | 128K |
| SWE-bench Verified | 89.4% | 79.6% | 88.6% |
| GPQA Diamond | 92.0% | 89.9% | 93.6% |
| Thinking mode | Adaptive, always on | Standard | Standard |
| Restrictions | Classified cyber / bio / distillation requests refuse explicitly | Standard | Standard |
Is the mid-tier upgrade worth it?
Run the math on a real workload before switching. A coding agent burning 2M input tokens and 400K output tokens a day costs about $8 on Sonnet 5 (2 × $2 + 0.4 × $10), against $12 on Sonnet 4.6 (2 × $3 + 0.4 × $15) and $20 on Opus 4.8 (2 × $5 + 0.4 × $25). Because Anthropic kept $2/$10 as the standard rate, that $8 did not rise after August 31. The cost calculator will run this against your own volumes, and the Claude Sonnet 5 pricing breakdown covers the caching and batch math in full.
Where the upgrade clearly pays off: coding agents and long-running tool loops that were hitting Sonnet 4.6's 64K output ceiling, since 128K max output means far fewer truncated responses mid-task. Where it is a harder sell: workloads that are already comfortable on Sonnet 4.6 and do not need the extra output headroom or the coding bump — the lower price is welcome, but it is still a migration to test rather than a blind swap. And if your work depends on the GPQA-level reasoning ceiling where Opus 4.8 led at launch, Sonnet 5 did not close that gap; test Claude Opus 5, which replaced Opus 4.8 at the same $5 / $25, rather than defaulting to the older Opus.
A crowded launch week
Sonnet 5 didn't ship in isolation. The same day, Anthropic's export-control review closed and Claude Fable 5 was restored to all customers, with AWS reinstating Bedrock access. OpenAI's GPT-5.6 was partner-gated at that point and later reached general availability on July 9. None of that changes the Sonnet 5 math directly, but it establishes the current availability timeline.