Quick answer

Anthropic shipped Claude Opus 4.9 on September 6 — just two days after GPT-5.7 narrowed OpenAI's gap on reasoning and agentic tool-use. Opus 4.9 extends the coding lead further (SWE-Bench Verified up to 91.2%, from 4.8's 89.7%) and pushes agentic tool-use reliability to 97%. Pricing is unchanged from Opus 4.8: $12.50 input / $75 output per million tokens, with the same 88% cache discount.

This is the fastest turnaround between a competitor release and an Anthropic response we have seen in 2026. Opus 4.6, 4.7, and 4.8 each shipped roughly two months apart; Opus 4.9 arrived just over two weeks after 4.8's three-month mark, and only two days after GPT-5.7. That timing alone tells you how much Anthropic is treating the coding-and-agents lead as core to its identity right now.

What actually changed in Opus 4.9

  • SWE-Bench Verified rises to 91.2%, up from 4.8's 89.7% — still the highest published score of any model, and the gap to GPT-5.7 widens rather than narrows
  • Agentic tool-use reliability hits 97% (from 96%), a smaller jump than 4.7-to-4.8 but still meaningful at scale for tools like Claude Code and Devin
  • Extended thinking mode gets faster token throughput at the same 60-minute ceiling introduced in 4.8 — same depth, less wall-clock time
  • No pricing change from Opus 4.8: input stays at $12.50/M, output at $75/M, cached input at $1.50/M (88% off)
  • Context window unchanged at 500,000 tokens — Anthropic again prioritised depth and reliability over a bigger number

GPT-5.7 vs Opus 4.9 — where things actually stand now

  • Coding and multi-file refactoring: Opus 4.9 leads clearly, and the gap widened rather than closed this week
  • General reasoning and everyday chat: close enough that ecosystem and pricing matter more than the benchmark charts
  • Agentic reliability: Opus 4.9's 97% edges out GPT-5.7, meaningful for long unsupervised agent runs
  • Access: both are now broadly available — no waitlist, no government-review gate on either side
  • Price: GPT-5.7 is meaningfully cheaper ($5/$15 vs $12.50/$75 per million tokens) — for high-volume, less-demanding workloads that difference adds up fast

The practical takeaway from a single week with two major releases: nothing about your model choice needs to change unless you were specifically waiting to see whether GPT-5.7 closed the gap. It didn't — Opus 4.9 immediately widened it back out on the tasks Anthropic has been quietly winning all year.

Should you switch or upgrade?

If you are already on Opus 4.8, the upgrade to 4.9 is automatic in claude.ai and a one-line model-string change in the API, with no pricing change — there is no real reason not to switch. If you are deciding between GPT-5.7 and Opus 4.9 for coding-heavy work, Opus 4.9 remains the stronger, more expensive choice; for high-volume, cost-sensitive workloads where the gap doesn't matter as much for your task, GPT-5.7's pricing is the more compelling option.

Bottom line

Opus 4.9 is Anthropic's fastest response yet, and it works — the coding and agentic lead GPT-5.7 narrowed two days ago is now wider than before. For most builders, the practical framework hasn't changed: Opus for the hardest coding and agentic work, a cheaper model like GPT-5.7 or GPT-5 mini for high-volume, lower-stakes tasks. What changed this week is how fast both labs are now willing to ship in response to each other — expect this cadence to keep compressing.