Quick answer
OpenAI released GPT-5.7 on September 3 — a broadly-available update (unlike the access-restricted GPT-5.6 Sol variant) targeting the coding and reasoning gap Claude Opus 4.8 opened up in June. Early benchmark results show GPT-5.7 closing most of that gap on general reasoning, and pulling roughly even on agentic tool-use reliability, but Opus 4.8 still holds a clear edge on the largest, most complex codebases. Pricing is unchanged from GPT-5 at $5/$15 per million input/output tokens.
Back in June, when Anthropic shipped Claude Opus 4.8, we wrote that "the next OpenAI release will need to retake the lead, or the narrative around 'the best AI' shifts decisively to Anthropic." What actually happened next was messier than a straight rematch: GPT-5.6 shipped in three variants (Sol, Terra, Luna) with Sol's frontier capability gated behind a new US government review process, meaning most developers and businesses never actually got hands-on access to OpenAI's strongest model. Luna got a solid coding update in late August. Now, three months after Opus 4.8, GPT-5.7 is OpenAI's first genuinely broadly-available answer.
What actually changed in GPT-5.7
- Reasoning benchmarks (GPQA Diamond, AIME-style math) move from trailing Opus 4.8 by several points to trailing by roughly one point — a real but not decisive gain
- Agentic tool-use reliability improves noticeably, closing most of the gap Opus 4.8 opened with its 96%-reliability jump in June
- Multi-file coding and refactoring improves versus GPT-5.6 Luna, but independent developer testing still puts Opus 4.8 ahead on the largest, most complex codebases
- No pricing change — GPT-5.7 stays at GPT-5's existing $5/$15 per million token rate, with the same 50%-off caching discount
- Unlike GPT-5.6 Sol, GPT-5.7 ships to all ChatGPT Plus and API users immediately — no waitlist, no government-review gate
So does it retake the lead?
Mostly no, but it closes the gap enough that "which model is better" now genuinely depends on the task rather than having one clear universal answer. For everyday chat, writing, and small-to-medium coding tasks, GPT-5.7 and Opus 4.8 are close enough that ecosystem and pricing preference matter more than raw capability. For the largest, most complex engineering work — and for agentic reliability at the margin — Opus 4.8 still has a real, measurable edge.
The more interesting story than "who's ahead this week" is that GPT-5.7 is broadly available where GPT-5.6 Sol was not. A model that's slightly behind but universally accessible is often more useful in practice than a technically-stronger one gated behind a review process most developers will never clear.
Should you switch?
- Already on Opus 4.8 for coding: stay put. GPT-5.7 narrows the gap but does not close it on the work Opus 4.8 is strongest at
- On GPT-5 or GPT-5.6 Luna already: upgrade is free/automatic and worth it — better reasoning and tool-use at no extra cost
- Evaluating both from scratch: for agentic engineering work at scale, test both on your actual codebase rather than trusting either vendor's benchmark chart
Related reading
Bottom line
GPT-5.7 is a real, broadly-available step forward that closes most — not all — of the gap Claude Opus 4.8 opened in June. It doesn't retake the lead outright, but it ends the period where OpenAI's strongest publicly-usable model was clearly behind. Expect Anthropic's response within weeks, not months, given how tight the race has become.

