OpenAI releases GPT-5.5: 82.7% Terminal-Bench refreshes SOTA, Agent programming capability has greatly improved
OpenAI released GPT-5.5 on April 23, 2026, reaching a SOTA accuracy of 82.7% on Terminal-Bench 2.0. It is priced at $5/ million input tokens and has begun to be pushed to Plus/Pro users.
OpenAI releases GPT-5.5: 82.7% Terminal-Bench refreshes SOTA, Agent programming capabilities have greatly improved
Core conclusion
OpenAI released GPT-5.5 on April 23, 2026, claiming it to be the "smartest and easiest to use" model. The new model achieves 82.7% SOTA accuracy on Terminal-Bench 2.0 and 73.1% in Expert-SWE internal evaluation, while significantly reducing token consumption while maintaining the same inference latency as GPT-5.4. The model has begun to be gradually pushed to Plus/Pro/Business/Enterprise users in ChatGPT and Codex, and the API price is set at $5/ million input tokens and $30/ million output tokens.
Background and trigger events
| Field | Content |
|---|---|
| Time | 2026-04-23 |
| Location/channel | OpenAI official blog release |
| Main players | OpenAI, NVIDIA, Anthropic (competitive product comparison) |
Event Highlights:
- GPT-5.5 is the successor version of GPT-5.4, focusing on improving agentic coding, computer use and knowledge work capabilities
- Same latency as GPT-5.4 but smarter and more token efficient
- Analyze several weeks of production traffic patterns through Codex, write custom heuristic algorithms to optimize GPU load balancing, and increase token generation speed by 20%+
- Based on NVIDIA GB200/GB300 NVL72 system training and inference
- The system card has been released, and the biological/chemical and cyber security capabilities have been evaluated as High (not Critical)
Key Impact
Technical Impact Table
| Dimensions | Impact content |
|---|---|
| Agent Programming | Terminal-Bench 2.0 82.7% (+7.6pp vs GPT-5.4), SWE-Bench Pro 58.6% (+0.9pp), Expert-SWE 73.1% (+4.6pp) |
| Knowledge Work | GDPval 84.9% (win rate), OSWorld-Verified 78.7%, Tau2-bench Telecom 98.0% |
| Scientific Research | GeneBench 25.0% (+6pp), BixBench 80.5% (+6.5pp), FrontierMath Tier 1-3 51.7% (+4.1pp) |
| Long context | 1M token context window, MRCR 8-needle 512K-1M reaches 74.0% (vs GPT-5.4 36.6%) |
| Abstract Reasoning | ARC-AGI-2 Verified 85.0% (+11.7pp vs GPT-5.4) |
Industry Impact Table
| Aspects | Subject to change |
|---|---|
| Competitive landscape | Pricing is between GPT-5.4 and Claude Opus 4.7, output pricing is 20% higher but token efficiency is higher |
| Code Assistant Marketplace | Codex users to gradually receive GPT-5.5 support within hours, 400K context windows |
| Developer Workflow | Testers say "the first programming model with serious conceptual clarity", significantly improved understanding of large code bases |
| AI Security | Launching a more stringent cybersecurity classifier and expanding the Trusted Access for Cyber project |
Adaptation suggestions
To Developers:
- GPT-5.5 can generate tokens 1.5x faster in Codex via Fast mode (cost 2.5x)
- API pricing $5/ million inputs, $30/ million outputs, Batch/Flex half price, Priority 2.5x
- GPT-5.5 Pro API pricing $30/ million inputs, $180/ million outputs, for higher accuracy requirements
For users:
- GPT-5.5 Thinking in ChatGPT has been opened to Plus/Pro users
- GPT-5.5 Pro is only available to Pro/Business/Enterprise users
Points worthy of attention:
- OpenAI is building a "global infrastructure for agentic AI"
- GPT-5.5 is used to improve its own inference infrastructure - the model helps improve the system that serves it, forming a positive cycle
- Mathematical discovery ability: GPT-5.5 assisted in discovering a new proof of Ramsey number, which was later verified by Lean. It was the first public example of its participation in core mathematics research.
Tool entry
This release involves multiple AI ecological tools:
- OpenAI — publisher of GPT-5.5, models available via ChatGPT and Codex
- Anthropic — Claude Opus 4.7 has been compared in several benchmarks
- NVIDIA — Provides GB200/GB300 NVL72 system to support GPT-5.5 training and inference
- Codex — OpenAI’s AI programming agent, GPT-5.5 demonstrates the strongest agentic programming capabilities in Codex
- n8n — Workflow automation tool, closely related to the AI agent ecosystem
Industry reaction
- Dan Shipper (Founder and CEO, Every): "The first programming model with serious conceptual clarity."
- Justin Boitano (VP of Enterprise AI, NVIDIA): "GPT-5.5 delivers the sustained performance needed to perform intensive work...enabling teams to move from natural language prompts to end-to-end feature releases."
- Brandon White (co-founder and CEO of Axiom Bio): "If OpenAI continues to develop like this, the foundation of drug discovery will have changed by the end of the year."
- HN Community: GPT-5.5 topped the HN homepage with 1416 points. Developers are concerned about its substantial lead of 82.7% on Terminal-Bench (compared to 69.4% for Claude Opus 4.7), but are concerned that it still lags behind Claude Opus 4.7 on SWE-Bench Pro (64.3% vs 58.6%). Several users noticed that ARC-AGI-3 scores were not published.
Action guidance
Want to learn more about practical applications? Please check out:
- Tutorial: How to configure GPT-5.5 in Codex for efficient programming
- Case: A real-life case of using GPT-5.5 to automate and optimize development workflow
Topic hub
AI Coding Tools Hub (2026)
From Copilot pricing changes to Claude Code + DeepSeek cost-saving setups—one place to compare tools, read explainers, and follow tutorials.
Explore AI Coding Tools Hub (2026) →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
DeepSeek + Claude Code Micro SaaS
Run multiple small products on cheap inference
Claude Code bug bounty
Productize agent skills into security services