WayToClawEarn
High impactOpenAI 官方/Hacker News

SWE-bench benchmark is saturated: OpenAI confirms AI coding capability assessment enters new stage

OpenAI officially confirmed that SWE-bench Verified is no longer able to distinguish cutting-edge AI coding model capabilities. The current top score is 93.9% (Anthropic), with 59.4% of remaining questions having testing flaws. The co-founder of SWE-bench announced that the Multilingual/Multimodal version will be open source. What does this mean for users of AI automation tools?

WayToClawEarn EditorialPublished Apr 27, 2026Updated Aug 8, 2026

Editorial review of public sources · AI-assisted drafting. How we work · Original source

Core conclusion

OpenAI officially issued a document confirming that the SWE-bench Verified benchmark has been "saturated" by cutting-edge models such as Claude - the current best model score is 93.9% (achieved by Anthropic). OpenAI found in its review that at least 59.4% of the 27.6% remaining open problems had test case flaws, meaning that the correct code for those problems was incorrectly judged as failing. This discovery marks a new stage in the evaluation of AI coding capabilities, where traditional single benchmarks are no longer able to distinguish the gaps between leading-edge models.

Key Points

  • Time of incident: 2026-04-26
  • Affected objects: AI coding model developers, AI automation tool users, and benchmarking communities
  • Core Change: SWE-bench co-founder Ofir Press confirmed that the benchmark is saturated, and SWE-bench Multilingual and Multimodal versions will be open source soon

Background and trigger events

On April 26, 2026, OpenAI published the article "SWE-bench Verified no longer measures frontier coding capabilities" on the official blog, which immediately appeared on the homepage of Hacker News and received 249 likes and 141 comments within 12 hours. SWE-bench co-founder Ofir Press subsequently responded personally on HN, confirming that SWE-bench Verified has reached saturation at 93.9% (achieved by Anthropic), and noted that SWE-bench Multilingual and SWE-bench Multimodal will be open sourced within a month as next-generation benchmarks.

Key Impact

DimensionsChangeWhat it means to usRecommended actions
Model evaluationSWE-bench Verified saturationSingle benchmark cannot distinguish top modelsFocus on new benchmark SWE-bench Multilingual
Test reliability59.4% of remaining questions have test flawsThe true capabilities of some models may be underestimatedCombining multiple benchmarks to cross-validate model capabilities
Development toolsAI coding capabilities are approaching the ceilingAutomated code generation has entered a practical periodIntegrate AI coding tools into workflow
Market competitionAll cutting-edge models convergeDifferentiation shifts from "can it" to "is it good"Focus on API cost, latency and maintainability

Adaptation suggestions

The saturation of SWE-bench is a landmark event - it shows that AI coding capabilities have crossed the practical threshold. For developers using AI automation tools, this means:

  • Integrate AI coding capabilities instantly into your daily development process, no more waiting and watching
  • Focus on actual task performance rather than benchmark scores and choose the model that best suits your scenario
  • Establish an automated testing pipeline to verify the quality of AI-generated code
  • For users of AI Agent tools, both OpenAI and Claude are mature. The difference lies in the cost and efficiency of specific scenarios.

Task List (Example)

  • Evaluate the actual performance of currently used AI coding tools (Claude Code / ChatGPT / DeepSeek)
  • Add code review link to automated workflow to balance efficiency and quality
  • Pay attention to the release of SWE-bench Multilingual and prepare for benchmark migration in advance

Extended interpretation

Behind the saturation of SWE-bench is the paradigm shift in the field of AI coding from "can it run through" to "can it be practical?" In the past year, from Claude Code to OpenAI’s o-series models, AI’s GitHub Issue resolution rate has jumped from 20% to 90%+. Now, the real challenge is not whether AI can write code, but how to embed AI coding capabilities into the complete development workflow.

AI

Related extended information

Tool entry (trigger tool floating card)

In this area, OpenAI and Claude (Claude Code) are the two current contenders for AI coding capabilities. DeepSeek also has a great performance. If you want to automate code work through AI Agent tools, n8n and OpenClaw are both good choices. Hermes Agent is also actively supporting AI coding scenarios.

Internal link guidance

View source →

Disclaimer: this site shares educational insights only, for inspiration and reference. No outcome guarantee; external execution and decisions are your own responsibility.