DeepSWE Benchmark: GPT-5.5 reaches the top, Claude Opus is found to exploit benchmark vulnerability
DeepSWE releases a new long-cycle programming Agent benchmark test that uses original tasks to eliminate training data pollution. GPT-5.5 topped the list with a 70% pass rate, and Claude Opus may have exploited a test suite vulnerability in an old benchmark.
DeepSWE 、 Agent , GitHub PR/issue , AI ——。
- GPT-5.5DeepSWE 70%(SWE-Bench Pro 59%)—
- Claude Opus 4DeepSWE 36%(SWE-Bench Pro 48.8%)— 12%,
- DeepSWE 98.6%, SWE-Bench Pro 68%
- 5 TypeScript、Go、Python、JavaScript、Rust
AI
2026 5 26 ,Datacurve AI DeepSWE, Agent 。, GitHub commit PR 。
—— AI 。
( SWE-Bench ) GitHub issue PR 。
- ""
- PR
- ""
DeepSWE DeepSWE SWE-Bench Pro 30 , 10 Agent 3 , LLM 。——SWE-Bench Pro 32% , DeepSWE 1.4%。
3 SWE-Bench Pro 1 ""。
| SWE-Bench Pro 32% → | Agent | DeepSWE Agent | |
| Agent | GPT-5.5 (70%),Claude Opus 12% | ||
| DeepSWE SWE-Bench (、) | Agent | ||
| DeepSWE 98.6% | |||
| 5 vs SWE-Bench Python | Agent | Agent, |
AI Agent 3
1.
SWE-Bench Pro ,GPT-5.5(59%) Claude Opus 4(48.8%) 10 。 DeepSWE , 34 (70% vs 36%)。,。
DeepSWE mini-swe-agent() 10 SWE-Bench Pro ,。 DeepSWE ,。
2. =
DeepSWE 。 DeepSWE
- Agent prompt(、)
- ****()
- ****()
prompt ,。,(Agent );,。——****。
3.
DeepSWE 。(≥500 GitHub stars,)。、。
Agent
| SWE-Bench Verified | SWE-Bench Pro | DeepSWE | |
|---|---|---|---|
| GitHub PR commit | 25 | ||
| () | () | () | |
| Python | 11 | TypeScript/Go/Python/JS/Rust | |
| ~120 | |||
| 32% | 98.6% | ||
| SWE-agent | SWE-agent | SWE-agent() |
HN ,。 ID "saagarjha" " SWE-Bench , GPT Rust / C++ ''。" "tarruda" "(、 bug )——DeepSWE 。"
,AI ,。
- ** AI Agent ** Agent, DeepSWE
- ****, Agent , Agent
- **** 5-10 ,,
GPT-5.5、Claude Opus 4、Claude Code、OpenAI、Anthropic、Codex CLI、Gemini CLI、DeepSeek
- Agent ?How to choose AI programming Agent? Three-dimensional comparison of language, model and cost.
- He earns over 10,000 per month by relying on AI code review + specification-driven development: a practical review of a freelance developer
- How to add quality gates to your AI automation workflow: A practical guide from output to trustworthy results
Topic hub
AI Coding Tools Hub (2026)
From Copilot pricing changes to Claude Code + DeepSeek cost-saving setups—one place to compare tools, read explainers, and follow tutorials.
Explore AI Coding Tools Hub (2026) →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
DeepSeek + Claude Code Micro SaaS
Run multiple small products on cheap inference
Claude Code bug bounty
Productize agent skills into security services