ProgramBench benchmark test released: The strongest AI model cannot rebuild a program from scratch
The Meta Superintelligence Laboratory, in partnership with Stanford and Harvard, released the ProgramBench benchmark, which requires AI models to reconstruct a complete code base from binaries. Test results showed that all top models, including Claude Opus 4.7 and GPT 5.4, achieved a 0% solution rate, revealing the fundamental limitations of current AI programming capabilities.
Core conclusion
On May 7, 2026, Meta Superintelligence Labs, together with Stanford University and Harvard University, released a new AI programming benchmark test - ProgramBench. The test requires the AI agent to rebuild the complete code base from scratch based solely on compiled binaries and documentation. The results are astounding: All top AI models, including Claude Opus 4.7, GPT 5.4, Gemini 3.1 Pro, all have a 0% solution rate.
Key Points
- Release time: 2026-05-07
- Publisher: Meta Super Intelligence Laboratory + Stanford University + Harvard University
- Core Findings: 0% solution rate for all models in 200 tasks
- BEST RESULT: Claude Opus 4.7 only "almost solved" 3.0% of the tests
- Implications for users of AI tools: AI coding is far from being able to replace human developers
Background: What is ProgramBench?
ProgramBench is a new AI programming benchmark developed by the original SWE-bench team (John Yang, Kilian Lieret, etc.). Unlike existing benchmarks, ProgramBench does not test "fixing bugs" or "adding features" but rather tests whether the AI can reconstruct an entire program through reverse engineering without having any reference to the source code.
In each test task, the AI agent receives an executable (binary) and its documentation, and then must rewrite the complete code base that implements the executable. The AI cannot see any of the original source code, nor can it decompile binaries—it can only infer the logic of the code by running the program and observing the output.
SEO Keywords: AI programming benchmark, ProgramBench, AI code generation capability assessment, LLM reverse engineering
Key findings: 200 tasks, 0% resolution rate
ProgramBench used the mini-SWE-agent framework to evaluate all mainstream AI models. The following is the complete ranking:
| Ranking | Models | Solve Rate | "Almost Solved" Rate |
|---|---|---|---|
| 1 | Claude Opus 4.7 | 0% | 3.0% |
| 2 | Claude Opus 4.6 | 0% | 2.5% |
| 3 | Claude Sonnet 4.6 | 0% | 1.0% |
| 4 | GPT 5.4 | 0% | 0.0% |
| 5 | Gemini 3.1 Pro | 0% | 0.0% |
| 6 | Gemini 3 Flash | 0% | 0.0% |
| 7 | Claude Haiku 4.5 | 0% | 0.0% |
| 8 | GPT 5.4 mini | 0% | 0.0% |
| 9 | GPT 5 mini | 0% | 0.0% |
Data source: ProgramBench | Paper: arXiv:2605.03546
What does this mean?
ProgramBench’s 0% solution rate reveals a fundamental limitation of current AI programming capabilities:
1. AI is good at modification, not good at creation Existing AI programming tools (such as Claude Code, Cursor, Copilot) perform well in daily coding and can usually solve 30-50% of issues. But ProgramBench proves that they almost completely fail when faced with "from scratch" programming tasks. It's like a person can help you correct typos in an article, but he can't write a brand new paper.
2. Reverse engineering is still the domain of humans Inferring code logic from binary files requires understanding the overall architecture, data flow and business logic of the program - which are still the strengths of human programmers. AI is very capable of code completion at the micro level, but it is seriously lacking in macro program reconstruction.
3. "Almost solved" 3% implies the direction of progress The Claude Opus 4.7 passed more than 95% of the tests on 3% of the tasks, which is a weak but present signal. As reasoning capabilities improve, AI may eventually cross this threshold.
Implications for AI automated workflows
For users who use AI tools for automated content production and programming, ProgramBench has several important implications:
- Don’t overestimate AI’s ability to “think independently” — AI is good at working within existing frameworks, but not good at building from scratch
- Human-machine collaboration is still the best strategy — let AI be responsible for code completion and testing, and humans be responsible for architecture design
- Agentic architecture needs a more complete context — As Simon Willison said in a recent discussion, vibe coding and agentic engineering are converging, but they are still far from true intelligence.
- AI tools are positioned as "super assistants" rather than "replacers" - This is the most pragmatic usage mentality at present
Related extended information
Tool entry
Tool platforms appearing in the text: Claude Code, OpenAI, ChatGPT, Gemini, Copilot, Cursor
Internal link guidance
- Want to see AI programming tools in action? Watch: AI Agent Tools 2026 Complete Tutorial: 5 Tools to Build an Automated Pipeline in 30 Minutes
- Real case: Someone used Claude Code to start a business from scratch in 48 hours to earn monthly income $9,000, Claude Code 48 hours to start a business: one person + US$29 monthly fee, monthly income in 3 months $9,000
Next action
- Try out the ProgramBench test suite to evaluate your own AI programming workflows
- Follow the follow-up work of the SWE-bench team - their research in the field of agentic programming is at the forefront
- Subscribe to waytoclawearn.com for more AI tool reviews and automation tutorials
Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
n8n + OpenAI affiliate site
Automate content and affiliate monetization
Claude + n8n automation agency
Charge monthly for agent workflow builds