SWE-bench
Science
Measuring whether coding agents can repair repository-level scientific software while preserving its scientific contracts.
- Tasks
00 (HARD70)- Repositories
00 - Domains
00
Leaderboard
Pass@1 versus mean token consumption per task. Select a point or model to inspect the configuration.
Labels show every configuration at or above 20% Pass@1, plus every highlighted model. Click multiple legend items or chart points to compare them; click again to remove a highlight.
Model results
| Rank | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | Codex | 100.00% | 83.18% | 80.11% | 95.95% | 50.42% | 46.15% | 57.14% | 44.44% | |
| 02 | Claude Code | 96.64% | 75.11% | 68.60% | 97.37% | 47.90% | 38.46% | 65.31% | 27.78% | |
| 03 | Claude Code | 100.00% | 73.16% | 65.77% | 96.58% | 42.02% | 26.92% | 57.14% | 44.44% | |
| 04 | Codex | 98.32% | 75.57% | 69.98% | 96.30% | 40.34% | 34.62% | 46.94% | 38.89% | |
| 05 | Kimi Code | 98.32% | 66.34% | 57.55% | 94.94% | 35.29% | 25.00% | 44.90% | 38.89% | |
| 06 | Codex | 94.12% | 63.61% | 53.81% | 97.53% | 31.93% | 17.31% | 46.94% | 33.33% | |
| 07 | Claude Code | 100.00% | 65.75% | 56.46% | 97.36% | 29.41% | 25.00% | 38.78% | 16.67% | |
| 08 | Codex | 93.28% | 61.89% | 51.09% | 94.92% | 24.37% | 11.54% | 36.73% | 27.78% | |
| 09 | Claude Code | 98.32% | 61.41% | 52.34% | 95.74% | 23.53% | 19.23% | 26.53% | 27.78% | |
| 10 | Claude Code | 97.48% | 55.96% | 42.68% | 96.67% | 21.01% | 9.62% | 32.65% | 22.22% | |
| 11 | Claude Code | 100.00% | 58.77% | 47.67% | 95.02% | 19.33% | 11.54% | 28.57% | 16.67% | |
| 12 | Claude Code | 98.32% | 54.70% | 39.20% | 96.59% | 16.81% | 7.69% | 24.49% | 22.22% | |
| 13 | Claude Code | 96.64% | 51.35% | 37.90% | 93.55% | 15.13% | 9.62% | 24.49% | 5.56% | |
| 14 | Claude Code | 96.64% | 50.39% | 36.81% | 92.88% | 15.13% | 5.77% | 22.45% | 22.22% | |
| 15 | Claude Code | 90.68% | 52.35% | 40.21% | 91.05% | 15.13% | 11.54% | 18.37% | 16.67% | |
| 16 | Codex | 96.64% | 51.79% | 38.33% | 95.16% | 14.29% | 5.77% | 24.49% | 11.11% | |
| 17 | Claude Code | 88.98% | 39.11% | 21.27% | 91.33% | 7.56% | 3.85% | 12.24% | 5.56% |
Pass@1 requires every applicable private test to pass. Issue, Expert, and Engineering report Pass@1 for the three task paradigms.
Hard70
Models above 20% on the full benchmark, evaluated on the Hard70 subset.
As of August 28, 2026, 16:31 UTC, we estimated task difficulty using the 12 models with complete results for all 119 tasks available at that time: Claude Opus 5 Max, GLM-5.2, Qwen3.8-27B, GPT-5.6-sol, Nex-N2-Pro, DeepSeek V4 Flash, Intern-S2-Preview-397B, Qwen3.6-35B-A3B, Nex-N2-mini, agents-a1, BigBang-v1, and Qwen3.5-9B. For each task, a model receives a binary reward of 1 if it passes the complete task evaluation and 0 otherwise. We average these rewards across the 12 models and select the 70 tasks with the lowest mean reward as this subset.
| Rank | Model | Agent | Pass@1 | Input tokens / task |
|---|---|---|---|---|
| 01 | GPT-6 Astra (max) | Codex | 30.00% | 1.404M |
| 02 | Claude-Opus-5 (max) | Claude Code | 21.43% | 7.048M |
| 03 | DeepSeek-V4-Pro (max) | Claude Code | 20.00% | 11.430M |
| 04 | GPT-5.6-sol (max) | Codex | 14.29% | 3.921M |
| 05 | Kimi-K3 (max) | Kimi Code | 7.14% | 4.926M |
| 06 | GLM-5.2 (max) | Codex | 5.71% | 3.345M |
| 06 | Qwen3.8-27B (xhigh) | Claude Code | 5.71% | 5.024M |
| 08 | Nex N2 | Codex | 2.86% | 8.033M |
| 09 | DeepSeek-V4-flash (max) | Claude Code | 1.43% | 21.990M |
| 09 | Intern-S2-Preview-397B (max) | Claude Code | 1.43% | 8.028M |
Tasks in Hard70 (70)
001, 002, 003, 005, 006, 007, 008, 011, 014, 019, 020, 033, 035, 036, 037, 039, 040, 041, 042, 043, 044, 046, 047, 049, 050, 051, 052, 054, 055, 059, 060, 061, 063, 064, 065, 066, 068, 070, 071, 072, 073, 075, 076, 079, 080, 082, 083, 085, 086, 087, 088, 089, 090, 092, 094, 095, 096, 097, 098, 099, 102, 104, 109, 110, 111, 113, 114, 116, 117, 119