Scientific software engineering benchmark

SWE-bench
Science

Measuring whether coding agents can repair repository-level scientific software while preserving its scientific contracts.

Tasks
0(HARD70)
Repositories
0
Domains
0
01

Leaderboard

Pass@1 versus mean token consumption per task. Select a point or model to inspect the configuration.

Claude-Opus-5 (max)47.90% Pass@17.048M input
Interactive scatter plot of model Pass@1 against mean token use.

Labels show every configuration at or above 20% Pass@1, plus every highlighted model. Click multiple legend items or chart points to compare them; click again to remove a highlight.

02

Model results

17 configurations
Rank
01Codex100.00%83.18%80.11%95.95%50.42%46.15%57.14%44.44%
02Claude Code96.64%75.11%68.60%97.37%47.90%38.46%65.31%27.78%
03Claude Code100.00%73.16%65.77%96.58%42.02%26.92%57.14%44.44%
04Codex98.32%75.57%69.98%96.30%40.34%34.62%46.94%38.89%
05Kimi Code98.32%66.34%57.55%94.94%35.29%25.00%44.90%38.89%
06Codex94.12%63.61%53.81%97.53%31.93%17.31%46.94%33.33%
07Claude Code100.00%65.75%56.46%97.36%29.41%25.00%38.78%16.67%
08Codex93.28%61.89%51.09%94.92%24.37%11.54%36.73%27.78%
09Claude Code98.32%61.41%52.34%95.74%23.53%19.23%26.53%27.78%
10Claude Code97.48%55.96%42.68%96.67%21.01%9.62%32.65%22.22%
11Claude Code100.00%58.77%47.67%95.02%19.33%11.54%28.57%16.67%
12Claude Code98.32%54.70%39.20%96.59%16.81%7.69%24.49%22.22%
13Claude Code96.64%51.35%37.90%93.55%15.13%9.62%24.49%5.56%
14Claude Code96.64%50.39%36.81%92.88%15.13%5.77%22.45%22.22%
15Claude Code90.68%52.35%40.21%91.05%15.13%11.54%18.37%16.67%
16Codex96.64%51.79%38.33%95.16%14.29%5.77%24.49%11.11%
17Claude Code88.98%39.11%21.27%91.33%7.56%3.85%12.24%5.56%

Pass@1 requires every applicable private test to pass. Issue, Expert, and Engineering report Pass@1 for the three task paradigms.

03

Hard70

Models above 20% on the full benchmark, evaluated on the Hard70 subset.

As of August 28, 2026, 16:31 UTC, we estimated task difficulty using the 12 models with complete results for all 119 tasks available at that time: Claude Opus 5 Max, GLM-5.2, Qwen3.8-27B, GPT-5.6-sol, Nex-N2-Pro, DeepSeek V4 Flash, Intern-S2-Preview-397B, Qwen3.6-35B-A3B, Nex-N2-mini, agents-a1, BigBang-v1, and Qwen3.5-9B. For each task, a model receives a binary reward of 1 if it passes the complete task evaluation and 0 otherwise. We average these rewards across the 12 models and select the 70 tasks with the lowest mean reward as this subset.

RankModelAgentPass@1Input tokens / task
01GPT-6 Astra (max)Codex
30.00%
1.404M
02Claude-Opus-5 (max)Claude Code
21.43%
7.048M
03DeepSeek-V4-Pro (max)Claude Code
20.00%
11.430M
04GPT-5.6-sol (max)Codex
14.29%
3.921M
05Kimi-K3 (max)Kimi Code
7.14%
4.926M
06GLM-5.2 (max)Codex
5.71%
3.345M
06Qwen3.8-27B (xhigh)Claude Code
5.71%
5.024M
08Nex N2Codex
2.86%
8.033M
09DeepSeek-V4-flash (max)Claude Code
1.43%
21.990M
09Intern-S2-Preview-397B (max)Claude Code
1.43%
8.028M
Tasks in Hard70 (70)

001, 002, 003, 005, 006, 007, 008, 011, 014, 019, 020, 033, 035, 036, 037, 039, 040, 041, 042, 043, 044, 046, 047, 049, 050, 051, 052, 054, 055, 059, 060, 061, 063, 064, 065, 066, 068, 070, 071, 072, 073, 075, 076, 079, 080, 082, 083, 085, 086, 087, 088, 089, 090, 092, 094, 095, 096, 097, 098, 099, 102, 104, 109, 110, 111, 113, 114, 116, 117, 119