Compare models.
Every model on the bench, side by side on the same 30 tasks. Toggle entries off to focus the comparison.
By harness
The full matrix. Cell tint scales with the value; the strongest tint wins the row.
| harness | ||
|---|---|---|
| 50.0%15/30 | 69.0%20/29 | |
| 53.3%16/30 | 72.4%21/29 | |
| 50.0%15/30 | 65.5%19/29 | |
| 56.7%17/30 | 72.4%21/29 | |
| 56.7%17/30 | 65.5%19/29 | |
| 53.3%16/30 | — | |
| 60.0%18/30 | 72.4%21/29 |
| Model | |||||||
|---|---|---|---|---|---|---|---|
| 50.0%15/30 | 53.3%16/30 | 50.0%15/30 | 56.7%17/30 | 56.7%17/30 | 53.3%16/30 | 60.0%18/30 | |
| 69.0%20/29 | 72.4%21/29 | 65.5%19/29 | 72.4%21/29 | 65.5%19/29 | — | 72.4%21/29 |
Overall
Each column shows that model's selected results on Claude Code, across the same 30 tasks. Choose a harness to compare results across the suite.
Success with retries
Each bar is one model's success on the picked harness, with observed pass@1 and projected pass@2, pass@3 and pass^3. These are expected rates, not measured repeat runs. One scale whichever harness is picked, sorted by pass@3.
Modelpass@3pass^3
GPT-6 Astra97.0%32.8%
DeepSeek V4 Pro55.0%45.0%
0%50%100%
- pass@1
- pass@2
- pass@3
- pass^3
Projected retries · pass@1 observed · every model on the picked harness · sorted by pass@3
Shape by harness
Success rate on every harness, one polygon per entry.
DeepSeek V4 ProGPT-6 Astra
50%69%
53%72%
50%66%
57%72%
57%66%
53%
60%72%