Compare harnesses
Every harness on the bench, side by side on the same 30 tasks.
By model
The full matrix. Cell tint scales with the value; the strongest tint wins the row.
| model | |||||||
|---|---|---|---|---|---|---|---|
| 50.0%15/30 | 53.3%16/30 | 50.0%15/30 | 56.7%17/30 | 56.7%17/30 | 53.3%16/30 | 60.0%18/30 | |
| 69.0%20/29 | 72.4%21/29 | 65.5%19/29 | 72.4%21/29 | 65.5%19/29 | — | 72.4%21/29 |
| Harness | ||
|---|---|---|
| 50.0%15/30 | 69.0%20/29 | |
| 53.3%16/30 | 72.4%21/29 | |
| 50.0%15/30 | 65.5%19/29 | |
| 56.7%17/30 | 72.4%21/29 | |
| 56.7%17/30 | 65.5%19/29 | |
| 53.3%16/30 | — | |
| 60.0%18/30 | 72.4%21/29 |
Overall
Each column shows that harness's selected results on GPT-6 Astra, across the same 30 tasks. Choose a model to compare results across the suite.
Success with retries
Each bar is one harness's success on the picked model, with observed pass@1 and projected pass@2, pass@3 and pass^3. These are expected rates, not measured repeat runs. One scale whichever model is picked, sorted by pass@3.
Harnesspass@3pass^3
Codex97.9%38.0%
Hermes Agent97.9%38.0%Command Code97.9%38.0%
Claude Code97.0%32.8%
OpenCode95.9%28.1%
Pi Agent95.9%28.1%
0%50%100%
- pass@1
- pass@2
- pass@3
- pass^3
Projected retries · pass@1 observed · every harness on the picked model · sorted by pass@3