Compare harnesses

Every harness on the bench, side by side on the same 30 tasks.

By model

The full matrix. Cell tint scales with the value; the strongest tint wins the row.

HarnessDeepSeek V4 ProGPT-6 Astra
Claude CodeAnthropic50.0%15/3069.0%20/29
CodexOpenAI53.3%16/3072.4%21/29
OpenCodeSST50.0%15/3065.5%19/29
Hermes AgentNous Research56.7%17/3072.4%21/29
Pi Agentpi.dev56.7%17/3065.5%19/29
DeepSeekDeepSeek53.3%16/30—
Command CodeCommand Code60.0%18/3072.4%21/29

Overall

Each column shows that harness's selected results on GPT-6 Astra, across the same 30 tasks. Choose a model to compare results across the suite.

0.0%
Claude Code
0.0%
Codex
0.0%
OpenCode
0.0%
Hermes Agent
0.0%
Pi Agent
0.0%
Command Code

Success with retries

Each bar is one harness's success on the picked model, with observed pass@1 and projected pass@2, pass@3 and pass^3. These are expected rates, not measured repeat runs. One scale whichever model is picked, sorted by pass@3.

Harnesspass@3pass^3
  1. Codex97.9%38.0%
  2. Hermes Agent97.9%38.0%
  3. Command Code97.9%38.0%
  4. Claude Code97.0%32.8%
  5. OpenCode95.9%28.1%
  6. Pi Agent95.9%28.1%
0%50%100%
  • pass@1
  • pass@2
  • pass@3
  • pass^3

Projected retries · pass@1 observed · every harness on the picked model · sorted by pass@3