Harnesses
The agent around the model: tool loop, retry policy, context strategy, verifier. The same model performs differently across harnesses, so every harness gets its own ranking.
Every harness on the bench
#NameModelsBest successBest time / task
- 01
Codex
72%Best time / task151s - 02
Hermes Agent72%Best time / task122s - 03
Command Code
72%Best time / task159s - 04
Claude Code
69%Best time / task194s - 05
OpenCode
66%Best time / task126s - 06
Pi Agent
66%Best time / task127s - 07
DeepSeeklocked
53%Best time / task388s
Ranked by models run, then best success · locked = vendor-locked to its own models · Composio Benchmark