Coding-agent harness benchmarks
Compares coding-agent harnesses while holding the model and task fixed. Each board is one result-sealed bundle.
Scores are not comparable across bundles: task sets, trial counts, and timeout caps differ.
OpenBench: gpt-5.6 across 7 harnesses
1 caveat(s) from the release page
- Archived task definitions differ from or are unavailable in the current checkout for add-feature, make-ci-green, taskflow, terminal-bench/schemelike-metacircular-eval, webcore. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| cursor | 85.7% 72.2%–93.3% | 36/42 | 42.9s | 5,519 | 72 | 5,446 | 407,067 | 45,776 | nativeharness_reported | proxy 0/42 · native 42/42 |
| devin | 81.0% 66.7%–90.0% | 34/42 | 68.4s | — | — | — | — | — | unavailable | proxy 0/42 · native 0/42 |
| grokbuild | 81.0% 66.7%–90.0% | 34/42 | 86.3s | 46,216 | 38,727 | 7,489 | 368,079 | 0 | proxyproxy_measured | proxy 42/42 · native 42/42 |
| opencode | 81.0% 66.7%–90.0% | 34/42 | 60.5s | 42,634 | 36,708 | 5,927 | 313,916 | 0 | nativevendor_split | proxy 0/42 · native 42/42 |
| claude | 76.2% 61.5%–86.5% | 32/42 | 58.8s | 40,255 | 34,455 | 5,800 | 153,488 | 0 | proxyproxy_measured | proxy 42/42 · native 42/42 |
| pi | 76.2% 61.5%–86.5% | 32/42 | 39.7s | 37,602 | 32,916 | 4,686 | 111,408 | 0 | proxyproxy_measured | proxy 42/42 · native 42/42 |
| codex | 73.8% 58.9%–84.7% | 31/42 | 94.6s | 117,107 | 102,174 | 14,933 | 1,164,511 | 0 | proxyproxy_measured | proxy 42/42 · native 39/42 |
OpenBench: deepseek across 5 harnesses
5 caveat(s) from the release page
- Archived first-party matrix published without rerunning models. The source contains 225 rows across 5 harnesses, 15 tasks, and 3 trials.
- The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
- The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
- Harness and host/container version stamps vary in the archive. Token columns use the complete counting-proxy lane for every displayed harness.
- Archived task definitions differ from or are unavailable in the current checkout for add-feature, make-ci-green, taskflow, terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval, webcore. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| claude | 81.1% 65.8%–90.5% | 30/37 | 47.6s | 78,508 | 33,949 | 44,558 | 1,918,016 | 0 | proxyproxy_measured | proxy 37/37 · native 36/37 |
| pi | 75.7% 59.9%–86.6% | 28/37 | 24.5s | 62,946 | 33,190 | 29,756 | 1,601,774 | 0 | proxyproxy_measured | proxy 37/37 · native 34/37 |
| opencode | 70.3% 54.2%–82.5% | 26/37 | 40.3s | 68,354 | 37,065 | 31,288 | 1,529,383 | 0 | proxyproxy_measured | proxy 37/37 · native 36/37 |
| grokbuild | 67.6% 51.5%–80.4% | 25/37 | 27.7s | 120,124 | 74,115 | 46,010 | 1,579,745 | 0 | proxyproxy_measured | proxy 37/37 · native 37/37 |
| codex | 64.9% 48.8%–78.2% | 24/37 | 27.2s | 86,414 | 36,972 | 49,442 | 1,119,568 | 0 | proxyproxy_measured | proxy 37/37 · native 32/37 |
OpenBench: grok-4.5 across 4 harnesses
5 caveat(s) from the release page
- Archived first-party matrix published without rerunning models. The source contains 180 rows across 4 harnesses, 15 tasks, and 3 trials.
- The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
- The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
- No counting-proxy telemetry was retained for this run. Complete native split telemetry is available only where the board reports 100% native coverage.
- Archived task definitions differ from or are unavailable in the current checkout for add-feature, make-ci-green, taskflow, terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval, webcore. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| cursor | 89.7% 76.4%–95.9% | 35/39 | 25.9s | 36,707 | 28,600 | 8,106 | 247,504 | 0 | nativeharness_reported | proxy 0/39 · native 39/39 |
| grokbuild | 87.2% 73.3%–94.4% | 34/39 | 42.3s | 58,603 | 49,845 | 8,758 | 212,932 | — | nativeunknown | proxy 0/39 · native 39/39 |
| pi | 79.5% 64.5%–89.2% | 31/39 | 25.3s | — | — | — | — | — | unavailable | proxy 0/39 · native 38/39 |
| opencode | 76.9% 61.7%–87.4% | 30/39 | 21.2s | 55,008 | 45,494 | 9,514 | 250,735 | 0 | nativevendor_split | proxy 0/39 · native 39/39 |
OpenBench showcase: BYO aider vs pi vs opencode (deepseek-v4-flash)
4 caveat(s) from the release page
- Harness mode is not apples-to-apples: aider ran as a one-shot --message invocation, while pi and opencode ran full agentic loops. Token comparison reflects that difference in harness mode, not a pure like-for-like agent loop contest.
- One aider cell excluded: 1 of 9 aider cells was infrastructure-classified and is excluded from solve-rate denominators (aider reported as 8/8). pi and opencode are 9/9.
- Uniform total-token basis: totals include uncached input, output, and cache reads from split fields; the vendor aggregate is never used. Values are 6,192 vs 48,262 vs 111,380 tokens/solve (aider vs pi vs opencode).
- Verifiable bundle: this release includes results.jsonl and provenance.json. Re-check digests with obench verify docs/releases/2026-07-20-aider-showcase.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| aider | 100.0% 67.6%–100.0% | 8/8 | 39.0s | 5,856 | 2,308 | 3,548 | 336 | 0 | proxyproxy_measured | proxy 8/8 · native 0/8 |
| opencode | 100.0% 67.6%–100.0% | 8/8 | 31.7s | 13,669 | 10,328 | 3,342 | 91,600 | 0 | proxyproxy_measured | proxy 8/8 · native 8/8 |
| pi | 100.0% 67.6%–100.0% | 8/8 | 26.4s | 7,742 | 4,928 | 2,814 | 35,456 | 0 | proxyproxy_measured | proxy 8/8 · native 8/8 |
OpenBench: GLM 5.2 quarantine-safe archive
5 caveat(s) from the release page
- Quarantine-safe derived bundle: all 15 terminal-bench/cancel-async-tasks rows were excluded because its load-sensitive checker is binding-quarantined. No model was rerun.
- The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
- The retained archive lacks image digests, version-source labels, and an explicit per-row timeout field; contemporaneous version-stamp drift is documented in the bundle README.
- The unified board ranks only task/trial cells shared by every harness. The release page also shows the per-arm countable view, so its denominators differ.
- Archived task definitions differ from or are unavailable in the current checkout for terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
| Harness↕ | Solve rate · Wilson 95%↓ 050100% | Solved↕ | Median wall↕ | Fresh tokens/solve↕ | Uncached input/solve↕ | Output/solve↕ | Cache-read/solve↕ | Cache-write/solve↕ | Telemetry source / basis | Telemetry coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| pi | 100.0% 67.6%–100.0% | 8/8 | 693.5s | — | — | — | — | — | unavailable | proxy 0/8 · native 6/8 |
| opencode | 87.5% 52.9%–97.8% | 7/8 | 661.2s | 94,197 | 54,762 | 39,435 | 973,166 | 0 | nativevendor_split | proxy 0/8 · native 8/8 |
| claude | 62.5% 30.6%–86.3% | 5/8 | 569.0s | — | — | — | — | — | unavailable | proxy 0/8 · native 4/8 |
| grokbuild | 62.5% 30.6%–86.3% | 5/8 | 267.6s | — | — | — | — | — | unavailable | proxy 0/8 · native 0/8 |
| codex | 50.0% 21.5%–78.5% | 4/8 | 801.1s | — | — | — | — | — | unavailable | proxy 0/8 · native 4/8 |
No boards match the current filters.
Not ranked (2)
No result-sealed results.jsonl.
- 2026-07-02-m3no results.jsonl (HTML-only release page)
- 2026-07-20-kimi-k3no results.jsonl (HTML-only release page)