OpenBenchleaderboards

Coding-agent harness benchmarks

Compares coding-agent harnesses while holding the model and task fixed. Each board is one result-sealed bundle.

Scores are not comparable across bundles: task sets, trial counts, and timeout caps differ.

OpenBench: gpt-5.6 across 7 harnesses

kind releasedate 2026-07-21denominators matched (task, trial)matched rows 294common task/trials 42results SHA 546f167576aatask-set SHA d949d46e8aae
gpt-5.6-solcaveats disclosed
1 caveat(s) from the release page
  • Archived task definitions differ from or are unavailable in the current checkout for add-feature, make-ci-green, taskflow, terminal-bench/schemelike-metacircular-eval, webcore. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
cursor
85.7%
72.2%–93.3%
36/4242.9s5,519725,446407,06745,776
nativeharness_reported
proxy 0/42 · native 42/42
devin
81.0%
66.7%–90.0%
34/4268.4s
unavailable
proxy 0/42 · native 0/42
grokbuild
81.0%
66.7%–90.0%
34/4286.3s46,21638,7277,489368,0790
proxyproxy_measured
proxy 42/42 · native 42/42
opencode
81.0%
66.7%–90.0%
34/4260.5s42,63436,7085,927313,9160
nativevendor_split
proxy 0/42 · native 42/42
claude
76.2%
61.5%–86.5%
32/4258.8s40,25534,4555,800153,4880
proxyproxy_measured
proxy 42/42 · native 42/42
pi
76.2%
61.5%–86.5%
32/4239.7s37,60232,9164,686111,4080
proxyproxy_measured
proxy 42/42 · native 42/42
codex
73.8%
58.9%–84.7%
31/4294.6s117,107102,17414,9331,164,5110
proxyproxy_measured
proxy 42/42 · native 39/42

OpenBench: deepseek across 5 harnesses

kind releasedate 2026-07-20denominators matched (task, trial)matched rows 185common task/trials 37results SHA ec6b9d58fd3atask-set SHA 1c4805e7fcd0
deepseek-v4-flashcaveats disclosed
5 caveat(s) from the release page
  • Archived first-party matrix published without rerunning models. The source contains 225 rows across 5 harnesses, 15 tasks, and 3 trials.
  • The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
  • The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
  • Harness and host/container version stamps vary in the archive. Token columns use the complete counting-proxy lane for every displayed harness.
  • Archived task definitions differ from or are unavailable in the current checkout for add-feature, make-ci-green, taskflow, terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval, webcore. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
claude
81.1%
65.8%–90.5%
30/3747.6s78,50833,94944,5581,918,0160
proxyproxy_measured
proxy 37/37 · native 36/37
pi
75.7%
59.9%–86.6%
28/3724.5s62,94633,19029,7561,601,7740
proxyproxy_measured
proxy 37/37 · native 34/37
opencode
70.3%
54.2%–82.5%
26/3740.3s68,35437,06531,2881,529,3830
proxyproxy_measured
proxy 37/37 · native 36/37
grokbuild
67.6%
51.5%–80.4%
25/3727.7s120,12474,11546,0101,579,7450
proxyproxy_measured
proxy 37/37 · native 37/37
codex
64.9%
48.8%–78.2%
24/3727.2s86,41436,97249,4421,119,5680
proxyproxy_measured
proxy 37/37 · native 32/37

OpenBench: grok-4.5 across 4 harnesses

kind releasedate 2026-07-20denominators matched (task, trial)matched rows 156common task/trials 39results SHA 3905ac339eeatask-set SHA 1c4805e7fcd0
grok-4.5caveats disclosed
5 caveat(s) from the release page
  • Archived first-party matrix published without rerunning models. The source contains 180 rows across 4 harnesses, 15 tasks, and 3 trials.
  • The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
  • The unified board ranks only task/trial cells shared by every harness. The historical release page retains its older per-arm countable view, so its denominators differ.
  • No counting-proxy telemetry was retained for this run. Complete native split telemetry is available only where the board reports 100% native coverage.
  • Archived task definitions differ from or are unavailable in the current checkout for add-feature, make-ci-green, taskflow, terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval, webcore. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
cursor
89.7%
76.4%–95.9%
35/3925.9s36,70728,6008,106247,5040
nativeharness_reported
proxy 0/39 · native 39/39
grokbuild
87.2%
73.3%–94.4%
34/3942.3s58,60349,8458,758212,932
nativeunknown
proxy 0/39 · native 39/39
pi
79.5%
64.5%–89.2%
31/3925.3s
unavailable
proxy 0/39 · native 38/39
opencode
76.9%
61.7%–87.4%
30/3921.2s55,00845,4949,514250,7350
nativevendor_split
proxy 0/39 · native 39/39

OpenBench showcase: BYO aider vs pi vs opencode (deepseek-v4-flash)

kind communitydate 2026-07-20denominators matched (task, trial)matched rows 24common task/trials 8results SHA 677642cc0649task-set SHA 8b757f8abf38
deepseek-v4-flashcaveats disclosed
4 caveat(s) from the release page
  • Harness mode is not apples-to-apples: aider ran as a one-shot --message invocation, while pi and opencode ran full agentic loops. Token comparison reflects that difference in harness mode, not a pure like-for-like agent loop contest.
  • One aider cell excluded: 1 of 9 aider cells was infrastructure-classified and is excluded from solve-rate denominators (aider reported as 8/8). pi and opencode are 9/9.
  • Uniform total-token basis: totals include uncached input, output, and cache reads from split fields; the vendor aggregate is never used. Values are 6,192 vs 48,262 vs 111,380 tokens/solve (aider vs pi vs opencode).
  • Verifiable bundle: this release includes results.jsonl and provenance.json. Re-check digests with obench verify docs/releases/2026-07-20-aider-showcase.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
aider
100.0%
67.6%–100.0%
8/839.0s5,8562,3083,5483360
proxyproxy_measured
proxy 8/8 · native 0/8
opencode
100.0%
67.6%–100.0%
8/831.7s13,66910,3283,34291,6000
proxyproxy_measured
proxy 8/8 · native 8/8
pi
100.0%
67.6%–100.0%
8/826.4s7,7424,9282,81435,4560
proxyproxy_measured
proxy 8/8 · native 8/8

OpenBench: GLM 5.2 quarantine-safe archive

kind releasedate 2026-07-09denominators matched (task, trial)matched rows 40common task/trials 8results SHA 5434f3f8aa66task-set SHA 46075fa27b81
glm-5.2caveats disclosed
5 caveat(s) from the release page
  • Quarantine-safe derived bundle: all 15 terminal-bench/cancel-async-tasks rows were excluded because its load-sensitive checker is binding-quarantined. No model was rerun.
  • The bundled results SHA is sealed. Current-tree task checks report expected post-run definition drift; recorded task digests are retained as run-time fingerprints.
  • The retained archive lacks image digests, version-source labels, and an explicit per-row timeout field; contemporaneous version-stamp drift is documented in the bundle README.
  • The unified board ranks only task/trial cells shared by every harness. The release page also shows the per-arm countable view, so its denominators differ.
  • Archived task definitions differ from or are unavailable in the current checkout for terminal-bench/feal-differential-cryptanalysis, terminal-bench/llm-inference-batching-scheduler, terminal-bench/schemelike-metacircular-eval. The result seal and bundled evidence still verify; recorded task digests remain the run-time fingerprints.
HarnessSolve rate · Wilson 95%
050100%
SolvedMedian wallFresh tokens/solveUncached input/solveOutput/solveCache-read/solveCache-write/solveTelemetry source / basisTelemetry coverage
pi
100.0%
67.6%–100.0%
8/8693.5s
unavailable
proxy 0/8 · native 6/8
opencode
87.5%
52.9%–97.8%
7/8661.2s94,19754,76239,435973,1660
nativevendor_split
proxy 0/8 · native 8/8
claude
62.5%
30.6%–86.3%
5/8569.0s
unavailable
proxy 0/8 · native 4/8
grokbuild
62.5%
30.6%–86.3%
5/8267.6s
unavailable
proxy 0/8 · native 0/8
codex
50.0%
21.5%–78.5%
4/8801.1s
unavailable
proxy 0/8 · native 4/8

Not ranked (2)

No result-sealed results.jsonl.

  • 2026-07-02-m3
    no results.jsonl (HTML-only release page)
  • 2026-07-20-kimi-k3
    no results.jsonl (HTML-only release page)

AI gateway benchmarks

Compares request latency, throughput, and reliability across AI gateways under separately scheduled cold and warm conditions. Select a model below; each bundle keeps its own request counts and matched-block denominators.

Gateway Bench measures request-level transport and serving telemetry. Cold and warm denominators are separate and are never merged.

GPT-5.6 Sol Gateway Bench, 100 requests per route

date 2026-08-20model match rolling_aliascold blocks 50/50warm blocks 50/50requests 700verified commit 73c569ebf06aexperiment 960f93e76135

Evidence depth: 50 cold + 50 warm matched blocks per route.

Run note: Five routes, each with 50 cold and 50 warm measured requests. Warm measurements include a separate primer. All routes use GPT-5.6 Sol with medium reasoning; temperature and top_p are omitted so provider defaults apply. Gateway routes lock to OpenAI with fallbacks and response caching disabled. Cloudflare returned HTTP 402 for 25 of 100 measured requests; availability is included in its composite score.

Supplemental run note: Ramp Router adds 50 cold and 50 warm measured requests collected on 2026-08-20. Its direct-relative deltas use the fresh direct OpenAI control from the same supplemental run, not the 2026-08-18 control. Ramp returns the served GPT-5.6 Sol model ID but not an upstream-provider field; OpenBench derives OpenAI ownership from the admitted model identity. Supplemental evidence (2026-08-20): results.jsonl; verified commit b7606d1f450f; experiment 86d9f9d74624

Gateway leaderboard

OpenBench Composite compares TTFT and output throughput with the matched direct-provider arm; the previous score used fixed absolute ceilings. Direct performance equals 75 before request success is applied; higher is better, cost is excluded, and Direct OpenAI is an unranked reference.

1
OpenRouter
100% request success
77.1composite
2
Vercel
100% request success
75.1composite
3
Concentrate
100% request success
72.1composite
4
Cloudflare
75% request success
56.3composite
5
Ramp Router
100% request success
51.6composite

Cold requests

Complete blocks: 50/50. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct OpenAI1.146s / 2.406s1.920s / 3.722s0.361s / 0.785s0.361s / 0.833s51.2 / 69.8total 84.5 / 95.6
cached 0.0 / 0.0
cache write 0.0 / 0.0
Cloudflare1.394s / 2.339s2.289s / 3.518s0.708s / 1.197s0.708s / 1.198s51.5 / 63.7total 85.0 / 96.9
cached 0.0 / 0.0
cache write 0.0 / 0.0
Concentrate1.177s / 2.598s2.025s / 3.674s0.626s / 0.999s0.626s / 0.999s51.0 / 65.1total 84.5 / 96.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter1.040s / 1.931s1.887s / 2.894s0.375s / 0.774s0.503s / 0.858s51.4 / 60.4total 85.5 / 97.0
cached 0.0 / 0.0
cache write — / — (0/50)
Ramp Router1.269s / 6.410s2.755s / 7.027s1.268s / 6.408s1.268s / 6.409s37.5 / 67.1total 83.5 / 94.1
cached 0.0 / 0.0
cache write 0.0 / 0.0
Vercel1.044s / 2.490s1.970s / 3.295s1.019s / 2.480s1.043s / 2.488s50.9 / 66.5total 85.0 / 95.1
cached 0.0 / 0.0
cache write 0.0 / 0.0

Warm requests

Complete blocks: 50/50. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct OpenAI1.022s / 2.321s1.909s / 3.979s0.351s / 0.683s0.351s / 0.683s47.9 / 70.5total 85.0 / 93.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
Cloudflare1.044s / 1.817s1.966s / 2.995s0.505s / 0.743s0.505s / 0.744s49.6 / 69.0total 86.0 / 96.0
cached 0.0 / 0.0
cache write 0.0 / 0.0
Concentrate1.159s / 2.713s2.069s / 3.607s0.557s / 0.932s0.557s / 0.933s49.7 / 64.2total 87.5 / 93.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter1.068s / 2.436s1.884s / 3.558s0.410s / 0.799s0.514s / 0.799s54.0 / 73.7total 85.0 / 95.5
cached 0.0 / 0.0
cache write — / — (0/50)
Ramp Router1.693s / 6.233s2.947s / 6.790s1.692s / 6.233s1.693s / 6.233s40.1 / 85.6total 83.0 / 94.0
cached 0.0 / 0.0
cache write 0.0 / 0.0
Vercel1.087s / 2.386s1.912s / 3.782s1.074s / 2.289s1.087s / 2.295s51.3 / 69.0total 84.0 / 93.0
cached 0.0 / 0.0
cache write 0.0 / 0.0

Cold setup

Connection setup phases for cold requests only.

RouteDNS p50 / p95TCP p50 / p95TLS p50 / p95
Direct OpenAI0.013s / 0.173s0.070s / 0.218s0.078s / 0.216s
Cloudflare0.002s / 0.060s0.072s / 0.219s0.082s / 0.228s
Concentrate0.002s / 0.004s0.068s / 0.202s0.082s / 0.211s
OpenRouter0.002s / 0.078s0.070s / 0.193s0.080s / 0.199s
Ramp Router0.002s / 0.003s0.020s / 0.037s0.028s / 0.043s
Vercel0.003s / 0.157s0.071s / 0.201s0.089s / 0.208s

Paired request deltas

Every delta is gateway minus Direct OpenAI. These are latency metrics, so positive means slower/worse and negative means faster/better. Medians use complete paired blocks with bootstrap 95% intervals.

Gateway routeConditionΔ response headersΔ TTFT
Cloudflarecold
+0.307s
95% CI +0.241s to +0.393s · paired 42/50
+0.179s
95% CI -0.033s to +0.494s · paired 42/50
Cloudflarewarm
+0.159s
95% CI +0.140s to +0.221s · paired 33/50
+0.116s
95% CI -0.017s to +0.198s · paired 33/50
Concentratecold
+0.234s
95% CI +0.189s to +0.291s · paired 50/50
+0.072s
95% CI -0.149s to +0.256s · paired 50/50
Concentratewarm
+0.205s
95% CI +0.157s to +0.224s · paired 50/50
+0.162s
95% CI -0.029s to +0.433s · paired 50/50
OpenRoutercold
+0.019s
95% CI -0.004s to +0.039s · paired 50/50
-0.139s
95% CI -0.252s to +0.061s · paired 50/50
OpenRouterwarm
+0.061s
95% CI +0.015s to +0.087s · paired 50/50
+0.038s
95% CI -0.128s to +0.165s · paired 50/50
Vercelcold
+0.613s
95% CI +0.573s to +0.745s · paired 50/50
+0.008s
95% CI -0.219s to +0.089s · paired 50/50
Vercelwarm
+0.720s
95% CI +0.641s to +0.855s · paired 50/50
+0.088s
95% CI -0.080s to +0.308s · paired 50/50
Ramp Routercold
+0.812s
95% CI +0.584s to +1.336s · paired 50/50
+0.254s
95% CI +0.036s to +0.542s · paired 50/50
Ramp Routerwarm
+1.295s
95% CI +0.886s to +1.541s · paired 50/50
+0.546s
95% CI +0.279s to +0.877s · paired 50/50

Completion integrity

Provider-reported completion reasons. Natural stop means the response ended normally; length means the provider reported a length-based termination, commonly an output or context limit; missing means no finish reason was reported; other covers any remaining explicit reason. Warm-primer natural stop shows whether the separate connection-warming request ended normally.

RouteConditionMeasured natural-stopMeasured lengthMeasured missingMeasured otherWarm-primer natural-stop
Direct OpenAICold00500
Direct OpenAIWarm005000/50
CloudflareCold00500
CloudflareWarm005000/50
ConcentrateCold00500
ConcentrateWarm005000/50
OpenRouterCold00500
OpenRouterWarm005000/50
Ramp RouterCold00500
Ramp RouterWarm005000/50
VercelCold00500
VercelWarm005000/50

DeepSeek V4 Flash Gateway Bench, 100 requests per route

date 2026-08-03model match rolling_aliascold blocks 50/50warm blocks 50/50requests 500verified commit e0e5dd2978c7experiment 34873bc948a0

Evidence depth: 50 cold + 50 warm matched blocks per route.

Run note: Five routes, each with 50 cold and 50 warm measured requests. Warm measurements include a separate primer. Gateway routes lock to DeepSeek with fallbacks and response caching disabled. Token counts are route-reported; Cloudflare reported a different input-token denominator from the other routes.

Gateway leaderboard

OpenBench Composite compares TTFT and output throughput with the matched direct-provider arm; the previous score used fixed absolute ceilings. Direct performance equals 75 before request success is applied; higher is better, cost is excluded, and Direct DeepSeek is an unranked reference.

1
OpenRouter
100% request success
72.5composite
2
Concentrate
100% request success
68.2composite
3
Cloudflare
100% request success
56.4composite
4
Vercel
100% request success
38.2composite

Cold requests

Complete blocks: 50/50. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct DeepSeek0.853s / 1.144s1.272s / 1.601s0.345s / 0.424s0.356s / 0.424s79.6 / 97.3total 149.5 / 163.1
cached 0.0 / 0.0
cache write — / — (0/50)
Cloudflare1.022s / 4.238s1.619s / 5.156s1.022s / 4.237s1.022s / 4.238s100.8 / 352.3total 83.5 / 96.5
cached 0.0 / 0.0
cache write — / — (0/50)
Concentrate1.105s / 1.539s1.491s / 2.014s0.558s / 0.735s0.558s / 0.735s80.8 / 103.1total 148.0 / 162.6
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter0.978s / 1.212s1.418s / 1.760s0.446s / 0.556s0.522s / 0.579s76.9 / 103.1total 147.5 / 171.6
cached 0.0 / 0.0
cache write 0.0 / 0.0
Vercel1.563s / 6.838s1.960s / 7.135s0.502s / 0.626s0.513s / 0.631s81.9 / 106.9total 149.0 / 167.6
cached 0.0 / 0.0
cache write — / — (0/50)

Warm requests

Complete blocks: 50/50. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct DeepSeek0.864s / 1.153s1.274s / 1.630s0.338s / 0.417s0.346s / 0.424s81.8 / 102.5total 147.5 / 171.6
cached 0.0 / 0.0
cache write — / — (0/50)
Cloudflare1.002s / 2.793s1.665s / 5.138s1.002s / 2.792s1.002s / 2.792s85.2 / 306.6total 82.0 / 119.2
cached 0.0 / 0.0
cache write — / — (0/50)
Concentrate1.036s / 1.262s1.426s / 1.680s0.478s / 0.603s0.478s / 0.603s80.6 / 105.2total 148.0 / 158.0
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter0.917s / 1.172s1.390s / 1.653s0.379s / 0.477s0.462s / 0.480s78.1 / 105.8total 149.0 / 176.6
cached 0.0 / 0.0
cache write 0.0 / 0.0
Vercel1.328s / 8.706s1.786s / 9.123s0.428s / 0.527s0.437s / 0.531s81.6 / 187.6total 147.5 / 164.6
cached 0.0 / 0.0
cache write — / — (0/50)

Cold setup

Connection setup phases for cold requests only.

RouteDNS p50 / p95TCP p50 / p95TLS p50 / p95
Direct DeepSeek0.002s / 0.031s0.006s / 0.007s0.009s / 0.013s
Cloudflare0.001s / 0.021s0.012s / 0.013s0.016s / 0.019s
Concentrate0.001s / 0.002s0.012s / 0.014s0.017s / 0.022s
OpenRouter0.001s / 0.021s0.011s / 0.014s0.015s / 0.019s
Vercel0.002s / 0.043s0.006s / 0.008s0.023s / 0.027s

Paired request deltas

Every delta is gateway minus Direct DeepSeek. These are latency metrics, so positive means slower/worse and negative means faster/better. Medians use complete paired blocks with bootstrap 95% intervals.

Gateway routeConditionΔ response headersΔ TTFT
Cloudflarecold
+0.676s
95% CI +0.582s to +0.811s · paired 50/50
+0.145s
95% CI +0.053s to +0.357s · paired 50/50
Cloudflarewarm
+0.668s
95% CI +0.459s to +0.865s · paired 50/50
+0.168s
95% CI -0.047s to +0.301s · paired 50/50
Concentratecold
+0.203s
95% CI +0.185s to +0.224s · paired 50/50
+0.248s
95% CI +0.159s to +0.350s · paired 50/50
Concentratewarm
+0.143s
95% CI +0.123s to +0.156s · paired 50/50
+0.172s
95% CI +0.076s to +0.218s · paired 50/50
OpenRoutercold
+0.098s
95% CI +0.082s to +0.112s · paired 50/50
+0.129s
95% CI +0.073s to +0.216s · paired 50/50
OpenRouterwarm
+0.038s
95% CI +0.032s to +0.050s · paired 50/50
+0.058s
95% CI -0.038s to +0.113s · paired 50/50
Vercelcold
+0.132s
95% CI +0.121s to +0.160s · paired 50/50
+0.712s
95% CI +0.286s to +1.318s · paired 50/50
Vercelwarm
+0.088s
95% CI +0.069s to +0.109s · paired 50/50
+0.445s
95% CI +0.226s to +1.229s · paired 50/50

Completion integrity

Provider-reported completion reasons. Natural stop means the response ended normally; length means the provider reported a length-based termination, commonly an output or context limit; missing means no finish reason was reported; other covers any remaining explicit reason. Warm-primer natural stop shows whether the separate connection-warming request ended normally.

RouteConditionMeasured natural-stopMeasured lengthMeasured missingMeasured otherWarm-primer natural-stop
Direct DeepSeekCold50000
Direct DeepSeekWarm5000050/50
CloudflareCold50000
CloudflareWarm5000050/50
ConcentrateCold50000
ConcentrateWarm5000050/50
OpenRouterCold50000
OpenRouterWarm5000050/50
VercelCold50000
VercelWarm5000050/50

Kimi K3 Gateway Bench, 100 requests per route

date 2026-07-28model match rolling_aliascold blocks 50/50warm blocks 50/50requests 500verified commit c45441997c81experiment 9748aef6256d

Evidence depth: 50 cold + 50 warm matched blocks per route.

Run note: Five routes, each with 50 cold and 50 warm measured requests. Warm measurements include a separate primer. Routes use rolling Kimi K3 aliases.

Gateway leaderboard

OpenBench Composite compares TTFT and output throughput with the matched direct-provider arm; the previous score used fixed absolute ceilings. Direct performance equals 75 before request success is applied; higher is better, cost is excluded, and Direct Moonshot is an unranked reference.

1
OpenRouter
100% request success
75.3composite
2
Vercel
100% request success
72.3composite
3
Concentrate
100% request success
72.2composite
4
Cloudflare
100% request success
69.5composite

Cold requests

Complete blocks: 50/50. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct Moonshot2.257s / 3.493s6.932s / 12.391s2.256s / 3.493s2.256s / 3.493s34.6 / 40.9total 263.0 / 370.1
cached 115.0 / 117.5
cache write — / — (0/50)
Cloudflare2.831s / 5.256s6.788s / 13.251s2.831s / 5.256s2.831s / 5.256s34.8 / 48.0total 227.5 / 379.6
cached 116.0 / 117.5 (30/50)
cache write — / — (0/50)
Concentrate2.699s / 3.839s5.707s / 10.723s2.699s / 3.839s2.699s / 3.839s36.2 / 45.3total 224.0 / 351.0
cached 114.0 / 117.0
cache write 0.0 / 0.0
OpenRouter2.568s / 3.392s6.011s / 10.152s2.537s / 3.377s2.538s / 3.377s35.1 / 48.0total 237.0 / 353.3
cached 0.0 / 117.0
cache write 0.0 / 0.0
Vercel2.734s / 4.504s6.541s / 11.249s2.726s / 4.494s2.734s / 4.504s34.7 / 48.1total 244.5 / 351.6
cached 113.0 / 117.5
cache write — / — (0/50)

Warm requests

Complete blocks: 50/50. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct Moonshot2.351s / 3.710s6.798s / 11.104s2.350s / 3.710s2.351s / 3.710s34.2 / 56.6total 256.0 / 372.1
cached 115.0 / 119.0 (30/50)
cache write — / — (0/50)
Cloudflare2.344s / 4.020s5.966s / 9.987s2.344s / 4.020s2.344s / 4.020s34.9 / 44.0total 237.0 / 342.0
cached 116.0 / 119.0 (42/50)
cache write — / — (0/50)
Concentrate2.514s / 3.691s6.677s / 10.043s2.513s / 3.691s2.514s / 3.691s34.9 / 44.4total 251.5 / 357.9
cached 115.0 / 118.0
cache write 0.0 / 0.0
OpenRouter2.436s / 2.988s6.580s / 9.581s2.423s / 2.982s2.423s / 2.982s34.1 / 41.1total 260.5 / 347.1
cached 0.0 / 118.0
cache write 0.0 / 0.0
Vercel2.459s / 3.138s6.320s / 10.413s2.459s / 3.132s2.459s / 3.138s35.1 / 46.0total 252.0 / 361.2
cached 115.0 / 119.0
cache write — / — (0/50)

Cold setup

Connection setup phases for cold requests only.

RouteDNS p50 / p95TCP p50 / p95TLS p50 / p95
Direct Moonshot0.002s / 0.037s0.012s / 0.013s0.018s / 0.022s
Cloudflare0.002s / 0.021s0.011s / 0.013s0.016s / 0.020s
Concentrate0.001s / 0.056s0.011s / 0.012s0.018s / 0.021s
OpenRouter0.001s / 0.022s0.011s / 0.012s0.016s / 0.019s
Vercel0.002s / 0.135s0.005s / 0.007s0.024s / 0.053s

Paired request deltas

Every delta is gateway minus Direct Moonshot. These are latency metrics, so positive means slower/worse and negative means faster/better. Medians use complete paired blocks with bootstrap 95% intervals.

Gateway routeConditionΔ response headersΔ TTFT
Cloudflarecold
+0.413s
95% CI +0.222s to +0.674s · paired 50/50
+0.412s
95% CI +0.222s to +0.674s · paired 50/50
Cloudflarewarm
+0.067s
95% CI -0.068s to +0.159s · paired 50/50
+0.067s
95% CI -0.068s to +0.159s · paired 50/50
Concentratecold
+0.249s
95% CI +0.086s to +0.393s · paired 50/50
+0.251s
95% CI +0.077s to +0.393s · paired 50/50
Concentratewarm
+0.236s
95% CI +0.110s to +0.281s · paired 50/50
+0.236s
95% CI +0.110s to +0.303s · paired 50/50
OpenRoutercold
+0.259s
95% CI +0.071s to +0.409s · paired 50/50
+0.271s
95% CI +0.090s to +0.420s · paired 50/50
OpenRouterwarm
+0.187s
95% CI -0.065s to +0.347s · paired 50/50
+0.193s
95% CI -0.077s to +0.378s · paired 50/50
Vercelcold
+0.297s
95% CI +0.118s to +0.492s · paired 50/50
+0.306s
95% CI +0.127s to +0.498s · paired 50/50
Vercelwarm
+0.117s
95% CI -0.056s to +0.271s · paired 50/50
+0.122s
95% CI -0.060s to +0.282s · paired 50/50

Completion integrity

Provider-reported completion reasons. Natural stop means the response ended normally; length means the provider reported a length-based termination, commonly an output or context limit; missing means no finish reason was reported; other covers any remaining explicit reason. Warm-primer natural stop shows whether the separate connection-warming request ended normally.

RouteConditionMeasured natural-stopMeasured lengthMeasured missingMeasured otherWarm-primer natural-stop
Direct MoonshotCold50000
Direct MoonshotWarm5000050/50
CloudflareCold50000
CloudflareWarm5000050/50
ConcentrateCold50000
ConcentrateWarm5000050/50
OpenRouterCold50000
OpenRouterWarm5000050/50
VercelCold50000
VercelWarm5000050/50

GPT-4o mini Gateway Bench

date 2026-07-27model match rolling_aliascold blocks 30/30warm blocks 30/30requests 300verified commit 6d1de84d6c96experiment e3649dbc6dc4

Evidence depth: 30 cold + 30 warm matched blocks per route.

Gateway leaderboard

OpenBench Composite compares TTFT and output throughput with the matched direct-provider arm; the previous score used fixed absolute ceilings. Direct performance equals 75 before request success is applied; higher is better, cost is excluded, and Direct OpenAI is an unranked reference.

1
OpenRouter
100% request success
85.7composite
2
Cloudflare
100% request success
75.9composite
3
Concentrate
100% request success
71.2composite
4
Vercel
100% request success
62.2composite

Cold requests

Complete blocks: 30/30. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct OpenAI0.527s / 1.356s0.735s / 1.763s0.232s / 0.887s0.232s / 0.888s92.3 / 123.2total 51.0 / 54.0
cached 0.0 / 0.0
cache write 0.0 / 0.0
Cloudflare0.656s / 1.055s0.812s / 1.215s0.341s / 0.767s0.341s / 0.767s84.6 / 113.7total 51.0 / 53.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
Concentrate0.763s / 0.946s0.930s / 1.803s0.404s / 0.538s0.404s / 0.538s84.9 / 123.0total 51.0 / 54.0
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter0.471s / 0.675s0.662s / 0.845s0.459s / 0.668s0.470s / 0.668s100.6 / 156.9total 54.0 / 56.5
cached 0.0 / 0.0
cache write — / — (0/30)
Vercel0.789s / 2.494s0.951s / 2.719s0.773s / 2.485s0.783s / 2.490s87.9 / 345.3total 51.0 / 54.0
cached 0.0 / 0.0
cache write 0.0 / 0.0

Warm requests

Complete blocks: 30/30. TTFT begins when the measured request is sent; connection setup is reported separately.

RouteTTFT p50 / p95Stream total p50 / p95Response headers p50 / p95First body byte p50 / p95Throughput tok/s p50 / p95Total / cached / cache-write tokens p50 / p95
Direct OpenAI0.504s / 0.977s0.656s / 1.221s0.184s / 0.436s0.184s / 0.437s95.1 / 127.2total 52.0 / 54.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
Cloudflare0.561s / 0.739s0.754s / 0.893s0.260s / 0.410s0.260s / 0.410s96.6 / 126.4total 52.0 / 54.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
Concentrate0.704s / 0.933s0.896s / 1.096s0.361s / 0.436s0.361s / 0.436s85.3 / 131.6total 52.0 / 54.5
cached 0.0 / 0.0
cache write 0.0 / 0.0
OpenRouter0.433s / 0.655s0.587s / 0.764s0.423s / 0.485s0.432s / 0.632s105.9 / 415.5total 54.5 / 57.0
cached 0.0 / 0.0
cache write — / — (0/30)
Vercel0.572s / 1.605s0.726s / 2.458s0.559s / 1.571s0.572s / 1.601s94.4 / 449.2total 52.5 / 55.0
cached 0.0 / 0.0
cache write 0.0 / 0.0

Cold setup

Connection setup phases for cold requests only.

RouteDNS p50 / p95TCP p50 / p95TLS p50 / p95
Direct OpenAI0.002s / 0.003s0.011s / 0.012s0.016s / 0.023s
Cloudflare0.001s / 0.002s0.011s / 0.012s0.017s / 0.023s
Concentrate0.001s / 0.002s0.011s / 0.014s0.018s / 0.023s
OpenRouter0.001s / 0.002s0.011s / 0.012s0.016s / 0.020s
Vercel0.002s / 0.148s0.006s / 0.007s0.024s / 0.028s

Paired request deltas

Every delta is gateway minus Direct OpenAI. These are latency metrics, so positive means slower/worse and negative means faster/better. Medians use complete paired blocks with bootstrap 95% intervals.

Gateway routeConditionΔ response headersΔ TTFT
Cloudflarecold
+0.090s
95% CI +0.072s to +0.120s · paired 30/30
+0.095s
95% CI +0.011s to +0.163s · paired 30/30
Cloudflarewarm
+0.081s
95% CI +0.050s to +0.112s · paired 30/30
+0.037s
95% CI +0.003s to +0.116s · paired 30/30
Concentratecold
+0.169s
95% CI +0.143s to +0.190s · paired 30/30
+0.193s
95% CI +0.097s to +0.235s · paired 30/30
Concentratewarm
+0.168s
95% CI +0.143s to +0.204s · paired 30/30
+0.202s
95% CI +0.142s to +0.230s · paired 30/30
OpenRoutercold
+0.233s
95% CI +0.212s to +0.244s · paired 30/30
-0.088s
95% CI -0.176s to -0.009s · paired 30/30
OpenRouterwarm
+0.230s
95% CI +0.202s to +0.252s · paired 30/30
-0.054s
95% CI -0.109s to -0.018s · paired 30/30
Vercelcold
+0.508s
95% CI +0.440s to +0.596s · paired 30/30
+0.227s
95% CI +0.171s to +0.297s · paired 30/30
Vercelwarm
+0.348s
95% CI +0.328s to +0.465s · paired 30/30
+0.058s
95% CI +0.014s to +0.214s · paired 30/30

Not published (1)

Did not pass public Gateway Bench verification.

  • 2026-07-28-kimi-k3-managed-30
    bundle verification failed: public manifest does not match schema

OpenBench benchmark results

Digest-verified benchmark releases, methodology, and project information.

Releases

First-party bundles.

Packs

Versioned task and harness packs.

  • openbench/core-smoke@1.0.0 tasks
    Apache-2.0 · data/packs/openbench-core-smoke · 3b1e576398a8
    Tiny polarity-checked smoke tasks (make-it-run, fix-failing-test)

OpenBench benchmark results

Digest-verified benchmark releases, methodology, and project information.

What is being measured

OpenBench runs two benchmark families. They share a task contract and a checker, and no denominators.

Harness Bench

Varies the coding-agent harness — the CLI that wraps a model in a run loop, tool set, and permission policy — while holding the model and task fixed. An arm is (harness, model). A task is solved when its checker.sh exits 0; the harness's own claim of success is never trusted.

Gateway Bench

Measures one model request at a time under separately scheduled cold and warm transport conditions. It reports request success, route verification, transport and stream phase timing, throughput, usage, and per-request cost. It is not a coding-agent outcome benchmark. Gateway Bench requests and Harness Bench cells are never pooled or compared as one denominator.

The Gateway Bench leaderboard scores each gateway relative to the matched direct-provider arm for the same model; the previous score used fixed absolute latency and throughput ceilings. Direct performance anchors each metric at 75, with 25% cold TTFT median, 20% cold TTFT p95, 25% warm TTFT median, 20% warm TTFT p95, and 10% warm median output throughput; the weighted result is multiplied by request success. TTFT starts when the measured request is sent; cold DNS, TCP, and TLS setup are reported separately. Cost is excluded, the direct-provider arm is unranked, and the detailed measurements below the score remain the factual record.

Denominators and intervals

  • Harness Bench denominators are countable cells. Infrastructure and rate-limit failures are excluded; other failures, including timeouts, stay in the denominator.
  • Harness Bench uses Wilson 95% intervals over matched (task, trial) cells whenever a bundle has two or more arms.
  • Gateway Bench displays complete cold and warm block counts separately. Availability uses a Wilson 95% interval over every attempted measured request, so gateway errors such as HTTP 429 responses and timeouts remain in that denominator. Phase summaries use successful, route-verified requests and retain metric-specific coverage. Paired deltas use complete gateway/direct blocks and bootstrap 95% intervals.

Efficiency and cost

  • Median wall time is taken among solved cells only.
  • Each Harness Bench arm uses one complete split-token lane across all matched result rows, preferring proxy telemetry and otherwise using native telemetry. Fresh tokens are uncached input plus output; cache reads and cache writes remain separate. Incomplete lanes produce no token metrics and report their row coverage instead.
  • Each per-solve token figure sums traffic from every matched attempt, including failed attempts, then divides by the number of solved cells. It measures attempted traffic required per solve, not the average size of successful attempts alone.
  • Harness $/solve appears only for models with a configured price.
  • Gateway Bench response headers are the time until HTTP response headers. First body byte and semantic TTFT are reported separately; response headers are not labeled TTFB.
  • Gateway Bench measured cost is the frozen-list request estimate. Charged cost is separately reported billing evidence. Each retains its own request coverage, as do total, cached-input, and cache-write token readings.
  • Harness defaults are not clamped.

Comparability

  • Cells from different bundles are never blended. Each board is one bundle; cross-bundle ranking on different task sets is not supported.
  • Every ranked bundle ships results.jsonl plus a provenance digest and is re-verified before it appears here. Digests show tamper-evidence, not absence of cherry-picking.
  • Results cover only the included tasks, trials, model deployments, harness versions, and timeout caps.

Reproducing a board

Every board links its results.jsonl. Re-check a bundle with obench verify <bundle> (harness) or obench gateway probe verify <bundle> (gateway), and rebuild this page with obench site build.

OpenBench benchmark results

Digest-verified benchmark releases, methodology, and project information.

Contact

Want to add a gateway or harness, submit results, report a problem, or share an idea? Reach Matthew through either channel.