Self-reported. Reproducible.
Every number on the home page comes from one run of the script at the bottom of this page, pointed at four providers from the same machine, in the same hour, with the same prompts. Nothing is averaged across runs or cherry-picked from several. If you get different numbers, we want to hear about it: contact@earthruntime.com.
provocapier/docs/bench-20260923.md at commit f4d6037Results
Our fleet
| Model | Aggregate | Per stream | TTFT p50 | TTFT p95 | Latency p50 | Latency p95 | Errors |
|---|---|---|---|---|---|---|---|
qwen3.6-35b | 2,464 tok/s | 78.0 tok/s | 765 ms | 1,481 ms | 2.47 s | 3.56 s | 0% |
gpt-oss-120b | 2,276 tok/s | 52.8 tok/s | 537 ms | 1,224 ms | 3.11 s | 4.57 s | 0% |
qwen3.8-27b | 1,608 tok/s | 30.1 tok/s | 508 ms | 1,095 ms | 4.70 s | 8.59 s | 0% |
Reasoning was off for both Qwen models (the server default) and on at the provider default effort for gpt-oss-120b. Zero errors across all 3,104 requests. | |||||||
Every model in our fleet has an OpenRouter fallback configured, so a flapping worker could silently turn one of our rows into an OpenRouter row. Each model's runs were bracketed by a read of its own workers' request counters to prove that did not happen: 796, 794, 784 and 784 requests landed on our own hardware against ~776 expected, the small excess being health-check traffic. All four passed.
Against the marketplaces
| Endpoint | Aggregate | TTFT p50 | Latency p50 | Latency p95 | Errors | Dominant error |
|---|---|---|---|---|---|---|
| earthruntime | 2,464 tok/s | 765 ms | 2.47 s | 3.56 s | 0% | none |
| OpenRouter, default routing | 1,534 tok/s | 541 ms | 1.65 s | 4.68 s | 0% | none |
| OpenRouter to AtlasCloud | 1,327 tok/s | 887 ms | 1.87 s | 3.18 s | 0% | none |
| OpenRouter to Parasail | 3,042 tok/s | 560 ms | 2.18 s | 2.74 s | 17.6% | HTTP 429, 100% of its failures |
| Parasail's throughput is the highest number in this table and it is not a win: aggregate throughput is successful tokens divided by wall clock, and an endpoint that refuses 17.6% of the load finishes a smaller job sooner. The same 429 pattern showed up in our 2026-06 and 2026-07 runs. | ||||||
On the other three models, against OpenRouter's default routing: we are ahead on gpt-oss-120b (2,276 vs 1,438 tok/s) and behind on qwen3.8-27b (1,608 vs 1,786 tok/s).
Where we do not lead
OpenRouter's default routing beats us on median latency, 1.65 s against our 2.47 s. We are ahead on the p95 tail (3.56 s against 4.68 s), on aggregate throughput, and on not dropping requests. If your workload is one request at a time and you care about the average case, that gap is real and it is in their favour. If it is many requests at once and you care about the worst case, it is in ours.
OpenRouter is also ahead on qwen3.8-27b throughput. We publish it in the table above rather than omitting the row.
Test environment
- Request origin: a single AWS EC2
t3.xlarge(4 vCPU, 15 GB) inus-east-2c, outside our own network. The same machine issued every request to every endpoint. - Client: Python 3.12.3,
httpx0.26.0, asyncio, streaming. - Concurrency: a fixed pool of 64 requests in flight for the whole run. Not a ramp, not a burst.
- Volume: 256 requests per run, 3 runs per endpoint, median of the 3 reported for every metric. An 8-request warmup at c=4 before each endpoint, not counted.
- Workload: 32 unique prompts cycled in order,
max_tokens128. Every endpoint saw the identical sequence. - Timeout: 300 s. A request still open at 300 s is an error.
- Marketplace settings: OpenRouter on default routing, plus two runs pinned to named providers. Qwen rows sent
reasoning: {enabled: false}to match our own default-off configuration. Slugs:qwen/qwen3.6-35b-a3b,qwen/qwen3.8-27b,openai/gpt-oss-120b.
Why our TTFT is 765 ms here and 73 ms on your laptop
Time to first token in these tables is measured at 64 concurrent streams, which is the worst case in this document, not what a single caller experiences. At one stream it is 73 ms; at 16 it is 204 ms. Quote the number at the concurrency you care about.
We decomposed the 765 ms rather than leaving it as a mystery, using vLLM's own histograms on either side of a run:
| Layer | TTFT |
|---|---|
| Worker side, from vLLM's histogram | 94.5 ms |
| of which queued | 0.0 ms |
| of which prefill | 71.2 ms |
| Client TTFB, 64 parallel curl, no Python | 419 ms |
| Client TTFT, our Python asyncio harness | 765 ms |
| Requests spend zero time queued and prefill is about 71 ms, so the models, the GPUs and the parallelism are not the cause. Roughly 325 ms is added by the gateway and a further 300 ms by the benchmark client itself at 64 streams. Loopback (768 ms) and the remote vantage (765 ms) agree, so the internet path is not implicated either. | |
An earlier version of this page published 196 ms. That figure was measured against a different stack: the FastAPI gateway that was retired on 2026-09-22. The two numbers do not describe the same system, so we replaced it rather than defending it.
What got worse
| Metric | Published before | This run | Verdict |
|---|---|---|---|
| Aggregate throughput | 2,237 tok/s | 2,464 tok/s | better by 10% |
| Error rate | 0% | 0% | unchanged |
| Latency p95 | 3.6 s | 3.56 s | unchanged |
| Latency p50 | 2.2 s | 2.47 s | worse by 12% |
| TTFT p50 | 196 ms | 765 ms | worse by 3.9x, see above |
Definitions
- Latency โ wall-clock from sending the request to receiving the final stream chunk, successful requests only.
- Time to first token (TTFT) โ wall-clock from sending the request to the first stream chunk carrying
content(orreasoning_content). - Median / P95 โ the 50th and 95th percentile of the per-request distribution, linearly interpolated.
- Aggregate throughput โ total completion tokens across successful requests divided by the wall-clock of the whole run. Token counts come from the provider's
usageblock when it streams one; otherwise one per content chunk, and the output says which was used. - Error rate โ errored requests รท total requests. Defined below.
What counts as an error
Any request that did not return a complete, well-formed answer. Concretely, in is_error():
- a transport failure โ connection reset, TLS failure, or the 120 s timeout;
- a non-2xx HTTP status (429 rate-limits and 5xx alike โ the caller can't use either);
- a 2xx that streamed nothing;
- a stream that ended without a
finish_reason(lengthcounts as complete: the request asked for a cap and got it).
An errored request contributes no latency sample and no tokens. That is deliberate: a fast 429 would otherwise improve a provider's median.
Reproduce it
Get a key from the home page, then run the same harness that produced every number above. It needs the 32-prompt workload file and writes one JSON summary per endpoint:
# Python 3.12, httpx 0.26 pip install httpx echo "$EARTHRUNTIME_KEY" > key.txt python bench64.py --base-url https://api.earthruntime.com/v1 --key-file key.txt \ --model qwen3.6-35b --label ours_qwen36 \ --prompts bench_prompts.json --out ours_qwen36.json \ --concurrency 64 --requests 256 --max-tokens 128 --runs 3 # the same run against a marketplace echo "$OPENROUTER_KEY" > or_key.txt python bench64.py --base-url https://openrouter.ai/api/v1 --key-file or_key.txt \ --model "qwen/qwen3.6-35b-a3b" --label or_qwen36 \ --prompts bench_prompts.json --out or_qwen36.json \ --concurrency 64 --requests 256 --max-tokens 128 --runs 3 # pin a named provider instead of default routing python bench64.py ... --provider AtlasCloud
A 256-request run at 128 max tokens costs a few cents on our side, so the free grant covers several. To match our Qwen rows exactly, add --extra '{"reasoning":{"enabled":false}}' on the OpenRouter runs: our Qwen models serve with reasoning off by default and the comparison is otherwise unfair to them.
The script
bench64.py, verbatim, exactly as run. Standard library plus httpx.
#!/usr/bin/env python3
"""c=64 streaming benchmark, matched to bench_results_overnight_20260605T031811Z.
Shape of record: concurrency 64, 256 requests, max_tokens 128, 32 unique prompts
(scripts/bench.py's 20 VARIED + 12 CODING), warmup first, N runs, median across runs.
Streams every response so TTFT is real. Output tokens = content chunks (same
convention as the published CSVs, where output_tokens == chunks).
"""
from __future__ import annotations
import argparse, asyncio, json, statistics, sys, time
import httpx
def pct(xs, p):
if not xs: return None
xs = sorted(xs); k = (len(xs) - 1) * p
f = int(k); c = min(f + 1, len(xs) - 1)
return xs[f] + (xs[c] - xs[f]) * (k - f)
async def one(client, url, headers, body, idx):
t0 = time.perf_counter(); ttft = None; chunks = 0; usage = None; err = None; hdrs = None
try:
async with client.stream("POST", url, headers=headers, json=body) as r:
hdrs = {k.lower(): v for k, v in r.headers.items() if k.lower().startswith(("x-", "server"))}
if r.status_code != 200:
err = f"http{r.status_code}:{(await r.aread())[:120].decode('utf-8','replace')}"
else:
async for line in r.aiter_lines():
if not line.startswith("data:"): continue
data = line[5:].strip()
if data == "[DONE]": break
try: obj = json.loads(data)
except Exception: continue
if obj.get("usage"): usage = obj["usage"]
for ch in obj.get("choices") or []:
d = ch.get("delta") or {}
piece = d.get("content") or d.get("reasoning_content") or d.get("reasoning") or ""
if piece:
if ttft is None: ttft = (time.perf_counter() - t0) * 1000
chunks += 1
except Exception as e:
err = f"{type(e).__name__}:{str(e)[:120]}"
total = (time.perf_counter() - t0) * 1000
out_tok = (usage or {}).get("completion_tokens") or chunks
return {"idx": idx, "ttft_ms": ttft, "total_ms": total, "output_tokens": out_tok,
"chunks": chunks, "err": err, "headers": hdrs if idx == 0 else None}
async def run_once(url, headers, model, prompts, n, conc, max_tokens, extra):
sem = asyncio.Semaphore(conc)
limits = httpx.Limits(max_connections=conc + 8, max_keepalive_connections=conc + 8)
async with httpx.AsyncClient(timeout=httpx.Timeout(300.0, connect=30.0), limits=limits) as client:
async def guarded(i):
async with sem:
body = {"model": model, "messages": [{"role": "user", "content": prompts[i % len(prompts)]}],
"max_tokens": max_tokens, "temperature": 0, "stream": True,
"stream_options": {"include_usage": True}}
body.update(extra)
return await one(client, url, headers, body, i)
t0 = time.perf_counter()
rows = await asyncio.gather(*[guarded(i) for i in range(n)])
wall = time.perf_counter() - t0
return rows, wall
def summarize(rows, wall):
ok = [r for r in rows if not r["err"] and r["ttft_ms"] is not None]
bad = [r for r in rows if r["err"] or r["ttft_ms"] is None]
ttfts = [r["ttft_ms"] for r in ok]; totals = [r["total_ms"] for r in ok]
per_stream = [r["output_tokens"] / max(1e-6, (r["total_ms"] - r["ttft_ms"]) / 1000) for r in ok if r["output_tokens"] and r["total_ms"] > r["ttft_ms"]]
return {"n": len(rows), "ok": len(ok), "errors": len(bad), "error_rate": len(bad) / len(rows) if rows else None,
"wall_s": wall, "out_tokens_total": sum(r["output_tokens"] for r in ok),
"agg_out_tok_s": sum(r["output_tokens"] for r in ok) / wall if wall else None,
"per_stream_tok_s_median": statistics.median(per_stream) if per_stream else None,
"ttft_ms_median": statistics.median(ttfts) if ttfts else None, "ttft_ms_p95": pct(ttfts, .95),
"latency_ms_median": statistics.median(totals) if totals else None, "latency_ms_p95": pct(totals, .95),
"err_samples": [r["err"] for r in bad][:3]}
async def main():
ap = argparse.ArgumentParser()
ap.add_argument("--base-url", required=True); ap.add_argument("--key-file", required=True)
ap.add_argument("--model", required=True); ap.add_argument("--label", required=True)
ap.add_argument("--runs", type=int, default=3); ap.add_argument("--concurrency", type=int, default=64)
ap.add_argument("--requests", type=int, default=256); ap.add_argument("--max-tokens", type=int, default=128)
ap.add_argument("--provider", default=None, help="OpenRouter provider order, e.g. 'AtlasCloud'")
ap.add_argument("--extra", default="{}", help="extra JSON merged into each body")
ap.add_argument("--prompts", required=True); ap.add_argument("--out", required=True)
a = ap.parse_args()
key = open(a.key_file).read().strip()
prompts = json.load(open(a.prompts))
url = a.base_url.rstrip("/") + "/chat/completions" if a.base_url.rstrip("/").endswith("/v1") else a.base_url.rstrip("/") + "/v1/chat/completions"
headers = {"Authorization": f"Bearer {key}", "Content-Type": "application/json"}
extra = json.loads(a.extra)
if a.provider: extra["provider"] = {"order": [a.provider], "allow_fallbacks": False}
print(f"=== {a.label} | {a.model} | c={a.concurrency} n={a.requests} max_tokens={a.max_tokens} runs={a.runs} ===", flush=True)
print(f" warmup (c=4, n=8)...", flush=True)
await run_once(url, headers, a.model, prompts, 8, 4, a.max_tokens, extra)
runs = []
for i in range(a.runs):
rows, wall = await run_once(url, headers, a.model, prompts, a.requests, a.concurrency, a.max_tokens, extra)
s = summarize(rows, wall); s["run"] = i + 1; runs.append(s)
print(f" run {i+1}: agg {s['agg_out_tok_s']:.0f} tok/s | TTFT p50 {s['ttft_ms_median']:.0f}ms p95 {s['ttft_ms_p95']:.0f}ms | "
f"lat p50 {s['latency_ms_median']/1000:.2f}s p95 {s['latency_ms_p95']/1000:.2f}s | err {s['error_rate']*100:.1f}% | wall {wall:.1f}s"
if s["ok"] else f" run {i+1}: ALL FAILED err={s['err_samples']}", flush=True)
if i == 0 and rows and rows[0].get("headers"): print(f" provenance headers: {rows[0]['headers']}", flush=True)
await asyncio.sleep(5)
def med(field):
vs = [r[field] for r in runs if r.get(field) is not None]
return statistics.median(vs) if vs else None
med_summary = {k: med(k) for k in ("agg_out_tok_s", "per_stream_tok_s_median", "ttft_ms_median", "ttft_ms_p95",
"latency_ms_median", "latency_ms_p95", "error_rate")}
out = {"label": a.label, "model": a.model, "base_url": a.base_url, "provider": a.provider,
"config": {"concurrency": a.concurrency, "requests": a.requests, "max_tokens": a.max_tokens,
"runs": a.runs, "prompts": len(prompts), "extra": extra},
"runs": runs, "median_of_runs": med_summary, "ts": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())}
json.dump(out, open(a.out, "w"), indent=1)
m = med_summary
print(f" MEDIAN OF {a.runs} RUNS: agg {m['agg_out_tok_s']:.0f} tok/s | per-stream {m['per_stream_tok_s_median']:.1f} tok/s | "
f"TTFT p50 {m['ttft_ms_median']:.0f}ms p95 {m['ttft_ms_p95']:.0f}ms | lat p50 {m['latency_ms_median']/1000:.2f}s "
f"p95 {m['latency_ms_p95']/1000:.2f}s | err {m['error_rate']*100:.1f}%" if m["agg_out_tok_s"] else " MEDIAN: no successful runs", flush=True)
print(f" raw -> {a.out}", flush=True)
asyncio.run(main())
Caveats
- This is our run, from one origin, on one day. Marketplace routing changes hour to hour; so does load on our hardware. Treat the numbers as a snapshot and the script as the durable part.
- Marketplaces were tested with default routing. Pinning a specific upstream on a marketplace may change its numbers in either direction.
- Throughput is aggregate across 64 streams, not per-stream decode speed. Per-stream speed is roughly throughput รท concurrency for a saturated pool.
- Geography matters: an origin near a provider's hardware flatters its TTFT. The vantage was AWS us-east-2 (Ohio) and our hardware is in the Boston area, so the route to us is a real internet path and not a short one. The TTFT decomposition above shows the internet leg is not what dominates our number, but a marketplace whose chosen upstream happens to sit in us-east-2 gets a genuine advantage in that column.