benchmarks

Self-reported. Reproducible.

Every number on the home page comes from one run of the script at the bottom of this page, pointed at four providers from the same machine, in the same hour, with the same prompts. Nothing is averaged across runs or cherry-picked from several. If you get different numbers, we want to hear about it: contact@earthruntime.com.

Run: 2026-09-24, 05:58 to 06:14 UTC · Vantage: AWS EC2 t3.xlarge, us-east-2c, outside our network · Concurrency: 64 · Requests: 256 per run, 3 runs per endpoint, median reported · max_tokens: 128 · Source of record: provocapier/docs/bench-20260923.md at commit f4d6037

Results

Our fleet

Every public model, 64 concurrent requests, median of 3 runs.
ModelAggregatePer streamTTFT p50TTFT p95Latency p50Latency p95Errors
qwen3.6-35b2,464 tok/s78.0 tok/s765 ms1,481 ms2.47 s3.56 s0%
gpt-oss-120b2,276 tok/s52.8 tok/s537 ms1,224 ms3.11 s4.57 s0%
qwen3.8-27b1,608 tok/s30.1 tok/s508 ms1,095 ms4.70 s8.59 s0%
Reasoning was off for both Qwen models (the server default) and on at the provider default effort for gpt-oss-120b. Zero errors across all 3,104 requests.

Every model in our fleet has an OpenRouter fallback configured, so a flapping worker could silently turn one of our rows into an OpenRouter row. Each model's runs were bracketed by a read of its own workers' request counters to prove that did not happen: 796, 794, 784 and 784 requests landed on our own hardware against ~776 expected, the small excess being health-check traffic. All four passed.

Against the marketplaces

Same run, same vantage, same 32-prompt workload, on qwen3.6-35b and its OpenRouter equivalent qwen/qwen3.6-35b-a3b.
EndpointAggregateTTFT p50Latency p50Latency p95ErrorsDominant error
earthruntime2,464 tok/s765 ms2.47 s3.56 s0%none
OpenRouter, default routing1,534 tok/s541 ms1.65 s4.68 s0%none
OpenRouter to AtlasCloud1,327 tok/s887 ms1.87 s3.18 s0%none
OpenRouter to Parasail3,042 tok/s560 ms2.18 s2.74 s17.6%HTTP 429, 100% of its failures
Parasail's throughput is the highest number in this table and it is not a win: aggregate throughput is successful tokens divided by wall clock, and an endpoint that refuses 17.6% of the load finishes a smaller job sooner. The same 429 pattern showed up in our 2026-06 and 2026-07 runs.

On the other three models, against OpenRouter's default routing: we are ahead on gpt-oss-120b (2,276 vs 1,438 tok/s) and behind on qwen3.8-27b (1,608 vs 1,786 tok/s).

Where we do not lead

OpenRouter's default routing beats us on median latency, 1.65 s against our 2.47 s. We are ahead on the p95 tail (3.56 s against 4.68 s), on aggregate throughput, and on not dropping requests. If your workload is one request at a time and you care about the average case, that gap is real and it is in their favour. If it is many requests at once and you care about the worst case, it is in ours.

OpenRouter is also ahead on qwen3.8-27b throughput. We publish it in the table above rather than omitting the row.

Test environment

Why our TTFT is 765 ms here and 73 ms on your laptop

Time to first token in these tables is measured at 64 concurrent streams, which is the worst case in this document, not what a single caller experiences. At one stream it is 73 ms; at 16 it is 204 ms. Quote the number at the concurrency you care about.

We decomposed the 765 ms rather than leaving it as a mystery, using vLLM's own histograms on either side of a run:

Where the 765 ms goes, measured at c=64 on the control plane over loopback.
LayerTTFT
Worker side, from vLLM's histogram94.5 ms
of which queued0.0 ms
of which prefill71.2 ms
Client TTFB, 64 parallel curl, no Python419 ms
Client TTFT, our Python asyncio harness765 ms
Requests spend zero time queued and prefill is about 71 ms, so the models, the GPUs and the parallelism are not the cause. Roughly 325 ms is added by the gateway and a further 300 ms by the benchmark client itself at 64 streams. Loopback (768 ms) and the remote vantage (765 ms) agree, so the internet path is not implicated either.

An earlier version of this page published 196 ms. That figure was measured against a different stack: the FastAPI gateway that was retired on 2026-09-22. The two numbers do not describe the same system, so we replaced it rather than defending it.

What got worse

This run against the figures this page used to publish, qwen3.6-35b.
MetricPublished beforeThis runVerdict
Aggregate throughput2,237 tok/s2,464 tok/sbetter by 10%
Error rate0%0%unchanged
Latency p953.6 s3.56 sunchanged
Latency p502.2 s2.47 sworse by 12%
TTFT p50196 ms765 msworse by 3.9x, see above

Definitions

What counts as an error

Any request that did not return a complete, well-formed answer. Concretely, in is_error():

An errored request contributes no latency sample and no tokens. That is deliberate: a fast 429 would otherwise improve a provider's median.

Reproduce it

Get a key from the home page, then run the same harness that produced every number above. It needs the 32-prompt workload file and writes one JSON summary per endpoint:

# Python 3.12, httpx 0.26
pip install httpx
echo "$EARTHRUNTIME_KEY" > key.txt

python bench64.py --base-url https://api.earthruntime.com/v1 --key-file key.txt \
                  --model qwen3.6-35b --label ours_qwen36 \
                  --prompts bench_prompts.json --out ours_qwen36.json \
                  --concurrency 64 --requests 256 --max-tokens 128 --runs 3

# the same run against a marketplace
echo "$OPENROUTER_KEY" > or_key.txt
python bench64.py --base-url https://openrouter.ai/api/v1 --key-file or_key.txt \
                  --model "qwen/qwen3.6-35b-a3b" --label or_qwen36 \
                  --prompts bench_prompts.json --out or_qwen36.json \
                  --concurrency 64 --requests 256 --max-tokens 128 --runs 3

# pin a named provider instead of default routing
python bench64.py ... --provider AtlasCloud

A 256-request run at 128 max tokens costs a few cents on our side, so the free grant covers several. To match our Qwen rows exactly, add --extra '{"reasoning":{"enabled":false}}' on the OpenRouter runs: our Qwen models serve with reasoning off by default and the comparison is otherwise unfair to them.

The harness is not on GitHub yet. It is committed beside the report in a private repository and our public org does not exist at the time of writing. The full script is printed below, and the 32-prompt workload is a plain JSON array of strings. Email contact@earthruntime.com and we will send you both files directly.

The script

bench64.py, verbatim, exactly as run. Standard library plus httpx.

bench64.py
#!/usr/bin/env python3
"""c=64 streaming benchmark, matched to bench_results_overnight_20260605T031811Z.

Shape of record: concurrency 64, 256 requests, max_tokens 128, 32 unique prompts
(scripts/bench.py's 20 VARIED + 12 CODING), warmup first, N runs, median across runs.
Streams every response so TTFT is real. Output tokens = content chunks (same
convention as the published CSVs, where output_tokens == chunks).
"""
from __future__ import annotations
import argparse, asyncio, json, statistics, sys, time
import httpx

def pct(xs, p):
    if not xs: return None
    xs = sorted(xs); k = (len(xs) - 1) * p
    f = int(k); c = min(f + 1, len(xs) - 1)
    return xs[f] + (xs[c] - xs[f]) * (k - f)

async def one(client, url, headers, body, idx):
    t0 = time.perf_counter(); ttft = None; chunks = 0; usage = None; err = None; hdrs = None
    try:
        async with client.stream("POST", url, headers=headers, json=body) as r:
            hdrs = {k.lower(): v for k, v in r.headers.items() if k.lower().startswith(("x-", "server"))}
            if r.status_code != 200:
                err = f"http{r.status_code}:{(await r.aread())[:120].decode('utf-8','replace')}"
            else:
                async for line in r.aiter_lines():
                    if not line.startswith("data:"): continue
                    data = line[5:].strip()
                    if data == "[DONE]": break
                    try: obj = json.loads(data)
                    except Exception: continue
                    if obj.get("usage"): usage = obj["usage"]
                    for ch in obj.get("choices") or []:
                        d = ch.get("delta") or {}
                        piece = d.get("content") or d.get("reasoning_content") or d.get("reasoning") or ""
                        if piece:
                            if ttft is None: ttft = (time.perf_counter() - t0) * 1000
                            chunks += 1
    except Exception as e:
        err = f"{type(e).__name__}:{str(e)[:120]}"
    total = (time.perf_counter() - t0) * 1000
    out_tok = (usage or {}).get("completion_tokens") or chunks
    return {"idx": idx, "ttft_ms": ttft, "total_ms": total, "output_tokens": out_tok,
            "chunks": chunks, "err": err, "headers": hdrs if idx == 0 else None}

async def run_once(url, headers, model, prompts, n, conc, max_tokens, extra):
    sem = asyncio.Semaphore(conc)
    limits = httpx.Limits(max_connections=conc + 8, max_keepalive_connections=conc + 8)
    async with httpx.AsyncClient(timeout=httpx.Timeout(300.0, connect=30.0), limits=limits) as client:
        async def guarded(i):
            async with sem:
                body = {"model": model, "messages": [{"role": "user", "content": prompts[i % len(prompts)]}],
                        "max_tokens": max_tokens, "temperature": 0, "stream": True,
                        "stream_options": {"include_usage": True}}
                body.update(extra)
                return await one(client, url, headers, body, i)
        t0 = time.perf_counter()
        rows = await asyncio.gather(*[guarded(i) for i in range(n)])
        wall = time.perf_counter() - t0
    return rows, wall

def summarize(rows, wall):
    ok = [r for r in rows if not r["err"] and r["ttft_ms"] is not None]
    bad = [r for r in rows if r["err"] or r["ttft_ms"] is None]
    ttfts = [r["ttft_ms"] for r in ok]; totals = [r["total_ms"] for r in ok]
    per_stream = [r["output_tokens"] / max(1e-6, (r["total_ms"] - r["ttft_ms"]) / 1000) for r in ok if r["output_tokens"] and r["total_ms"] > r["ttft_ms"]]
    return {"n": len(rows), "ok": len(ok), "errors": len(bad), "error_rate": len(bad) / len(rows) if rows else None,
            "wall_s": wall, "out_tokens_total": sum(r["output_tokens"] for r in ok),
            "agg_out_tok_s": sum(r["output_tokens"] for r in ok) / wall if wall else None,
            "per_stream_tok_s_median": statistics.median(per_stream) if per_stream else None,
            "ttft_ms_median": statistics.median(ttfts) if ttfts else None, "ttft_ms_p95": pct(ttfts, .95),
            "latency_ms_median": statistics.median(totals) if totals else None, "latency_ms_p95": pct(totals, .95),
            "err_samples": [r["err"] for r in bad][:3]}

async def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--base-url", required=True); ap.add_argument("--key-file", required=True)
    ap.add_argument("--model", required=True); ap.add_argument("--label", required=True)
    ap.add_argument("--runs", type=int, default=3); ap.add_argument("--concurrency", type=int, default=64)
    ap.add_argument("--requests", type=int, default=256); ap.add_argument("--max-tokens", type=int, default=128)
    ap.add_argument("--provider", default=None, help="OpenRouter provider order, e.g. 'AtlasCloud'")
    ap.add_argument("--extra", default="{}", help="extra JSON merged into each body")
    ap.add_argument("--prompts", required=True); ap.add_argument("--out", required=True)
    a = ap.parse_args()
    key = open(a.key_file).read().strip()
    prompts = json.load(open(a.prompts))
    url = a.base_url.rstrip("/") + "/chat/completions" if a.base_url.rstrip("/").endswith("/v1") else a.base_url.rstrip("/") + "/v1/chat/completions"
    headers = {"Authorization": f"Bearer {key}", "Content-Type": "application/json"}
    extra = json.loads(a.extra)
    if a.provider: extra["provider"] = {"order": [a.provider], "allow_fallbacks": False}
    print(f"=== {a.label} | {a.model} | c={a.concurrency} n={a.requests} max_tokens={a.max_tokens} runs={a.runs} ===", flush=True)
    print(f"  warmup (c=4, n=8)...", flush=True)
    await run_once(url, headers, a.model, prompts, 8, 4, a.max_tokens, extra)
    runs = []
    for i in range(a.runs):
        rows, wall = await run_once(url, headers, a.model, prompts, a.requests, a.concurrency, a.max_tokens, extra)
        s = summarize(rows, wall); s["run"] = i + 1; runs.append(s)
        print(f"  run {i+1}: agg {s['agg_out_tok_s']:.0f} tok/s | TTFT p50 {s['ttft_ms_median']:.0f}ms p95 {s['ttft_ms_p95']:.0f}ms | "
              f"lat p50 {s['latency_ms_median']/1000:.2f}s p95 {s['latency_ms_p95']/1000:.2f}s | err {s['error_rate']*100:.1f}% | wall {wall:.1f}s"
              if s["ok"] else f"  run {i+1}: ALL FAILED err={s['err_samples']}", flush=True)
        if i == 0 and rows and rows[0].get("headers"): print(f"    provenance headers: {rows[0]['headers']}", flush=True)
        await asyncio.sleep(5)
    def med(field):
        vs = [r[field] for r in runs if r.get(field) is not None]
        return statistics.median(vs) if vs else None
    med_summary = {k: med(k) for k in ("agg_out_tok_s", "per_stream_tok_s_median", "ttft_ms_median", "ttft_ms_p95",
                                       "latency_ms_median", "latency_ms_p95", "error_rate")}
    out = {"label": a.label, "model": a.model, "base_url": a.base_url, "provider": a.provider,
           "config": {"concurrency": a.concurrency, "requests": a.requests, "max_tokens": a.max_tokens,
                      "runs": a.runs, "prompts": len(prompts), "extra": extra},
           "runs": runs, "median_of_runs": med_summary, "ts": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())}
    json.dump(out, open(a.out, "w"), indent=1)
    m = med_summary
    print(f"  MEDIAN OF {a.runs} RUNS: agg {m['agg_out_tok_s']:.0f} tok/s | per-stream {m['per_stream_tok_s_median']:.1f} tok/s | "
          f"TTFT p50 {m['ttft_ms_median']:.0f}ms p95 {m['ttft_ms_p95']:.0f}ms | lat p50 {m['latency_ms_median']/1000:.2f}s "
          f"p95 {m['latency_ms_p95']/1000:.2f}s | err {m['error_rate']*100:.1f}%" if m["agg_out_tok_s"] else "  MEDIAN: no successful runs", flush=True)
    print(f"  raw -> {a.out}", flush=True)

asyncio.run(main())

Caveats