# earthruntime > Qwen and GPT-OSS inference without marketplace overhead. We operate the hardware ourselves: direct pricing, predictable performance, and full visibility into the model you run. A product of Provocative Science Holdings, Inc. ## Models - `qwen3.6-35b` - Qwen 3.6 35B, served from Qwen's official FP8 checkpoint, 262K context. Reasoning off by default; toggle with chat_template_kwargs.enable_thinking. - `qwen3.8-27b` - Qwen 3.8 27B, Qwen official FP8 with FP8 KV, 262K context. Reasoning off by default; same toggle. - `gpt-oss-120b` - GPT-OSS 120B (MoE), checkpoint-native MXFP4 with FP8 KV, 131K context. Reasoning ON by default; control with reasoning_effort (low/medium/high). - Coming soon: DeepSeek V4 Flash (preview access not yet open). ## Precision Models are served from the model authors' own official checkpoints. We do not requantize anything ourselves. qwen3.6-35b and qwen3.8-27b run Qwen's official FP8 releases; gpt-oss-120b runs its checkpoint-native MXFP4. ## API - OpenAI-compatible endpoint: `POST https://api.earthruntime.com/v1/chat/completions` - Auth: `Authorization: Bearer $EARTHRUNTIME_KEY` - One key, one endpoint — change the model ID in the request to switch models. Example: ```bash curl https://api.earthruntime.com/v1/chat/completions \ -H "Authorization: Bearer $EARTHRUNTIME_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"qwen3.6-35b","messages":[{"role":"user","content":"Hello"}]}' ``` ## Pricing - Start with 100,000 free tokens. No card required. - Credit packs (no subscription): $1 (~1M tokens), $5 (~5M tokens), $20 (~20M tokens). - Full per-model input/output rates: https://earthruntime.com/pricing.md ## Performance Self-reported, run 2026-09-24 from AWS EC2 us-east-2c at 64 concurrent requests. Methodology and reproduction script: https://earthruntime.com/benchmarks - qwen3.6-35b: 2,464 tok/s aggregate, TTFT p50 765 ms at c=64 (73 ms at c=1), latency p50 2.47s / p95 3.56s, 0% errors - gpt-oss-120b: 2,276 tok/s, TTFT p50 537 ms, latency p50 3.11s / p95 4.57s, 0% errors - qwen3.8-27b: 1,608 tok/s, TTFT p50 508 ms, latency p50 4.70s / p95 8.59s, 0% errors - Zero errors across all 3,104 requests. - Versus OpenRouter default routing on qwen3.6-35b: we are +61% on aggregate throughput (2,464 vs 1,534) and ahead on p95 (3.56s vs 4.68s); OpenRouter is faster on the median (1.65s vs our 2.47s). Parasail posts 3,042 tok/s but drops 17.6% of requests with 429s. ## Infrastructure We operate inference hardware in cities and pair compute with direct-air carbon capture; recovered CO2 is supplied to nearby hospitality businesses. ## Docs - API reference and quickstart: https://earthruntime.com/docs - Benchmark methodology and reproduction script: https://earthruntime.com/benchmarks - Machine-readable pricing: https://earthruntime.com/pricing.md ## Contact - contact@earthruntime.com - https://earthruntime.com/