earthruntime API
One base URL, one bearer key, the OpenAI chat-completions shape. If your code already talks to api.openai.com, change two strings and it talks to us.
https://api.earthruntime.com/v1 · Auth Authorization: Bearer $EARTHRUNTIME_KEY · Models qwen3.6-35b qwen3.8-27b gpt-oss-120bQuickstart
- Get a key. Enter your email on the home page. A 6-digit code is emailed; paste it and the key is shown once, with 100,000 free tokens on it. Store it as
EARTHRUNTIME_KEY. - Make a request.
curl https://api.earthruntime.com/v1/chat/completions \ -H "Authorization: Bearer $EARTHRUNTIME_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "qwen3.6-35b", "messages": [{"role": "user", "content": "Explain why tail latency matters for production inference in three sentences."}] }'
Single-line, for terminals that eat backslashes:
curl https://api.earthruntime.com/v1/chat/completions -H "Authorization: Bearer $EARTHRUNTIME_KEY" -H "Content-Type: application/json" -d '{"model":"qwen3.6-35b","messages":[{"role":"user","content":"Hello"}]}'
Python (pip install openai):
from openai import OpenAI import os client = OpenAI(base_url="https://api.earthruntime.com/v1", api_key=os.environ["EARTHRUNTIME_KEY"]) r = client.chat.completions.create( model="qwen3.6-35b", messages=[{"role": "user", "content": "Hello"}], ) print(r.choices[0].message.content)
Node (npm i openai):
import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.earthruntime.com/v1", apiKey: process.env.EARTHRUNTIME_KEY }); const r = await client.chat.completions.create({ model: "gpt-oss-120b", messages: [{ role: "user", content: "Hello" }] }); console.log(r.choices[0].message.content);
Authentication
Every request carries the key as a bearer token:
Authorization: Bearer $EARTHRUNTIME_KEY
- Keys are shown once, at issue time, and never emailed. If you lose one, email contact@earthruntime.com and we rotate it.
- One key per email address. Credits bought with that address land on that key.
- Treat it like any secret: environment variables, not source control. The key is the only credential — there is no account password.
OpenAI compatibility
The API implements the OpenAI chat-completions request and response format. The official openai SDKs for Python and Node work unchanged with base_url / baseURL set; so do LangChain, LiteLLM, Vercel AI SDK, Continue, Cline, Open WebUI and anything else with an "OpenAI-compatible" provider setting.
| Endpoint | Status | Notes |
|---|---|---|
POST /v1/chat/completions | Supported | Streaming and non-streaming. Reference. |
GET /v1/models | Supported | Lists the served model ids. |
POST /v1/completions (legacy) | Supported | Returns text_completion. Prefer chat completions for new code. |
POST /v1/embeddings | Not offered | No embedding model is in the catalog; the route returns 404. |
| Tool / function calling | Supported | tools and tool_choice, on every served model. |
| JSON mode / structured output | Supported | response_format with a json_schema. |
| Vision / image input | Not supported | Not a feature we offer or test. qwen3.6-35b will in practice accept and describe an image, but that behaviour is undocumented, unsupported and may change without notice. Do not build on it. On the other models an image request may fail or be routed upstream, so send text. |
POST /v1/chat/completions
Request body, JSON. Fields follow OpenAI's semantics; the ones below are the ones people actually use.
| Field | Type | Notes |
|---|---|---|
model | string, required | qwen3.6-35b, qwen3.8-27b or gpt-oss-120b. An unknown id returns 404. |
messages | array, required | {role, content} objects. Roles: system, user, assistant. |
stream | boolean | Default false. See streaming. |
max_tokens | integer | Cap on completion tokens. No default: omit it and generation runs until the model stops on its own or fills the context window. Set it if you care about cost or tail latency. Prompt + completion must fit the context window. |
temperature, top_p | number | Standard sampling controls. Model defaults apply when omitted. |
stop | string | string[] | Up to 4 stop sequences. |
frequency_penalty, presence_penalty | number | Standard. |
seed | integer | Honoured. The same seed and body returns an identical completion; a different seed returns a different one. |
stream_options.include_usage | boolean | Honoured. The final chunk carries usage, including reasoning_tokens. |
Response — the standard shape:
{
"id": "chatcmpl-…",
"object": "chat.completion",
"model": "qwen3.6-35b",
"choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
"usage": { "prompt_tokens": 21, "completion_tokens": 84, "total_tokens": 105 }
}
usage is what you are billed on — see billing.
Streaming
Set "stream": true. The response is text/event-stream: one data: line per chunk, each a chat.completion.chunk object with a delta, terminated by data: [DONE]. This is the format every OpenAI SDK already parses.
curl -N https://api.earthruntime.com/v1/chat/completions \ -H "Authorization: Bearer $EARTHRUNTIME_KEY" -H "Content-Type: application/json" \ -d '{"model":"qwen3.6-35b","stream":true,"messages":[{"role":"user","content":"Count to five."}]}' data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]} data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"One"}}]} … data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]} data: [DONE]
Time to first token is 73 ms at a single stream and 765 ms at 64 concurrent streams; see where that gap comes from. Stream if your UI shows tokens as they arrive.
Reasoning toggle
Qwen 3.6 35B is served with reasoning off by default — it costs latency and output tokens, and most production calls don't need it. Turn it on per request when the task benefits (multi-step math, planning, hard code changes):
{
"model": "qwen3.6-35b",
"messages": [ … ],
"chat_template_kwargs": { "enable_thinking": true } // Qwen models
}
Reasoning is returned in reasoning_content on the delta or message, separately from content. It is counted inside completion_tokens and billed at the output rate.
The toggle differs by model family. There is no single parameter that works everywhere:
| Model | Default | How to change it |
|---|---|---|
qwen3.6-35b | off | chat_template_kwargs: {"enable_thinking": true} |
qwen3.8-27b | off | chat_template_kwargs: {"enable_thinking": true} |
gpt-oss-120b | on | reasoning_effort: low, medium or high. Any other value returns 400. There is no way to turn it off. |
Note that gpt-oss-120b reasons by default, so it bills more output tokens than you may expect. Effort scales roughly 32 / 65 / 156 completion tokens on a short prompt. | ||
GET /v1/models
Implemented. Returns the model ids you can pass as model. The list, with what each is served at:
| Model id | Model | Context | Precision | Reasoning |
|---|---|---|---|---|
qwen3.6-35b | Qwen 3.6 35B | 262,144 | Qwen official FP8 | Off by default; per-request toggle |
qwen3.8-27b | Qwen 3.8 27B | 262,144 | Qwen official FP8, FP8 KV cache | Off by default; per-request toggle |
gpt-oss-120b | GPT-OSS 120B (MoE) | 131,072 | Checkpoint-native MXFP4, FP8 KV cache | On by default; reasoning_effort |
| DeepSeek V4 Flash: coming soon; not yet listed. Per-token prices: pricing · pricing.md. | ||||
Errors
Errors use the OpenAI envelope — { "error": { "message", "type", "code" } } — with the HTTP status doing the real work:
| Status | Meaning | What to do |
|---|---|---|
400 | Malformed request, or prompt + max_tokens exceeds the context window. | Fix the body. The message says which field. |
401 | Missing or invalid key. | Check the Authorization header. Keys are case-sensitive. |
402 | Balance is zero. | Add credit with the email the key is tied to; requests succeed within a minute of payment. The body is {"error":{"message":"Insufficient balance: this key's prepaid credit is used up.","type":"auth_error","code":"402"}}. No Retry-After, since it is not a rate limit. |
404 | Unknown model id. | Use one of the ids above, exactly. |
429 | Rate limited. | Back off and retry; honour Retry-After when present. See limits. |
5xx | Our fault. | Retry with backoff. Persistent 5xx: check status and tell us. |
Rate limits
There is no published per-key request or token rate limit today. Your prepaid balance is the effective ceiling, and we do not emit 429.
64 concurrent requests on a single key is verified: our published benchmark ran 3,104 requests that way and returned zero errors. Treat anything above that as untested rather than blessed, and tell us if you need it so we can make room. We expect to introduce limits as the fleet fills, and will announce them before they take effect.
Credits & billing
- Free tier: every new key starts with a prepaid budget of $0.097, no card. That is a dollar budget, not a token counter: it is sized to deliver the 100,000 tokens we promise, and the odd figure is what 100,000 output tokens cost at the rate the promise was written against. Credit is spent per token at each model's own rate, so what you actually get depends on the model and your input/output mix. On
qwen3.6-35bit is about 108,000 output tokens, and ongpt-oss-120broughly 570,000. 100,000 is the floor across the fleet, never the ceiling. - Metering: billed per token on the
usageblock of each response —prompt_tokensat the model's input rate,completion_tokens(including any reasoning tokens) at its output rate. Rates: pricing. - Adding credit: pick a pack on the pricing section ($1 / $5 / $20), enter the email your key is tied to, pay via Stripe. Credits land on the key within a minute. No subscription, and credits do not expire.
- Zero balance: requests return
402until credit is added. Nothing is queued or throttled; the key stops working until it is topped up. - Invoices / higher volume: contact@earthruntime.com. Charges appear from Provocative Science Holdings, Inc.
Status & support
- GitHub org: github.com/provocative-science. The gateway and the benchmark harness are not public yet; ask us and we will send you the harness.
- Email: contact@earthruntime.com — a human reads it.
- Machine-readable summary of this site: /llms.txt.