docs

earthruntime API

One base URL, one bearer key, the OpenAI chat-completions shape. If your code already talks to api.openai.com, change two strings and it talks to us.

Base URL https://api.earthruntime.com/v1 · Auth Authorization: Bearer $EARTHRUNTIME_KEY · Models qwen3.6-35b qwen3.8-27b gpt-oss-120b

Quickstart

  1. Get a key. Enter your email on the home page. A 6-digit code is emailed; paste it and the key is shown once, with 100,000 free tokens on it. Store it as EARTHRUNTIME_KEY.
  2. Make a request.
curl https://api.earthruntime.com/v1/chat/completions \
  -H "Authorization: Bearer $EARTHRUNTIME_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.6-35b",
    "messages": [{"role": "user", "content": "Explain why tail latency matters for production inference in three sentences."}]
  }'

Single-line, for terminals that eat backslashes:

curl https://api.earthruntime.com/v1/chat/completions -H "Authorization: Bearer $EARTHRUNTIME_KEY" -H "Content-Type: application/json" -d '{"model":"qwen3.6-35b","messages":[{"role":"user","content":"Hello"}]}'

Python (pip install openai):

from openai import OpenAI
import os

client = OpenAI(base_url="https://api.earthruntime.com/v1", api_key=os.environ["EARTHRUNTIME_KEY"])
r = client.chat.completions.create(
    model="qwen3.6-35b",
    messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)

Node (npm i openai):

import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.earthruntime.com/v1", apiKey: process.env.EARTHRUNTIME_KEY });
const r = await client.chat.completions.create({ model: "gpt-oss-120b", messages: [{ role: "user", content: "Hello" }] });
console.log(r.choices[0].message.content);

Authentication

Every request carries the key as a bearer token:

Authorization: Bearer $EARTHRUNTIME_KEY

OpenAI compatibility

The API implements the OpenAI chat-completions request and response format. The official openai SDKs for Python and Node work unchanged with base_url / baseURL set; so do LangChain, LiteLLM, Vercel AI SDK, Continue, Cline, Open WebUI and anything else with an "OpenAI-compatible" provider setting.

What is and isn't supported.
EndpointStatusNotes
POST /v1/chat/completionsSupportedStreaming and non-streaming. Reference.
GET /v1/modelsSupportedLists the served model ids.
POST /v1/completions (legacy)SupportedReturns text_completion. Prefer chat completions for new code.
POST /v1/embeddingsNot offeredNo embedding model is in the catalog; the route returns 404.
Tool / function callingSupportedtools and tool_choice, on every served model.
JSON mode / structured outputSupportedresponse_format with a json_schema.
Vision / image inputNot supportedNot a feature we offer or test. qwen3.6-35b will in practice accept and describe an image, but that behaviour is undocumented, unsupported and may change without notice. Do not build on it. On the other models an image request may fail or be routed upstream, so send text.

POST /v1/chat/completions

Request body, JSON. Fields follow OpenAI's semantics; the ones below are the ones people actually use.

FieldTypeNotes
modelstring, requiredqwen3.6-35b, qwen3.8-27b or gpt-oss-120b. An unknown id returns 404.
messagesarray, required{role, content} objects. Roles: system, user, assistant.
streambooleanDefault false. See streaming.
max_tokensintegerCap on completion tokens. No default: omit it and generation runs until the model stops on its own or fills the context window. Set it if you care about cost or tail latency. Prompt + completion must fit the context window.
temperature, top_pnumberStandard sampling controls. Model defaults apply when omitted.
stopstring | string[]Up to 4 stop sequences.
frequency_penalty, presence_penaltynumberStandard.
seedintegerHonoured. The same seed and body returns an identical completion; a different seed returns a different one.
stream_options.include_usagebooleanHonoured. The final chunk carries usage, including reasoning_tokens.

Response — the standard shape:

{
  "id": "chatcmpl-…",
  "object": "chat.completion",
  "model": "qwen3.6-35b",
  "choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
  "usage": { "prompt_tokens": 21, "completion_tokens": 84, "total_tokens": 105 }
}

usage is what you are billed on — see billing.

Streaming

Set "stream": true. The response is text/event-stream: one data: line per chunk, each a chat.completion.chunk object with a delta, terminated by data: [DONE]. This is the format every OpenAI SDK already parses.

curl -N https://api.earthruntime.com/v1/chat/completions \
  -H "Authorization: Bearer $EARTHRUNTIME_KEY" -H "Content-Type: application/json" \
  -d '{"model":"qwen3.6-35b","stream":true,"messages":[{"role":"user","content":"Count to five."}]}'

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"One"}}]}
…
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]

Time to first token is 73 ms at a single stream and 765 ms at 64 concurrent streams; see where that gap comes from. Stream if your UI shows tokens as they arrive.

Reasoning toggle

Qwen 3.6 35B is served with reasoning off by default — it costs latency and output tokens, and most production calls don't need it. Turn it on per request when the task benefits (multi-step math, planning, hard code changes):

{
  "model": "qwen3.6-35b",
  "messages": [ … ],
  "chat_template_kwargs": { "enable_thinking": true }   // Qwen models
}

Reasoning is returned in reasoning_content on the delta or message, separately from content. It is counted inside completion_tokens and billed at the output rate.

The toggle differs by model family. There is no single parameter that works everywhere:

ModelDefaultHow to change it
qwen3.6-35boffchat_template_kwargs: {"enable_thinking": true}
qwen3.8-27boffchat_template_kwargs: {"enable_thinking": true}
gpt-oss-120bonreasoning_effort: low, medium or high. Any other value returns 400. There is no way to turn it off.
Note that gpt-oss-120b reasons by default, so it bills more output tokens than you may expect. Effort scales roughly 32 / 65 / 156 completion tokens on a short prompt.

GET /v1/models

Implemented. Returns the model ids you can pass as model. The list, with what each is served at:

Model idModelContextPrecisionReasoning
qwen3.6-35bQwen 3.6 35B262,144Qwen official FP8Off by default; per-request toggle
qwen3.8-27bQwen 3.8 27B262,144Qwen official FP8, FP8 KV cacheOff by default; per-request toggle
gpt-oss-120bGPT-OSS 120B (MoE)131,072Checkpoint-native MXFP4, FP8 KV cacheOn by default; reasoning_effort
DeepSeek V4 Flash: coming soon; not yet listed. Per-token prices: pricing · pricing.md.

Errors

Errors use the OpenAI envelope — { "error": { "message", "type", "code" } } — with the HTTP status doing the real work:

StatusMeaningWhat to do
400Malformed request, or prompt + max_tokens exceeds the context window.Fix the body. The message says which field.
401Missing or invalid key.Check the Authorization header. Keys are case-sensitive.
402Balance is zero.Add credit with the email the key is tied to; requests succeed within a minute of payment. The body is {"error":{"message":"Insufficient balance: this key's prepaid credit is used up.","type":"auth_error","code":"402"}}. No Retry-After, since it is not a rate limit.
404Unknown model id.Use one of the ids above, exactly.
429Rate limited.Back off and retry; honour Retry-After when present. See limits.
5xxOur fault.Retry with backoff. Persistent 5xx: check status and tell us.

Rate limits

There is no published per-key request or token rate limit today. Your prepaid balance is the effective ceiling, and we do not emit 429.

64 concurrent requests on a single key is verified: our published benchmark ran 3,104 requests that way and returned zero errors. Treat anything above that as untested rather than blessed, and tell us if you need it so we can make room. We expect to introduce limits as the fleet fills, and will announce them before they take effect.

Credits & billing

Status & support