API reference

Chat completions

The single endpoint you call. It mirrors OpenAI's /v1/chat/completions, so existing SDKs and tools work without changes.

POST __BASE__/v1/chat/completions

Authentication

Every request carries a bearer token in the Authorization header — either a generated client key or the master key:

Authorization: Bearer YOUR_API_KEY

Request body

FieldTypeNotes
messagesarrayRequired. Chat turns, each { "role", "content" }. role is system, user or assistant.
modelstringDefaults to auto. Ignored if the key is pinned to a model. See Models & keys.
streambooleanWhen true, the reply arrives as Server-Sent Events. Default false.
userstringOptional. Doubles as a sticky-routing key so one end-user's turns keep landing on the same upstream account.

Other OpenAI fields (temperature, max_tokens, top_p, …) are accepted for compatibility but ignored — the upstream ChatGPT account governs generation.

Message content

content is either a plain string, or an array of parts for multimodal input:

{
  "messages": [
    { "role": "system", "content": "You are concise." },
    { "role": "user", "content": "Summarise the plot of Dune in one line." }
  ]
}

To attach images, use the content-parts form documented in Vision.

Non-streaming response

With stream omitted or false, you get one JSON object in OpenAI's shape:

{
  "id": "chatcmpl-…",
  "object": "chat.completion",
  "created": 1733500000,
  "model": "auto",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hello! How can I help?" },
      "finish_reason": "stop"
    }
  ]
}

Streaming response

With "stream": true, the body is a text/event-stream of chat.completion.chunk deltas, terminated by data: [DONE] — identical to OpenAI, so SDK stream helpers just work:

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"!"}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Behind a reverse proxy, disable response buffering (in nginx / NPM: proxy_buffering off) or chunks arrive in one lump. The bundled deployment already sets this.

List models

GET __BASE__/v1/models

Returns msemax for ChatGPT routing and opencode/muse-spark-1.3-contributor-free for the local OpenCode bridge. A key's model pin takes priority — see Models & keys.

Errors

Errors use standard HTTP status codes with a JSON { "detail": … } body.

StatusMeaning
400Malformed JSON, or a broken image (bad base64 / unfetchable URL).
401Missing or invalid API key.
413Too many messages, prompt too large, body too large, or an image over the size cap.
429Per-IP rate limit hit. Back off and retry.
502Upstream ChatGPT error the gateway could not recover from.
503No healthy session in the pool. Add or revive an account.