Chat completions
The single endpoint you call. It mirrors OpenAI's /v1/chat/completions, so existing SDKs and tools work without changes.
Authentication
Every request carries a bearer token in the Authorization header — either a generated client key or the master key:
Authorization: Bearer YOUR_API_KEY
Request body
| Field | Type | Notes |
|---|---|---|
messages | array | Required. Chat turns, each { "role", "content" }. role is system, user or assistant. |
model | string | Defaults to auto. Ignored if the key is pinned to a model. See Models & keys. |
stream | boolean | When true, the reply arrives as Server-Sent Events. Default false. |
user | string | Optional. Doubles as a sticky-routing key so one end-user's turns keep landing on the same upstream account. |
Other OpenAI fields (temperature, max_tokens, top_p, …) are accepted for compatibility but ignored — the upstream ChatGPT account governs generation.
Message content
content is either a plain string, or an array of parts for multimodal input:
{
"messages": [
{ "role": "system", "content": "You are concise." },
{ "role": "user", "content": "Summarise the plot of Dune in one line." }
]
}To attach images, use the content-parts form documented in Vision.
Non-streaming response
With stream omitted or false, you get one JSON object in OpenAI's shape:
{
"id": "chatcmpl-…",
"object": "chat.completion",
"created": 1733500000,
"model": "auto",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello! How can I help?" },
"finish_reason": "stop"
}
]
}Streaming response
With "stream": true, the body is a text/event-stream of chat.completion.chunk deltas, terminated by data: [DONE] — identical to OpenAI, so SDK stream helpers just work:
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"!"}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]Behind a reverse proxy, disable response buffering (in nginx / NPM: proxy_buffering off) or chunks arrive in one lump. The bundled deployment already sets this.
List models
Returns msemax for ChatGPT routing and opencode/muse-spark-1.3-contributor-free for the local OpenCode bridge. A key's model pin takes priority — see Models & keys.
Errors
Errors use standard HTTP status codes with a JSON { "detail": … } body.
| Status | Meaning |
|---|---|
400 | Malformed JSON, or a broken image (bad base64 / unfetchable URL). |
401 | Missing or invalid API key. |
413 | Too many messages, prompt too large, body too large, or an image over the size cap. |
429 | Per-IP rate limit hit. Back off and retry. |
502 | Upstream ChatGPT error the gateway could not recover from. |
503 | No healthy session in the pool. Add or revive an account. |