Docs
Chat completions
POST
https://api.modelsheep.com/chat/completionsPOST
https://api.modelsheep.com/v1/chat/completionsCreates a model response for a list of chat messages. The endpoint is compatible with the OpenAI chat completions endpoint, so the official SDKs call it without changes.
Both paths are served and return the same response. The OpenAI SDKs append /v1 on their own.
Request Body
- arrayrequired
- The conversation so far, oldest first.
- modelstringrequired
- A public model name, as listed on /models.
- max_tokensinteger1 to the context window
- The maximum number of tokens in the response. Thinking text counts against it.
- temperaturenumber0 to 2
- Sampling temperature. Higher values give more varied output. It defaults to the value the model ships with, which differs per model.
- top_pnumber0 to 1
- Nucleus sampling, an alternative to temperature.
- stopstring or array
- One or more strings that end the response.
- streambooleandefault false
- When true the answer arrives as server-sent events, a piece at a time, instead of one object at the end.
- array
- Functions the model may call. The response then carries tool_calls instead of content.
- object
- The shape the response must take.
- object
- Options the model reads when it builds the prompt from your messages. The keys differ per model, and a model ignores a key it does not know. The page of each model lists the keys it reads.
- reasoning_effortstringnone to max
- How long the model thinks. none turns the thinking off. The levels differ per model, and the page of each model lists the ones it takes.
- thinking_token_budgetinteger
- The largest number of tokens a model may spend on the thinking step. It caps the thinking alone, not the answer after it.
- include_reasoningbooleandefault true
- Send false to keep the thinking text out of the response. The model still thinks, and those tokens are still billed.
- continue_final_messagebooleandefault false
- Continues the final message instead of starting a new one. Use it when finish_reason came back as length: send the text you already have as a last message with the role assistant, and the model carries on from where it stopped. The response repeats that text before the new text, so cut the part you sent.
- add_generation_promptbooleandefault true
- Opens a new assistant turn at the end of the prompt. It must be false when continue_final_message is true.
Response (stream = false)
- idstring
- The id of this completion.
- objectstring
- The type of object returned.
- createdinteger
- The creation time in unix seconds.
- modelstring
- The model that produced the response.
- array
- The generated responses, one entry unless more were requested.
- object
- The token counts the call is priced on.
Response (stream = true)
Set stream to true and the answer arrives as server-sent events. Each event is a line that starts with data: and carries one chunk of JSON. Append every delta.content to what you already have. The stream ends with the literal data: [DONE], which is not JSON and must not be parsed as it.
- idstring
- The same id on every chunk of one answer.
- objectstring
- chat.completion.chunk, which is how a chunk is told apart from a finished answer.
- createdinteger
- Unix seconds, the same on every chunk.
- modelstring
- The model that answered.
- array
- One entry, carrying the piece written since the last chunk.
Request Example
from openai import OpenAI
client = OpenAI(
base_url="https://api.modelsheep.com",
api_key=API_KEY_MODELSHEEP,
)
response = client.chat.completions.create(
model="Qwen/Qwen3.5-35B-A3B-FP8",
messages=[
{
"role": "user",
"content": "Why do sheep count people?"
}
],
max_tokens=200,
)
print(response.choices[0].message.content)Response (stream = false)
One JSON object, sent once the model has finished.
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1789832000,
"model": "Qwen/Qwen3.5-35B-A3B-FP8",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Counting is a rhythm without a destination."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 14,
"completion_tokens": 41,
"total_tokens": 55
}
}Response (stream = true)
Server-sent events. Each event is a line that starts with data: and carries one piece of the answer. The last line is data: [DONE].
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"delta":{"content":"Counting"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"delta":{"content":" is a rhythm"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc","object":"chat.completion.chunk","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]