Docs
Completions
POST
https://api.modelsheep.com/completionsPOST
https://api.modelsheep.com/v1/completionsCreates a text completion for a prompt. The endpoint is compatible with the OpenAI completions endpoint and takes a plain prompt instead of chat messages.
Both paths are served and return the same response. The OpenAI SDKs append /v1 on their own.
Request Body
- promptstringrequired
- The text the model continues.
- modelstringrequired
- A public model name, as listed on /models.
- max_tokensinteger1 to the context window
- The maximum number of tokens in the response. Thinking text counts against it.
- temperaturenumber0 to 2
- Sampling temperature. Higher values give more varied output. It defaults to the value the model ships with, which differs per model.
- top_pnumber0 to 1
- Nucleus sampling, an alternative to temperature.
- stopstring or array
- One or more strings that end the response.
- streambooleandefault false
- When true the answer arrives as server-sent events, a piece at a time, instead of one object at the end.
Response (stream = false)
- idstring
- The id of this completion.
- objectstring
- The type of object returned.
- createdinteger
- The creation time in unix seconds.
- modelstring
- The model that produced the response.
- array
- The generated completions, one entry unless more were requested.
- object
- The token counts the call is priced on.
Response (stream = true)
Set stream to true and the answer arrives as server-sent events. Each event is a line that starts with data: and carries one chunk of JSON. This endpoint puts the piece in choices[0].text rather than in a delta. The stream ends with the literal data: [DONE], which is not JSON.
- idstring
- The same id on every chunk of one answer.
- objectstring
- text_completion.
- createdinteger
- Unix seconds, the same on every chunk.
- modelstring
- The model that answered.
- array
- One entry, carrying the piece written since the last chunk.
Request Example
from openai import OpenAI
client = OpenAI(
base_url="https://api.modelsheep.com",
api_key=API_KEY_MODELSHEEP,
)
response = client.completions.create(
model="Qwen/Qwen3.5-35B-A3B-FP8",
prompt="A sheep walks into a data centre and",
max_tokens=60,
)
print(response.choices[0].text)Response (stream = false)
One JSON object, sent once the model has finished.
{
"id": "cmpl-...",
"object": "text_completion",
"created": 1789832000,
"model": "Qwen/Qwen3.5-35B-A3B-FP8",
"choices": [
{
"index": 0,
"text": " asks the rack how it sleeps at night.",
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 9,
"completion_tokens": 11,
"total_tokens": 20
}
}Response (stream = true)
Server-sent events. Each event is a line that starts with data: and carries one piece of the answer. The last line is data: [DONE].
data: {"id":"cmpl-abc","object":"text_completion","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"text":"Counting","finish_reason":null}]}
data: {"id":"cmpl-abc","object":"text_completion","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"text":" is a rhythm","finish_reason":null}]}
data: {"id":"cmpl-abc","object":"text_completion","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"text":"","finish_reason":"stop"}]}
data: [DONE]