ModelSheep

Docs

Completions

POSThttps://api.modelsheep.com/completions
POSThttps://api.modelsheep.com/v1/completions

Creates a text completion for a prompt. The endpoint is compatible with the OpenAI completions endpoint and takes a plain prompt instead of chat messages.

Both paths are served and return the same response. The OpenAI SDKs append /v1 on their own.

Request Body

promptstringrequired
The text the model continues.
modelstringrequired
A public model name, as listed on /models.
max_tokensinteger1 to the context window
The maximum number of tokens in the response. Thinking text counts against it.
temperaturenumber0 to 2
Sampling temperature. Higher values give more varied output. It defaults to the value the model ships with, which differs per model.
top_pnumber0 to 1
Nucleus sampling, an alternative to temperature.
stopstring or array
One or more strings that end the response.
streambooleandefault false
When true the answer arrives as server-sent events, a piece at a time, instead of one object at the end.

Response (stream = false)

idstring
The id of this completion.
objectstring
The type of object returned.
createdinteger
The creation time in unix seconds.
modelstring
The model that produced the response.
array
The generated completions, one entry unless more were requested.
object
The token counts the call is priced on.

Response (stream = true)

Set stream to true and the answer arrives as server-sent events. Each event is a line that starts with data: and carries one chunk of JSON. This endpoint puts the piece in choices[0].text rather than in a delta. The stream ends with the literal data: [DONE], which is not JSON.

idstring
The same id on every chunk of one answer.
objectstring
text_completion.
createdinteger
Unix seconds, the same on every chunk.
modelstring
The model that answered.
array
One entry, carrying the piece written since the last chunk.

Request Example

from openai import OpenAI

client = OpenAI(
    base_url="https://api.modelsheep.com",
    api_key=API_KEY_MODELSHEEP,
)

response = client.completions.create(
    model="Qwen/Qwen3.5-35B-A3B-FP8",
    prompt="A sheep walks into a data centre and",
    max_tokens=60,
)

print(response.choices[0].text)

Response (stream = false)

One JSON object, sent once the model has finished.

{
  "id": "cmpl-...",
  "object": "text_completion",
  "created": 1789832000,
  "model": "Qwen/Qwen3.5-35B-A3B-FP8",
  "choices": [
    {
      "index": 0,
      "text": " asks the rack how it sleeps at night.",
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 9,
    "completion_tokens": 11,
    "total_tokens": 20
  }
}

Response (stream = true)

Server-sent events. Each event is a line that starts with data: and carries one piece of the answer. The last line is data: [DONE].

data: {"id":"cmpl-abc","object":"text_completion","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"text":"Counting","finish_reason":null}]}

data: {"id":"cmpl-abc","object":"text_completion","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"text":" is a rhythm","finish_reason":null}]}

data: {"id":"cmpl-abc","object":"text_completion","created":1789832000,"model":"Qwen/Qwen3.5-35B-A3B-FP8","choices":[{"index":0,"text":"","finish_reason":"stop"}]}

data: [DONE]