Inference
Quick start
Get your first response in a few lines. The inference API is OpenAI-compatible, so you can point any OpenAI SDK at the base URL below and authenticate with your API key.
https://api.nebux.cloud/v1Installation
npm install openaiimport OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-api-key",
baseURL: "https://api.nebux.cloud/v1",
});
const completion = await client.chat.completions.create({
model: "qwen3-14b",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Recommend me a book" },
{ role: "assistant", content: "Try The Shadow of the Wind" },
{ role: "user", content: "And a similar one?" },
],
});
console.log(completion.choices[0].message.content);{
"id": "chatcmpl-8f3a2b1c9d",
"object": "chat.completion",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "If you liked it, try Nada, by Carmen Laforet."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 47,
"completion_tokens": 13,
"total_tokens": 60
}
}Authentication
All requests must include your API key in the Authorization header as a Bearer token.
Generate and manage your API keys from the dashboard. Manage API keys
Keep your API key secret. Never expose it in client-side code or commit it to version control.
Authorization: Bearer your-api-keyChat completions
Generates a model response for a conversation. Send a list of messages and receive the assistant's reply.
Parameters
modelstringrequiredID of the model to use.
messagesarrayrequiredThe conversation so far, as a list of messages.
rolestringThe role of the message author.
systemSets the assistant's behaviour and context.userA message from the end user.assistantA previous reply from the model.toolThe result of a tool call, sent back to the model.contentstringThe text content of the message.
namestringAn optional name to disambiguate participants with the same role.
temperaturenumberoptionalSampling temperature between 0 and 2. Higher values make the output more random; lower values more focused and deterministic.
top_pnumberoptionalNucleus sampling: consider only the tokens making up the top_p probability mass. Use this or temperature, not both.
max_tokensnumberoptionalThe maximum number of tokens to generate in the response.
stopstring | arrayoptionalUp to four sequences where generation stops. The stop sequence is not included in the output.
frequency_penaltynumberoptionalBetween -2 and 2. Positive values penalise tokens by how often they have already appeared, reducing repetition.
presence_penaltynumberoptionalBetween -2 and 2. Positive values penalise tokens that have already appeared, encouraging the model to introduce new topics.
seednumberoptionalIf set, the model samples as deterministically as possible so repeated requests with the same parameters return similar results.
nnumberoptionalHow many completions to generate for each request.
streambooleanoptionalStream the response as Server-Sent Events. See Streaming.
response_formatobjectoptionalConstrain the output format, e.g. to JSON. See JSON mode.
toolsarrayoptionalA list of functions the model may call. See Tools.
chat_template_kwargsobjectoptionalTurn reasoning on for this request. See Reasoning.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-api-key",
baseURL: "https://api.nebux.cloud/v1",
});
const completion = await client.chat.completions.create({
model: "qwen3-14b",
messages: [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Recommend me a book" },
{ role: "assistant", content: "Try The Shadow of the Wind" },
{ role: "user", content: "And a similar one?" },
],
});
console.log(completion.choices[0].message.content);{
"id": "chatcmpl-8f3a2b1c9d",
"object": "chat.completion",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "If you liked it, try Nada, by Carmen Laforet."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 47,
"completion_tokens": 13,
"total_tokens": 60
}
}Streaming
Set stream to true to receive the response as Server-Sent Events, with tokens arriving incrementally as they are generated. The stream ends with a data: [DONE] line.
Parameters
streambooleanrequiredSet to true to stream the response as Server-Sent Events.
stream_optionsobjectoptionalOptions that apply only when streaming.
include_usagebooleanEmit an extra final chunk with the token usage for the whole request.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-api-key",
baseURL: "https://api.nebux.cloud/v1",
});
const stream = await client.chat.completions.create({
model: "qwen3-14b",
messages: [{ role: "user", content: "Which running shoes do you recommend?" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}data: {
"id": "chatcmpl-8f3a2b1c9d",
"object": "chat.completion.chunk",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [{ "index": 0, "delta": { "role": "assistant" }, "finish_reason": null }]
}
data: {
"id": "chatcmpl-8f3a2b1c9d",
"object": "chat.completion.chunk",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [{ "index": 0, "delta": { "content": "For asphalt," }, "finish_reason": null }]
}
data: {
"id": "chatcmpl-8f3a2b1c9d",
"object": "chat.completion.chunk",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [{ "index": 0, "delta": { "content": " a neutral shoe with good cushioning" }, "finish_reason": null }]
}
data: {
"id": "chatcmpl-8f3a2b1c9d",
"object": "chat.completion.chunk",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [{ "index": 0, "delta": {}, "finish_reason": "stop" }]
}
data: [DONE]Reasoning
Let the model think before it answers. It is off by default: send chat_template_kwargs with enable_thinking set to true to turn it on for a request. The reasoning comes back in reasoning, beside content rather than inside it, so you can show it, log it or ignore it. It is billed as output tokens like the rest of the reply, and a reasoned answer can spend several times as many.
Parameters
chat_template_kwargsobjectrequiredExtra arguments for the model prompt template. Only enable_thinking is supported.
enable_thinkingbooleantrue asks the model to reason before answering. Defaults to false.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-api-key",
baseURL: "https://api.nebux.cloud/v1",
});
const res = await client.chat.completions.create({
model: "qwen3-14b",
messages: [
{ role: "user", content: "A coffee is 1.20. What are 7 with 15% off?" },
],
// chat_template_kwargs is outside the OpenAI library's types,
// so it travels as an extra request parameter.
// The other option is calling POST https://api.nebux.cloud/v1/chat/completions
// directly.
// @ts-expect-error
chat_template_kwargs: {"enable_thinking": true},
});
console.log(res.choices[0].message.content);{
"id": "chatcmpl-6b1d05e37a",
"object": "chat.completion",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"reasoning": "7 x 1.20 = 8.40. A 15% discount leaves 85%, so 8.40 x 0.85 = 7.14.",
"content": "7.14 EUR"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 22,
"completion_tokens": 61,
"total_tokens": 83
}
}JSON mode
Set response_format to json_object to constrain the model to emit syntactically valid JSON. Always instruct the model to produce JSON in a system or user message as well.
Parameters
response_formatobjectrequiredAn object specifying the format the model must output.
typestringSet to json_object to enable JSON mode.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-api-key",
baseURL: "https://api.nebux.cloud/v1",
});
const completion = await client.chat.completions.create({
model: "qwen3-14b",
messages: [
{ role: "system", content: "Classify each comment as good or bad, keeping its id. Reply as JSON." },
{ role: "user", content: "1: Great! 2: Awful 3: Good 4: Slow" },
],
response_format: { type: "json_object" },
});
console.log(JSON.parse(completion.choices[0].message.content));{
"id": "chatcmpl-2a7c4e9f10",
"object": "chat.completion",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "{\"good\": [{\"id\": 1, \"comment\": \"Great!\"}, {\"id\": 3, \"comment\": \"Good\"}], \"bad\": [{\"id\": 2, \"comment\": \"Awful\"}, {\"id\": 4, \"comment\": \"Slow\"}]}"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 58,
"completion_tokens": 51,
"total_tokens": 109
}
}Tools
Provide a list of functions the model can call. When the model decides to use one, it returns the call in tool_calls instead of a text reply; you run the function and send the result back as a tool message.
Parameters
toolsarrayrequiredA list of functions the model may call.
typestringThe tool type. Currently only function is supported.
functionobjectThe function definition.
namestringThe name of the function to call.
descriptionstringA description of what the function does, used by the model to decide when to call it.
parametersobjectThe function's parameters, described as a JSON Schema object.
tool_choicestring | objectoptionalControls tool use: auto, none, required, or a specific function to force.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-api-key",
baseURL: "https://api.nebux.cloud/v1",
});
const completion = await client.chat.completions.create({
model: "qwen3-14b",
messages: [{ role: "user", content: "What's the weather in València?" }],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Get the current weather for a location",
parameters: {
type: "object",
properties: { location: { type: "string" } },
required: ["location"],
},
},
},
],
});
console.log(completion.choices[0].message.tool_calls);{
"id": "chatcmpl-5b1d8e2a34",
"object": "chat.completion",
"created": 1719763200,
"model": "qwen3-14b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_a1b2c3",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"location\": \"València\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
],
"usage": {
"prompt_tokens": 82,
"completion_tokens": 19,
"total_tokens": 101
}
}Errors
Failed requests return a standard HTTP status code and a JSON body with an error object describing what went wrong.
Status codes
400invalid_request_errorThe request was malformed or a parameter was invalid.401invalid_api_keyThe API key is missing or invalid.404model_not_foundThe requested model does not exist or is not available to your organization.429rate_limit_exceededYou have exceeded your rate limit. Retry after a short back-off.500server_errorAn unexpected error occurred while processing the request.503engine_overloadedThe service is temporarily overloaded. Retry after a short back-off.{
"error": {
"message": "Incorrect API key provided.",
"type": "invalid_request_error",
"param": null,
"code": "invalid_api_key"
}
}Unsupported features
This API is compatible with OpenAI's chat completions interface. The following OpenAI features are not available.
- The Embeddings API.
- Assistants API (assistants, threads and runs).
- Fine-tuning.
- The Batch API.
- The Files and uploads API, and vector stores.
- Audio endpoints (speech synthesis and transcription).
- Image generation.
- The Moderations API.
- The Realtime API.