Easily integrate AI in your app Directly compatible with any OpenAI client

Point any OpenAI-compatible library to the Nebux base URL using your API key — no code to rewrite.

baseURL"https://api.openai.com/v1"
baseURL"https://api.nebux.cloud/v1"
1import OpenAI from "openai";
2
3const client = new OpenAI({
4 apiKey: "your-api-key",
5 baseURL: "https://api.nebux.cloud/v1",
6});

You decide what the model remembers

Each request stands alone: the model only knows the messages you include in it. To keep a conversation going, send the history along with the new message and choose which model answers.

System · helpful assistant
Recommend me a book
Try "To Kill a Mockingbird"
And a similar one?
7const completion = await client.chat.completions.create({
8 model: "qwen3-14b",
9 messages: [
10 { role: "system", content: "You are a helpful assistant." },
11 { role: "user", content: "Recommend me a book" },
12 { role: "assistant", content: "Try The Shadow of the Wind" },
13 { role: "user", content: "And a similar one?" },
14 ],
15});
16
17console.log(completion.choices[0].message.content);

Real-time streaming

Show the response as it's being written, token by token. Or, if you'd rather, get it whole in one go.

Which running shoes do you recommend?
AssistantFor asphalt, a neutral shoe with good cushioning
7const stream = await client.chat.completions.create({
8 model: "qwen3-14b",
9 messages: [{ role: "user", content: "Which running shoes do you recommend?" }],
10 stream: true,
11});
12
13for await (const chunk of stream) {
14 process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
15}

Guaranteed JSON output

Force responses into valid JSON that matches your schema.

› Classify these comments:
"Great!""Awful""Good""Slow"
{
  "good": [{ "id": 1, "comment": "Great!" }, { "id": 3, "comment": "Good" }],
  "bad": [{ "id": 2, "comment": "Awful" }, { "id": 4, "comment": "Slow" }]
}
7const completion = await client.chat.completions.create({
8 model: "qwen3-14b",
9 messages: [
10 { role: "system", content: "Classify each comment as good or bad, keeping its id. Reply as JSON." },
11 { role: "user", content: "1: Great! 2: Awful 3: Good 4: Slow" },
12 ],
13 response_format: { type: "json_object" },
14});
15
16console.log(JSON.parse(completion.choices[0].message.content));

Reasoning, only when you need it

Let the model think before it answers on harder tasks. The reasoning comes back separately from the answer, so you can show it, log it or drop it, and you only pay for it on the requests where you turn it on.

A coffee is €1.20. What are 7 with 15% off?
Reasoning
7 × 1.20 = 8.408.40 − 15% = 7.14
€7.14
7const res = await client.chat.completions.create({
8 model: "qwen3-14b",
9 messages: [
10 { role: "user", content: "A coffee is 1.20. What are 7 with 15% off?" },
11 ],
12 // chat_template_kwargs is outside the OpenAI library's types,
13 // so it travels as an extra request parameter.
14 // The other option is calling POST https://api.nebux.cloud/v1/chat/completions
15 // directly.
16 // @ts-expect-error
17 chat_template_kwargs: {"enable_thinking": true},
18});
19
20console.log(res.choices[0].message.content);

You decide how it answers

Temperature, top-p, max tokens, penalties and stop sequences. The same question, in the tone and the length you want.

Describe the product in one line
temperature 0
An OpenAI-compatible inference service.
temperature 1.2
Open models, served from Europe, one line of code away.
7const res = await client.chat.completions.create({
8 model: "qwen3-14b",
9 messages: [
10 { role: "user", content: "Describe the product in one line" },
11 ],
12 temperature: 0.7,
13 max_tokens: 1024,
14 top_p: 0.95,
15 frequency_penalty: 0.4,
16 stop: ["\n\n"],
17});
18
19console.log(res.choices[0].message.content);

Repeatable results

Set a seed and the temperature to zero to get consistent answers across runs. Ideal for classification, data extraction or testing without surprises.

seed: 42temperature: 0
›Classify this reviewpositive
›Classify this reviewpositive
Same output
7const res = await client.chat.completions.create({
8 model: "qwen3-14b",
9 messages: [
10 { role: "user", content: "Classify this review" },
11 ],
12 temperature: 0,
13 seed: 42,
14});
15
16console.log(res.choices[0].message.content);

Function calling

The model calls your tools and functions when needed to answer with the data they return.

What's the weather in València?
get_weather("València")
→ { "temp": 18, "sky": "sunny" }
It's 25 °C and sunny in València ☀️
7const completion = await client.chat.completions.create({
8 model: "qwen3-14b",
9 messages: [{ role: "user", content: "What's the weather in València?" }],
10 tools: [
11 {
12 type: "function",
13 function: {
14 name: "get_weather",
15 description: "Get the current weather for a location",
16 parameters: {
17 type: "object",
18 properties: { location: { type: "string" } },
19 required: ["location"],
20 },
21 },
22 },
23 ],
24});
25
26console.log(completion.choices[0].message.tool_calls);

Try and tune before writing code

Copy ready-to-use code in any of the languages we document and start building.

Open the playground

The whole API, documented

Every endpoint parameter, with examples: streaming, guaranteed JSON, function calling and how to handle errors.

Documentation

Ready to try AI Inference?

Or let us run it for you

Prefer to focus on your product? We maintain and evolve your systems as one more team inside your company.

Learn more about managed cloud