AI Inference

Give your app superpowers Integrate AI in a snap

Consume open-source language models through our OpenAI-compatible API.

  • Lower cost — just change the base URL and API key
  • Our own infrastructure in Europe

What can you build

Enrich your application with AI

Chat & conversational assistants

Chatbots and support agents that answer with your business context.

Text generation

Drafting, summaries, rewriting, proofreading and autocomplete suggestions.

Classification & sentiment analysis

Label tickets, emails or reviews and detect tone automatically.

Machine translation

Translate and adapt content across languages while keeping meaning and tone.

Code

Autocomplete and generation, explanation, review, refactoring and test generation.

Data extraction

Turn invoices, emails or documents into structured fields, ready for your database.

Capabilities

What our API can do

Easily integrate AI in your app

Point any OpenAI-compatible library to the Nebux base URL using your API key — no code to rewrite.

See the code

Real-time streaming

Show the response as it's being written, token by token. Or, if you'd rather, get it whole in one go.

See the code

Guaranteed JSON output

Force responses into valid JSON that matches your schema.

See the code

Function calling

The model calls your tools and functions when needed to answer with the data they return.

See the code

What to expect

No small print

Flexible billing, no lock-in

Pay as you go with clear per-token pricing — no monthly fees, hidden costs or reserved capacity you don't use.

Usage in detail

Always know what you're spending and on which models, month by month and against the month before.

Zero data retention

No prompt logging and no training on your data. What you send is never kept.

On our own infrastructure

The models run on our own hardware in our facilities in València, Spain; your prompts never pass through third-party providers.

Playground

Try and tune before writing code

Pricing

Pay only for the tokens you use

Qwen 3 14B

qwen3-14b

A token is about 4 characters and a million tokens is around 700,000 words — Don Quixote twice over, for about 22 cents.

Input
€0.12
per 1M tokens
Output
€0.22
per 1M tokens

Transparent per-token pricing — no subscriptions, no monthly fees.

Prices are shown without tax for .

Get started

Get started in minutes

  1. 1

    Create an account

    Sign up in seconds.

  2. 2

    Add a payment method

    You're billed in arrears, on a postpaid basis; you only pay for what you use.

  3. 3

    Get your API key

    Generate your key from the dashboard. Ready to use immediately.

  4. 4

    Plug in your SDK

    Already using an OpenAI-compatible SDK? Just swap the base URL and API key.

Ready to try AI Inference?

Or let us run it for you

Prefer to focus on your product? We maintain and evolve your systems as one more team inside your company.

Learn more about managed cloud