Skip to documentation
Better Infer / API documentation

Your stack.
Our inference.

From your first request to production. Connect your existing OpenAI client, stream a response, and put qwen-3.8 to work.

OpenAI-compatible · availableAnthropic · plannedHTTPS + JSON
Inference base URLhttps://api.betterinfer.com/v1
Compatible with the interface you know.Better Infer serves its own models through the OpenAI Chat Completions interface. Use a Better Infer key and model ID. Compatibility does not mean every OpenAI product, endpoint or model-specific feature is available here.

Quickstart

Start with an account, a positive credit balance and an API key. These examples make real, billable requests when you run them with your key. No request is sent by this documentation page.

01 / ACCOUNT

Create your account

Register or log in to open your dashboard.

02 / CREDIT

Add $5 or more

New accounts start without credit. Top up your balance before making requests.

03 / API KEY

Create a named key

Open API keys. Copy the full secret when it is created; it is only shown once.

1. Set your server environment

The examples consistently use BETTER_INFER_API_KEY. Replace the placeholder in your own environment. Use the same terminal to run the examples.

Set your API key
# Set this in your server environment; never commit it.
export BETTER_INFER_API_KEY="YOUR_BETTER_INFER_KEY"

2. Make a request

Choose your language. Python and Node.js use the openai package; cURL needs no SDK. The cURL examples use Bash syntax. Windows users can use the PowerShell example.

Your first completion
curl --fail-with-body "https://api.betterinfer.com/v1/chat/completions" \
  -H "Authorization: Bearer $BETTER_INFER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-3.8",
    "messages": [
      {"role": "user", "content": "Explain an API in one sentence."}
    ],
    "max_tokens": 128,
    "stream": false
  }'

The SDK base URL already includes /v1; do not append /chat/completions to the SDK configuration. When moving from another provider, update the base URL, API key and model ID.

3. Read the response

Read generated text from choices[0].message.content. The following is an illustrative response, not a live result. IDs, timestamps, content and token counts vary.

Example JSON response
{
  "id": "chatcmpl-example",
  "object": "chat.completion",
  "created": 1789776000,
  "model": "qwen-3.8",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "An API lets applications exchange data through a defined interface."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 18,
    "completion_tokens": 15,
    "total_tokens": 33
  }
}

Authentication

Every documented /v1 request uses bearer authentication, including GET /v1/models. The key owner must have a positive balance. Dashboard login cookies are not inference credentials.

Request headers
Authorization: Bearer sk_live_YOUR_KEY_SECRET
Content-Type: application/json
  • Send the full key after the exact Bearer prefix, including the space. An OpenAI key or a dashboard session will not work.
  • Store secrets in your server environment or secret manager. Do not include them in browser code, public repositories, URLs or client-visible logs.
  • Use a separate named key for each application or environment. Create and revoke keys from the dashboard.
  • If a secret is lost, create a replacement. If it is exposed, revoke it and deploy a replacement. A revoked key returns 401 on new requests.

Chat completions

POST/v1/chat/completions

Generate an assistant response from a conversation. Send JSON with a supported model ID and a messages array. For a multi-turn conversation, resend the relevant history with each request.

Chat completion request parameters
FieldTypeUsage
modelstring · requiredModel ID, for example qwen-3.8. Use the exact ID returned by the model list.
messagesarray · requiredConversation messages in order. Text messages use role and content.
messages[].rolestringUse system for instructions, user for input, assistant for prior responses and tool for tool results.
messages[].contentstringText content. Assistant tool-call messages may have null content; tool results need tool_call_id.
max_tokensinteger · optionalOutput token cap. Keep the prompt and requested output within the model’s context capacity.
temperaturenumber · optionalSampling control. Accepted range and default are determined by the serving model.
top_pnumber · optionalNucleus sampling control. Usually adjust this or temperature, rather than both.
stopstring or array · optionalStop sequence(s), where supported by the serving model.
streamboolean · optionalSet true for server-sent events. Omit or set false for one JSON response.
stream_options.include_usageboolean · optionalSet true to request token usage in the stream’s final usage chunk.
toolsarray · optionalFunction definitions. See the complete tool-calling example below.
tool_choicestring or object · optionalTool selection behavior. Model support and server configuration determine accepted options.
Model-dependent optionsOptional generation parameters are passed to the serving model. Do not assume support for OpenAI-only settings such as reasoning controls, strict structured outputs, multimodal inputs or hosted tools. Start with the minimal example and validate optional settings against your deployed model.
Multi-turn request body
{
  "model": "qwen-3.8",
  "messages": [
    {
      "role": "system",
      "content": "You explain concepts with short, practical examples."
    },
    {
      "role": "user",
      "content": "What is a cache?"
    },
    {
      "role": "assistant",
      "content": "A cache keeps frequently used data close by so it can be retrieved faster."
    },
    {
      "role": "user",
      "content": "Give me a web application example."
    }
  ],
  "max_tokens": 256,
  "temperature": 0.7
}

Response fields

Chat completion response fields
FieldMeaning
idCompletion identifier returned by the model server.
choices[].messageAssistant message. Inspect content for text or tool_calls for function requests.
choices[].finish_reasonCommon values are stop (finished), length (token limit) and tool_calls (tool result needed). Handle other values without crashing.
usageToken counts: prompt_tokens, completion_tokens and total_tokens. Cached input can appear in prompt_tokens_details.cached_tokens.

Streaming

Set stream: true on the chat endpoint to receive text/event-stream. Append each text delta as it arrives. A delta is a fragment, not the full answer, and some chunks contain no text.

Stream text and usage
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.betterinfer.com/v1",
    api_key=os.environ["BETTER_INFER_API_KEY"],
    timeout=120.0,  # Example timeout; adjust for your workload.
    max_retries=0,  # Keep retries explicit in your application.
)

stream = client.chat.completions.create(
    model="qwen-3.8",
    messages=[{"role": "user", "content": "Explain caching in three steps."}],
    max_tokens=256,
    stream=True,
    stream_options={"include_usage": True},
)

with stream:
    for chunk in stream:
        # The final usage chunk can have an empty choices array.
        if chunk.choices:
            text = chunk.choices[0].delta.content
            if text:
                print(text, end="", flush=True)
        if chunk.usage:
            print("\nUsage:", chunk.usage)
print()

Reading SSE directly

Each event is separated by a blank line. Decode the JSON after data:; stop on data: [DONE]. Network reads can split an event or combine several events, so use an SSE parser instead of parsing each TCP chunk as JSON.

Abbreviated SSE example
data: {"choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"choices":[{"index":0,"delta":{"content":"A cache"},"finish_reason":null}]}

data: {"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"choices":[],"usage":{"prompt_tokens":18,"completion_tokens":15,"total_tokens":33}}

data: [DONE]
  • Request stream_options.include_usage: true explicitly if you need usage in your client. The final usage chunk can have an empty choices array.
  • A stream interrupted before its end may never deliver final usage or [DONE]. Keep partial output separate from a completed answer.
  • Once streaming has begun, an HTTP status cannot be changed. Handle read failures and incomplete streams, not just the initial status.
  • If you relay through your own web server, disable response buffering and propagate cancellation when your user disconnects.

For the underlying event pattern, see the OpenAI streaming guide. The Better Infer endpoint and model stay as configured above.

Tool calling

Tools let the model request a function call. Your application executes that function and sends back the result. Better Infer does not run your code or access your systems on your behalf.

  1. Describe allowed functions with a name, description and JSON Schema parameters.
  2. Check the assistant message for tool_calls. A model may also answer without calling a tool.
  3. Validate the function name and parsed arguments, then run only the allowed function.
  4. Append the original assistant message and a role: "tool" message for every call, preserving its tool_call_id.
  5. Send the updated history to get the final answer.
Complete tool round trip · Python
import json
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.betterinfer.com/v1",
    api_key=os.environ["BETTER_INFER_API_KEY"],
    timeout=120.0,  # Example timeout; adjust for your workload.
    max_retries=0,  # Keep retries explicit in your application.
)

# A local, read-only demo function. Replace with your own implementation.
def get_shipping_days(country):
    return {"DE": 2, "CH": 3, "AT": 3}[country]

tools = [{
    "type": "function",
    "function": {
        "name": "get_shipping_days",
        "description": "Get estimated shipping days for a supported country.",
        "parameters": {
            "type": "object",
            "properties": {"country": {"type": "string", "enum": ["DE", "CH", "AT"]}},
            "required": ["country"],
            "additionalProperties": False,
        },
    },
}]
messages = [{"role": "user", "content": "How long does shipping to Switzerland take?"}]

first = client.chat.completions.create(
    model="qwen-3.8", messages=messages, tools=tools, max_tokens=256,
)
message = first.choices[0].message

if not message.tool_calls:
    print(message.content)
else:
    messages.append(message.model_dump(exclude_none=True))
    for call in message.tool_calls:
        # Never execute arbitrary names or trust model-generated arguments.
        if call.type != "function" or call.function.name != "get_shipping_days":
            raise ValueError("Unsupported tool")
        arguments = json.loads(call.function.arguments)
        if not isinstance(arguments, dict) or set(arguments) != {"country"}:
            raise ValueError("Unexpected tool arguments")
        country = arguments["country"]
        if not isinstance(country, str) or country not in ("DE", "CH", "AT"):
            raise ValueError("Unsupported country")
        messages.append({
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps({"days": get_shipping_days(country)}),
        })
    # Omit tools on this final call to keep this demo to one tool round.
    final = client.chat.completions.create(
        model="qwen-3.8", messages=messages, max_tokens=256,
    )
    print(final.choices[0].message.content)

This example uses fixed demo shipping estimates. In production, bound the number of tool rounds, handle invalid JSON and require your own authorization before a tool changes data. Tool availability depends on the selected model and its serving configuration; strict JSON Schema enforcement is not promised here.

With streaming, function arguments arrive in fragments. Assemble them per call index and parse only when complete. See the OpenAI function-calling guide for the message format.

Models

GET/v1/models

Discover model IDs before sending a generation request. This endpoint also requires a valid Better Infer key and a positive balance. The website currently advertises qwen-3.8; use the returned ID if your environment exposes a different deployment name.

List available models
curl --fail-with-body "https://api.betterinfer.com/v1/models" \
  -H "Authorization: Bearer $BETTER_INFER_API_KEY"
Illustrative model-list shape
{
  "object": "list",
  "data": [
    {
      "id": "qwen-3.8",
      "object": "model"
    }
  ]
}

Additional model metadata may be present. A model ID is not a guarantee of every optional capability. Do not infer context size, precision or parameter support from the name alone.

Text completions

POST/v1/completions

For existing prompt-based integrations, use the legacy text-completion interface. New conversational applications should use Chat Completions.

Text completion fields
FieldUsage
modelRequired model ID.
promptRequired text prompt. This interface does not accept a messages array.
max_tokensOptional output token cap.
temperatureOptional sampling parameter, subject to model support.
streamSet true for SSE. Text fragments use choices[].text rather than choices[].delta.content.
Legacy text completion
curl --fail-with-body "https://api.betterinfer.com/v1/completions" \
  -H "Authorization: Bearer $BETTER_INFER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-3.8",
    "prompt": "Three useful names for a caching library:",
    "max_tokens": 64,
    "temperature": 0.7
  }'

Read text from choices[0].text. Token counts use the same usage structure as Chat Completions. Model-specific prompt formatting still applies.

Usage & billing

Requests are authorized against your current credit balance. Completed requests are charged from reported token usage. Published rates below match the pricing shown on this website; all amounts are in USD, excluding VAT.

Published token pricing
Token categoryPrice per 1M tokens
Uncached input$0.09
Cached input$0.03
Output$0.30
Estimate the token cost
uncached_input = max(prompt_tokens - cached_tokens, 0)

cost_eur = (
    uncached_input * 0.09
    + cached_tokens * 0.03
    + completion_tokens * 0.3
) / 1_000_000

For example, 1,000 input tokens (200 cached) and 500 output tokens cost approximately $0.000228 at these rates. Cached tokens are a subset of input tokens, not extra tokens. Cache hits depend on the serving model and workload.

A balance check is not a hard per-request spending cap.Credit is checked before the request and debited after completion. An in-flight request, or several concurrent requests, can take the balance below zero. Further requests are blocked until the balance is positive again. Use output limits and bounded concurrency to control spend.

A request reported as failed is not charged. Billing reports are processed after requests finish, so dashboard updates are not necessarily instantaneous. A client-side timeout alone does not prove that generation failed; be careful when replaying a request.

Review your balance or add credit in Billing.

Errors & retries

Check the HTTP status before parsing the body. Authorization failures can have an empty body, proxy failures can be plain text, and model-server errors may be JSON. Do not require an OpenAI-shaped JSON error object for every failure.

HTTP errors and recovery
StatusMeaningWhat to do
400Invalid request or model parametersValidate JSON, required fields, token limits and optional settings. Correct the request before retrying.
401Missing, invalid or revoked keyCheck Authorization: Bearer and use the complete Better Infer key. Do not retry unchanged.
402Insufficient creditTop up the key owner’s balance. A new account needs credit before its first request.
404Unknown path or model, depending on serverVerify the base URL, endpoint and model ID. Avoid duplicated /v1 segments.
429Capacity or rate limit, if returned by the serving infrastructureReduce concurrency. Honor Retry-After if provided; otherwise use bounded exponential backoff.
502Upstream connection or response failureRetry a limited number of times with backoff. Check for repeated service failures.
503Service unavailable or model readiness timeoutAllow time for recovery or model startup, then retry with backoff.

No fixed requests-per-minute or tokens-per-minute limit is published here. Do not assume an OpenAI account’s limits apply to Better Infer. Rate-limit headers and Retry-After are not guaranteed.

Bounded retries

The example disables SDK retries and uses one explicit retry layer for transient HTTP statuses. Add Retry-After handling if your deployment provides it. Avoid retrying requests after partial streamed output, and account for possible duplicate work after ambiguous network failures.

Retry transient HTTP errors · Python
import random
import time
import openai
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.betterinfer.com/v1",
    api_key=os.environ["BETTER_INFER_API_KEY"],
    timeout=120.0,  # Example timeout; adjust for your workload.
    max_retries=0,  # Keep retries explicit in your application.
)

# One retry layer, three total attempts; no retries for auth or balance errors.
for attempt in range(3):
    try:
        completion = client.chat.completions.create(
            model="qwen-3.8",
            messages=[{"role": "user", "content": "Say hello."}],
            max_tokens=64,
        )
        print(completion.choices[0].message.content)
        break
    except openai.APIStatusError as error:
        if error.status_code not in (429, 502, 503) or attempt == 2:
            raise
        time.sleep(min(8, 2 ** attempt) + random.uniform(0, 0.5))
# Connection failures and timeouts propagate for the caller to investigate.
# A timed-out generation may already have completed; do not blindly replay it.

For diagnostics, record the status, model, request duration and a redacted error body. Do not log API keys or full prompts by default. Preserve any response identifier that is present; a specific request-ID header is not guaranteed.

Production checklist

  • Keep credentials on your server. Call Better Infer from your backend and expose only your own application endpoints to the browser.
  • Set timeouts deliberately. The first request after idle time can take longer while a model becomes ready. The 120-second SDK timeout in these examples is an application choice, not a latency promise.
  • Bound output and concurrency. Set a token budget, queue bursts and monitor your balance. A positive balance does not reserve enough credit for every concurrent request.
  • Handle all completion paths. Text, tool calls, token-limit finishes and interrupted streams need separate handling. Validate tool arguments before executing them.
  • Keep one retry policy. Coordinate SDK and application retries so nested retries do not multiply requests.
  • Test your exact request shape. OpenAI compatibility describes the documented interfaces, not support for every OpenAI model or feature.

Compatibility at a glance

API compatibility
InterfaceStatusIntegration
OpenAI Chat CompletionsAvailableUse the OpenAI SDK with the Better Infer base URL, key and model ID.
SSE streamingAvailableUse stream: true; explicitly request streamed usage if needed.
Function toolsModel-dependentUses the Chat Completions tool-call format. Validate your model’s configuration.
Model list / text completionsAvailableGET /v1/models and POST /v1/completions.
Responses, Assistants, files, images, audio, embeddingsNot part of this documented interfaceDo not assume these OpenAI endpoints or features are supported.
Anthropic MessagesPlannedNot available yet. See the roadmap section below.

The OpenAI Chat Completions reference provides protocol background. The supported Better Infer surface is the one documented on this page.

Anthropic compatibility

Planned · not available yet

Anthropic-style Messages compatibility is planned. The current public integration uses OpenAI-compatible endpoints; an Anthropic SDK cannot be pointed at the current base URL as a drop-in replacement.

Implementation details will be documented when available.The final base URL, authentication headers, version headers, request schema, streaming events and supported features are not published here yet. There is no release date to rely on.

Build against Chat Completions today. When the Messages interface ships, this section will contain its own tested quickstart and migration notes. Interface compatibility will not, by itself, imply availability of Anthropic-hosted Claude models.