Your stack.
Our inference.
From your first request to production. Connect your existing OpenAI client, stream a response, and put qwen-3.8 to work.
https://api.betterinfer.com/v1Quickstart
Start with an account, a positive credit balance and an API key. These examples make real, billable requests when you run them with your key. No request is sent by this documentation page.
Add $5 or more
New accounts start without credit. Top up your balance before making requests.
Create a named key
Open API keys. Copy the full secret when it is created; it is only shown once.
1. Set your server environment
The examples consistently use BETTER_INFER_API_KEY. Replace the placeholder in your own environment. Use the same terminal to run the examples.
# Set this in your server environment; never commit it.
export BETTER_INFER_API_KEY="YOUR_BETTER_INFER_KEY"2. Make a request
Choose your language. Python and Node.js use the openai package; cURL needs no SDK. The cURL examples use Bash syntax. Windows users can use the PowerShell example.
curl --fail-with-body "https://api.betterinfer.com/v1/chat/completions" \
-H "Authorization: Bearer $BETTER_INFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-3.8",
"messages": [
{"role": "user", "content": "Explain an API in one sentence."}
],
"max_tokens": 128,
"stream": false
}'The SDK base URL already includes /v1; do not append /chat/completions to the SDK configuration. When moving from another provider, update the base URL, API key and model ID.
3. Read the response
Read generated text from choices[0].message.content. The following is an illustrative response, not a live result. IDs, timestamps, content and token counts vary.
{
"id": "chatcmpl-example",
"object": "chat.completion",
"created": 1789776000,
"model": "qwen-3.8",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "An API lets applications exchange data through a defined interface."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 18,
"completion_tokens": 15,
"total_tokens": 33
}
}Authentication
Every documented /v1 request uses bearer authentication, including GET /v1/models. The key owner must have a positive balance. Dashboard login cookies are not inference credentials.
Authorization: Bearer sk_live_YOUR_KEY_SECRET
Content-Type: application/json- Send the full key after the exact
Bearerprefix, including the space. An OpenAI key or a dashboard session will not work. - Store secrets in your server environment or secret manager. Do not include them in browser code, public repositories, URLs or client-visible logs.
- Use a separate named key for each application or environment. Create and revoke keys from the dashboard.
- If a secret is lost, create a replacement. If it is exposed, revoke it and deploy a replacement. A revoked key returns
401on new requests.
Chat completions
/v1/chat/completionsGenerate an assistant response from a conversation. Send JSON with a supported model ID and a messages array. For a multi-turn conversation, resend the relevant history with each request.
| Field | Type | Usage |
|---|---|---|
model | string · required | Model ID, for example qwen-3.8. Use the exact ID returned by the model list. |
messages | array · required | Conversation messages in order. Text messages use role and content. |
messages[].role | string | Use system for instructions, user for input, assistant for prior responses and tool for tool results. |
messages[].content | string | Text content. Assistant tool-call messages may have null content; tool results need tool_call_id. |
max_tokens | integer · optional | Output token cap. Keep the prompt and requested output within the model’s context capacity. |
temperature | number · optional | Sampling control. Accepted range and default are determined by the serving model. |
top_p | number · optional | Nucleus sampling control. Usually adjust this or temperature, rather than both. |
stop | string or array · optional | Stop sequence(s), where supported by the serving model. |
stream | boolean · optional | Set true for server-sent events. Omit or set false for one JSON response. |
stream_options.include_usage | boolean · optional | Set true to request token usage in the stream’s final usage chunk. |
tools | array · optional | Function definitions. See the complete tool-calling example below. |
tool_choice | string or object · optional | Tool selection behavior. Model support and server configuration determine accepted options. |
{
"model": "qwen-3.8",
"messages": [
{
"role": "system",
"content": "You explain concepts with short, practical examples."
},
{
"role": "user",
"content": "What is a cache?"
},
{
"role": "assistant",
"content": "A cache keeps frequently used data close by so it can be retrieved faster."
},
{
"role": "user",
"content": "Give me a web application example."
}
],
"max_tokens": 256,
"temperature": 0.7
}Response fields
| Field | Meaning |
|---|---|
id | Completion identifier returned by the model server. |
choices[].message | Assistant message. Inspect content for text or tool_calls for function requests. |
choices[].finish_reason | Common values are stop (finished), length (token limit) and tool_calls (tool result needed). Handle other values without crashing. |
usage | Token counts: prompt_tokens, completion_tokens and total_tokens. Cached input can appear in prompt_tokens_details.cached_tokens. |
Streaming
Set stream: true on the chat endpoint to receive text/event-stream. Append each text delta as it arrives. A delta is a fragment, not the full answer, and some chunks contain no text.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.betterinfer.com/v1",
api_key=os.environ["BETTER_INFER_API_KEY"],
timeout=120.0, # Example timeout; adjust for your workload.
max_retries=0, # Keep retries explicit in your application.
)
stream = client.chat.completions.create(
model="qwen-3.8",
messages=[{"role": "user", "content": "Explain caching in three steps."}],
max_tokens=256,
stream=True,
stream_options={"include_usage": True},
)
with stream:
for chunk in stream:
# The final usage chunk can have an empty choices array.
if chunk.choices:
text = chunk.choices[0].delta.content
if text:
print(text, end="", flush=True)
if chunk.usage:
print("\nUsage:", chunk.usage)
print()Reading SSE directly
Each event is separated by a blank line. Decode the JSON after data:; stop on data: [DONE]. Network reads can split an event or combine several events, so use an SSE parser instead of parsing each TCP chunk as JSON.
data: {"choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"choices":[{"index":0,"delta":{"content":"A cache"},"finish_reason":null}]}
data: {"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"choices":[],"usage":{"prompt_tokens":18,"completion_tokens":15,"total_tokens":33}}
data: [DONE]- Request
stream_options.include_usage: trueexplicitly if you need usage in your client. The final usage chunk can have an emptychoicesarray. - A stream interrupted before its end may never deliver final usage or
[DONE]. Keep partial output separate from a completed answer. - Once streaming has begun, an HTTP status cannot be changed. Handle read failures and incomplete streams, not just the initial status.
- If you relay through your own web server, disable response buffering and propagate cancellation when your user disconnects.
For the underlying event pattern, see the OpenAI streaming guide. The Better Infer endpoint and model stay as configured above.
Tool calling
Tools let the model request a function call. Your application executes that function and sends back the result. Better Infer does not run your code or access your systems on your behalf.
- Describe allowed functions with a name, description and JSON Schema parameters.
- Check the assistant message for
tool_calls. A model may also answer without calling a tool. - Validate the function name and parsed arguments, then run only the allowed function.
- Append the original assistant message and a
role: "tool"message for every call, preserving itstool_call_id. - Send the updated history to get the final answer.
import json
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.betterinfer.com/v1",
api_key=os.environ["BETTER_INFER_API_KEY"],
timeout=120.0, # Example timeout; adjust for your workload.
max_retries=0, # Keep retries explicit in your application.
)
# A local, read-only demo function. Replace with your own implementation.
def get_shipping_days(country):
return {"DE": 2, "CH": 3, "AT": 3}[country]
tools = [{
"type": "function",
"function": {
"name": "get_shipping_days",
"description": "Get estimated shipping days for a supported country.",
"parameters": {
"type": "object",
"properties": {"country": {"type": "string", "enum": ["DE", "CH", "AT"]}},
"required": ["country"],
"additionalProperties": False,
},
},
}]
messages = [{"role": "user", "content": "How long does shipping to Switzerland take?"}]
first = client.chat.completions.create(
model="qwen-3.8", messages=messages, tools=tools, max_tokens=256,
)
message = first.choices[0].message
if not message.tool_calls:
print(message.content)
else:
messages.append(message.model_dump(exclude_none=True))
for call in message.tool_calls:
# Never execute arbitrary names or trust model-generated arguments.
if call.type != "function" or call.function.name != "get_shipping_days":
raise ValueError("Unsupported tool")
arguments = json.loads(call.function.arguments)
if not isinstance(arguments, dict) or set(arguments) != {"country"}:
raise ValueError("Unexpected tool arguments")
country = arguments["country"]
if not isinstance(country, str) or country not in ("DE", "CH", "AT"):
raise ValueError("Unsupported country")
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps({"days": get_shipping_days(country)}),
})
# Omit tools on this final call to keep this demo to one tool round.
final = client.chat.completions.create(
model="qwen-3.8", messages=messages, max_tokens=256,
)
print(final.choices[0].message.content)This example uses fixed demo shipping estimates. In production, bound the number of tool rounds, handle invalid JSON and require your own authorization before a tool changes data. Tool availability depends on the selected model and its serving configuration; strict JSON Schema enforcement is not promised here.
With streaming, function arguments arrive in fragments. Assemble them per call index and parse only when complete. See the OpenAI function-calling guide for the message format.
Models
/v1/modelsDiscover model IDs before sending a generation request. This endpoint also requires a valid Better Infer key and a positive balance. The website currently advertises qwen-3.8; use the returned ID if your environment exposes a different deployment name.
curl --fail-with-body "https://api.betterinfer.com/v1/models" \
-H "Authorization: Bearer $BETTER_INFER_API_KEY"{
"object": "list",
"data": [
{
"id": "qwen-3.8",
"object": "model"
}
]
}Additional model metadata may be present. A model ID is not a guarantee of every optional capability. Do not infer context size, precision or parameter support from the name alone.
Text completions
/v1/completionsFor existing prompt-based integrations, use the legacy text-completion interface. New conversational applications should use Chat Completions.
| Field | Usage |
|---|---|
model | Required model ID. |
prompt | Required text prompt. This interface does not accept a messages array. |
max_tokens | Optional output token cap. |
temperature | Optional sampling parameter, subject to model support. |
stream | Set true for SSE. Text fragments use choices[].text rather than choices[].delta.content. |
curl --fail-with-body "https://api.betterinfer.com/v1/completions" \
-H "Authorization: Bearer $BETTER_INFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-3.8",
"prompt": "Three useful names for a caching library:",
"max_tokens": 64,
"temperature": 0.7
}'Read text from choices[0].text. Token counts use the same usage structure as Chat Completions. Model-specific prompt formatting still applies.
Usage & billing
Requests are authorized against your current credit balance. Completed requests are charged from reported token usage. Published rates below match the pricing shown on this website; all amounts are in USD, excluding VAT.
| Token category | Price per 1M tokens |
|---|---|
| Uncached input | $0.09 |
| Cached input | $0.03 |
| Output | $0.30 |
uncached_input = max(prompt_tokens - cached_tokens, 0)
cost_eur = (
uncached_input * 0.09
+ cached_tokens * 0.03
+ completion_tokens * 0.3
) / 1_000_000For example, 1,000 input tokens (200 cached) and 500 output tokens cost approximately $0.000228 at these rates. Cached tokens are a subset of input tokens, not extra tokens. Cache hits depend on the serving model and workload.
A request reported as failed is not charged. Billing reports are processed after requests finish, so dashboard updates are not necessarily instantaneous. A client-side timeout alone does not prove that generation failed; be careful when replaying a request.
Review your balance or add credit in Billing.
Errors & retries
Check the HTTP status before parsing the body. Authorization failures can have an empty body, proxy failures can be plain text, and model-server errors may be JSON. Do not require an OpenAI-shaped JSON error object for every failure.
| Status | Meaning | What to do |
|---|---|---|
| 400 | Invalid request or model parameters | Validate JSON, required fields, token limits and optional settings. Correct the request before retrying. |
| 401 | Missing, invalid or revoked key | Check Authorization: Bearer and use the complete Better Infer key. Do not retry unchanged. |
| 402 | Insufficient credit | Top up the key owner’s balance. A new account needs credit before its first request. |
| 404 | Unknown path or model, depending on server | Verify the base URL, endpoint and model ID. Avoid duplicated /v1 segments. |
| 429 | Capacity or rate limit, if returned by the serving infrastructure | Reduce concurrency. Honor Retry-After if provided; otherwise use bounded exponential backoff. |
| 502 | Upstream connection or response failure | Retry a limited number of times with backoff. Check for repeated service failures. |
| 503 | Service unavailable or model readiness timeout | Allow time for recovery or model startup, then retry with backoff. |
No fixed requests-per-minute or tokens-per-minute limit is published here. Do not assume an OpenAI account’s limits apply to Better Infer. Rate-limit headers and Retry-After are not guaranteed.
Bounded retries
The example disables SDK retries and uses one explicit retry layer for transient HTTP statuses. Add Retry-After handling if your deployment provides it. Avoid retrying requests after partial streamed output, and account for possible duplicate work after ambiguous network failures.
import random
import time
import openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.betterinfer.com/v1",
api_key=os.environ["BETTER_INFER_API_KEY"],
timeout=120.0, # Example timeout; adjust for your workload.
max_retries=0, # Keep retries explicit in your application.
)
# One retry layer, three total attempts; no retries for auth or balance errors.
for attempt in range(3):
try:
completion = client.chat.completions.create(
model="qwen-3.8",
messages=[{"role": "user", "content": "Say hello."}],
max_tokens=64,
)
print(completion.choices[0].message.content)
break
except openai.APIStatusError as error:
if error.status_code not in (429, 502, 503) or attempt == 2:
raise
time.sleep(min(8, 2 ** attempt) + random.uniform(0, 0.5))
# Connection failures and timeouts propagate for the caller to investigate.
# A timed-out generation may already have completed; do not blindly replay it.For diagnostics, record the status, model, request duration and a redacted error body. Do not log API keys or full prompts by default. Preserve any response identifier that is present; a specific request-ID header is not guaranteed.
Production checklist
- Keep credentials on your server. Call Better Infer from your backend and expose only your own application endpoints to the browser.
- Set timeouts deliberately. The first request after idle time can take longer while a model becomes ready. The 120-second SDK timeout in these examples is an application choice, not a latency promise.
- Bound output and concurrency. Set a token budget, queue bursts and monitor your balance. A positive balance does not reserve enough credit for every concurrent request.
- Handle all completion paths. Text, tool calls, token-limit finishes and interrupted streams need separate handling. Validate tool arguments before executing them.
- Keep one retry policy. Coordinate SDK and application retries so nested retries do not multiply requests.
- Test your exact request shape. OpenAI compatibility describes the documented interfaces, not support for every OpenAI model or feature.
Compatibility at a glance
| Interface | Status | Integration |
|---|---|---|
| OpenAI Chat Completions | Available | Use the OpenAI SDK with the Better Infer base URL, key and model ID. |
| SSE streaming | Available | Use stream: true; explicitly request streamed usage if needed. |
| Function tools | Model-dependent | Uses the Chat Completions tool-call format. Validate your model’s configuration. |
| Model list / text completions | Available | GET /v1/models and POST /v1/completions. |
| Responses, Assistants, files, images, audio, embeddings | Not part of this documented interface | Do not assume these OpenAI endpoints or features are supported. |
| Anthropic Messages | Planned | Not available yet. See the roadmap section below. |
The OpenAI Chat Completions reference provides protocol background. The supported Better Infer surface is the one documented on this page.
Anthropic compatibility
Anthropic-style Messages compatibility is planned. The current public integration uses OpenAI-compatible endpoints; an Anthropic SDK cannot be pointed at the current base URL as a drop-in replacement.
Build against Chat Completions today. When the Messages interface ships, this section will contain its own tested quickstart and migration notes. Interface compatibility will not, by itself, imply availability of Anthropic-hosted Claude models.