Skip to content
Less overhead. More intelligence.

Big ideas.
Tiny token bills.

Put qwen-3.8 to work with an API you already know. Simple, pay-per-token inference that leaves more room for what you’re building.

Start with $5No subscriptionYour existing SDK

OpenAI-compatible
Streaming responses
Tool calling
Credit never expires
01 / Simple pricing

Small prices. Nothing hidden.

One model. Three clear rates.
Pay only for the tokens you use with qwen-3.8.

Input tokensPay as you go
$0.09
/ 1 million tokens

Everything you send to the model.

Cached input
$0.03
/ 1 million tokens

Less to pay when context is reused.

Output tokens
$0.30
/ 1 million tokens

Everything the model sends back.

Estimate with 75% input + 25% output. No cache savings included.

Estimated monthly cost$1.43

All prices in USD, excluding VAT. Actual cost depends on your token mix.

02 / Familiar by design

One request. No migration.

The same schema you already send to OpenAI. Swap two lines of config and your app keeps working.

Read the API docs ↗
curl https://api.betterinfer.com/v1/chat/completions \
  -H "Authorization: Bearer sk_live_a1b2c3d4..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-3.8",
    "messages": [{
      "role": "user",
      "content": "Explain rate limiting."
    }],
    "stream": true
  }'
Example · 200 OK · streamed response
▌
OpenAI chat completions
POST /v1/chat/completions
Supported
OpenAI completions
POST /v1/completions
Supported
OpenAI model list
GET /v1/models
Supported
Anthropic messages
Messages API · coming soon
Planned
03 / Ready when you are

A little goes a long way.

Pay for what you use. Credit never expires, and there is nothing to cancel.

Buys you roughly70.2M tokens
Chat messages78K messages
Expiresnever
1
Load credit
Pick an amount and pay. It clears as soon as the provider confirms, and the balance is yours to keep.
2
Grab a key
One click in the dashboard. Create as many as you like and revoke any of them instantly.
3
Swap the base URL
Keep your OpenAI client exactly as it is. Only the endpoint and the key change.
04 / Meet the model

qwen-3.8

live

The only model we serve right now. We would rather run one model well than list twenty we cannot keep warm.

Context window128K
Input$0.09 / 1M
Cached input$0.03 / 1M
Output$0.30 / 1M
Streamingyes
Tool callingyes

More soon

We add a model when enough people ask and we can price it below everyone else. Tell us what you need and we will tell you honestly whether it is worth hosting.

No silent swaps
The model behind qwen-3.8 does not change without a version bump and a changelog entry.
No quantisation surprises
We publish the precision we serve. If it changes, the model id changes with it.
No training on your data
Prompts and completions are dropped after the request completes.
05 / Good questions

A little more clarity.

Pricing, integration and your data — the essentials before your first request.

The short version

  1. $5 to get started
  2. Your existing OpenAI-compatible SDK
  3. No monthly subscription
Read the quickstart
We run open weights on our own rented hardware and we do not carry a sales team, a free tier or a per-seat dashboard. The margin we skip is the discount you get.
Built for your next idea

Less on tokens.
More on possibility.

Your next project starts with a key and $5 in credit. Keep your tools. Make something great.

Start building

No plans. No monthly commitment.