Big ideas.
Tiny token bills.
Put qwen-3.8 to work with an API you already know. Simple, pay-per-token inference that leaves more room for what you’re building.
Start with $5No subscriptionYour existing SDK
Small prices. Nothing hidden.
One model. Three clear rates.
Pay only for the tokens you use with qwen-3.8.
Everything you send to the model.
Less to pay when context is reused.
Everything the model sends back.
Estimate with 75% input + 25% output. No cache savings included.
All prices in USD, excluding VAT. Actual cost depends on your token mix.
One request. No migration.
The same schema you already send to OpenAI. Swap two lines of config and your app keeps working.
curl https://api.betterinfer.com/v1/chat/completions \
-H "Authorization: Bearer sk_live_a1b2c3d4..." \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-3.8",
"messages": [{
"role": "user",
"content": "Explain rate limiting."
}],
"stream": true
}'▌A little goes a long way.
Pay for what you use. Credit never expires, and there is nothing to cancel.
qwen-3.8
liveThe only model we serve right now. We would rather run one model well than list twenty we cannot keep warm.
More soon
We add a model when enough people ask and we can price it below everyone else. Tell us what you need and we will tell you honestly whether it is worth hosting.
A little more clarity.
Pricing, integration and your data — the essentials before your first request.
The short version
- $5 to get started
- Your existing OpenAI-compatible SDK
- No monthly subscription
Less on tokens.
More on possibility.
Your next project starts with a key and $5 in credit. Keep your tools. Make something great.
No plans. No monthly commitment.