// BACKED BY
Y COMBINATOR //
Run Open-source models. At lowest cost. Highest speed.
Open models the way they should always have been: ultra-fast and too cheap to meter.
/////////////////////////////////
@@**__**__**@@@@**__**__**@@
\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\
// WHY OPENSCALE //
The most efficient place to run open models
Cheaper, faster, and fully elastic, so you ship on open weights without managing a single GPU
COST
TYPICAL
OPENSCALE
=== === ===
Lowest cost
Up to 50% cheaper than typical providers. We serve open weights efficiently and pass the savings straight to you.
===
PAYMENTS
===
1M tokens
$0.39
10M tokens
$3.90
100M tokens
$39.00
1B tokens
$390.00
linear pricing, billed only for tokens used
Pay as you go
Only pay for the tokens you use. No idle GPUs, no seats, no minimums, no commitments.
<<<<<<<
SPEED
>>>>>>>
TYPICAL
OPENSCALE
=== === ===
Fastest inference
Industry-leading tokens per second on every model, with tuned kernels and warm capacity for low latency.
AUTOSCALING
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
^
Scale to infinity
Autoscaling that handles a single request or a million, with no capacity planning on your side.
// MODELS & PRICING //
The lowest prices for open models.
Run any open model for less than other gateways charge. Pay only for the tokens you use, with no GPUs to manage.
GLM 5.3 Flash
Z.ai

Kimi K3
Moonshot AI
// HOW IT WORKS //
From zero to inference in minutes
If you can call the OpenAI API, you can run on OpenScale. Three steps, no infrastructure to manage.
Pick a model
Choose from every major open model, all behind one OpenAI-compatible endpoint.
Point your code
Swap your base URL and key. Keep your existing OpenAI SDK, prompts and tooling exactly as they are.
Scale automatically
We handle batching, warm capacity and autoscaling. You ship, from the first request to millions.
// ZERO DATA RETENTION //
Your prompts are never stored
Every request is processed in memory and metered by token count. Nothing you send or receive is kept, on any model, in any region.
Nothing stored
Request and response bodies are never written to disk. Only 30 days of metadata is kept, for billing.
Gone after the reply
A request lives only as long as it takes to answer it. Once the tokens are counted, the content is discarded.
Every model, every region
The same guarantee on every model and in every region we serve from. Your data never trains anything.

