// BACKED BY

Y COMBINATOR //

Run Open-source models.At lowest cost. Highest speed.

Open models the way they should always have been: ultra-fast and too cheap to meter.

////////////////////

@@**__**__**@@

\\\\\\\\\\\\\\\\\\\\

Kimi
Z.ai
DeepSeek
MiMo
GLM
MiniMax

// WHY OPENSCALE //

The most efficient place to run open models

Cheaper, faster, and fully elastic, so you ship on open weights without managing a single GPU

COST

TYPICAL

OPENSCALE

=== === ===

Lowest cost

Up to 50% cheaper than typical providers. We serve open weights efficiently and pass the savings straight to you.

===

PAYMENTS

===

1M tokens

$0.39

10M tokens

$3.90

100M tokens

$39.00

1B tokens

$390.00

linear pricing, billed only for tokens used

Pay as you go

Only pay for the tokens you use. No idle GPUs, no seats, no minimums, no commitments.

<<<<<<<

SPEED

>>>>>>>

TYPICAL

OPENSCALE

=== === ===

Fastest inference

Industry-leading tokens per second on every model, with tuned kernels and warm capacity for low latency.

AUTOSCALING

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

^

Scale to infinity

Autoscaling that handles a single request or a million, with no capacity planning on your side.

// MODELS & PRICING //

The lowest prices for open models.

Run any open model for less than other gateways charge. Pay only for the tokens you use, with no GPUs to manage.

Qwen 3.6 35B A3B

Alibaba

IN / 1M
$0.005
CACHED / 1M
$0.0025
OUT / 1M
$0.045
95% OffBEST PROMO

DeepSeek V4.1 Flash

DeepSeek

IN / 1M
$0.12
CACHED / 1M
$0.006
OUT / 1M
$0.48
60% OffBEST PROMO

GLM 5.3 Flash

Z.ai

Coming soon
Coming soon

Kimi K3

Moonshot AI

Coming soon
Coming soon

// HOW IT WORKS //

From zero to inference in minutes

If you can call the OpenAI API, you can run on OpenScale. Three steps, no infrastructure to manage.

Pick a model

Choose from every major open model, all behind one OpenAI-compatible endpoint.

Point your code

Swap your base URL and key. Keep your existing OpenAI SDK, prompts and tooling exactly as they are.

Scale automatically

We handle batching, warm capacity and autoscaling. You ship, from the first request to millions.

// ZERO DATA RETENTION //

Your prompts are never stored

Every request is processed in memory and metered by token count. Nothing you send or receive is kept, on any model, in any region.

Nothing stored

Request and response bodies are never written to disk. Only 30 days of metadata is kept, for billing.

Gone after the reply

A request lives only as long as it takes to answer it. Once the tokens are counted, the content is discarded.

Every model, every region

The same guarantee on every model and in every region we serve from. Your data never trains anything.