Kourier uses Google Analytics to understand aggregate website usage. Analytics is optional and is not used for targeted advertising. Read our Privacy Policy.

Open models. One coding plan to rule them all.

Serve DeepSeek V4 Flash through one OpenAI-compatible endpoint, with Qwen 3.6 35B coming soon. Two simple plans, $50/month or $100/month. No per-token metering.

route.live
$routingidle

The open model lineup.

One API key, every model. Switch by changing one parameter.

View plans
Thinking...
Latency-last
0msp50 0ms / p90 0ms / n 01400ms

Built for teams that ship with open models.

01

Open models, open weights.

Every model we serve has public weights and a permissive license. No vendor lock-in, no model you cannot run yourself. We are the fast path to models you already trust.

02

We run the GPUs.

Inference at scale is hard: batching, prefix caching, failover, capacity planning, SLOs. We handle all of it so you get the model, not the pager.

03

Simple plans, no surprises.

Two plans, one price each. Pro at $50/month for hands-on coding, Max at $100/month for long sessions and heavy use. No seat fees, no per-token metering, no surprise bills.

Works with the tools you already use.

The endpoint is OpenAI-compatible. Set a base URL and key, and your editor, agent, or CLI talks to open models immediately. Up to 262,144 input tokens with smart compaction behind every request.

  • Cursor
  • Zed
  • OpenCode
  • Cline
  • Oh My Pi
AGENT SESSIONdeepseek-v4-flash - 278,528 ctx

103,600 / 278,528 tokens · 37%

  • System3,4001%
  • Tools2,140<1%
  • History96,80035%
  • Turn1,260<1%
  • Free174,92863%

event - session resumed - 51.8% of window in use

37% of window

Inference you can measure.

Measured fleet latency and throughput, plus a rolling seven-day average of DeepSeek V4 Flash output speed when sufficient traffic is available.

Latency

time to first token, 24h

Throughput

aggregate tokens/sec, 24h

DeepSeek V4 Flash token speed

7d avg 92.8 tok/s · n=10018
successful measurable streams only, hourly trace

One coding plan to rule them all.

Two simple plans. Pro at $50/month for hands-on coding, Max at $100/month for long sessions and heavy use.

Pro

Hands-on coding

wt 10.0
$50/mo

billed monthly

Get an API key

Max

Long sessions, heavy use

wt 15.0
$100/mo

billed monthly

Questions, answered.

DeepSeek V4 Flash is available today. Qwen 3.6 35B is coming soon and will use the same OpenAI-compatible API when it reaches production.
Two flat monthly plans with no request-volume or token-usage quotas. Pro is $50/month with 3 simultaneous requests; Max is $100/month with 7, priority routing, and upgraded support. No usage credits, rolling quotas, seat fees, or surprise bills.
No. Prompts and completions are processed in memory and discarded after the response returns. We keep only billing metadata: timestamp, token counts, and the model id. We never train on your data.
We do not currently offer a contractual uptime SLA. We are deploying independent monitoring and historical incident reporting for the public status page.
Yes. If you have an open-weights model you want served, contact us. We deploy it on dedicated GPUs with the same API and observability as our standard lineup.
Yes. The endpoint accepts the OpenAI chat completions schema, so any client built for OpenAI works unchanged. Point the base URL at us, set your key, and the same code runs.

Start building with open models.

Choose Pro or Max to get started, or use an Omega invitation.