Open models. One coding plan to rule them all.
Serve DeepSeek V4 Flash through one OpenAI-compatible endpoint, with Qwen 3.6 35B coming soon. Two simple plans, $50/month or $100/month. No per-token metering.
The open model lineup.
One API key, every model. Switch by changing one parameter.
Built for teams that ship with open models.
Open models, open weights.
Every model we serve has public weights and a permissive license. No vendor lock-in, no model you cannot run yourself. We are the fast path to models you already trust.
We run the GPUs.
Inference at scale is hard: batching, prefix caching, failover, capacity planning, SLOs. We handle all of it so you get the model, not the pager.
Simple plans, no surprises.
Two plans, one price each. Pro at $50/month for hands-on coding, Max at $100/month for long sessions and heavy use. No seat fees, no per-token metering, no surprise bills.
Works with the tools you already use.
The endpoint is OpenAI-compatible. Set a base URL and key, and your editor, agent, or CLI talks to open models immediately. Up to 262,144 input tokens with smart compaction behind every request.
Cursor
Zed
OpenCode
Cline
Oh My Pi
103,600 / 278,528 tokens · 37%
- System3,4001%
- Tools2,140<1%
- History96,80035%
- Turn1,260<1%
- Free174,92863%
event - session resumed - 51.8% of window in use
Inference you can measure.
Measured fleet latency and throughput, plus a rolling seven-day average of DeepSeek V4 Flash output speed when sufficient traffic is available.
Latency
time to first token, 24hThroughput
aggregate tokens/sec, 24hDeepSeek V4 Flash token speed
7d avg 92.8 tok/s · n=10018One coding plan to rule them all.
Two simple plans. Pro at $50/month for hands-on coding, Max at $100/month for long sessions and heavy use.
Questions, answered.
Start building with open models.
Choose Pro or Max to get started, or use an Omega invitation.