131+ free models. Smart routing. 4% markup.
The indie LLM router that tries free models first. OpenAI-compatible API, waterfall routing from free to cheap to premium. Built by one person + Claude, not 20 engineers in SF.
[ Flow / 01 ]
How waterfall routing works
Your request cascades down from free to paid until it gets a great answer. Most of the time, free is all you need.
01 / input
Send a request
Use our OpenAI-compatible API. Drop-in replacement — change one line of code and you are done.
02 / route
Smart routing
We try free models first (DeepSeek, Gemini, Llama), then cheap ones, then premium. Only escalate when needed.
03 / outcome
Save money
Most queries resolve on free models. You only pay when you actually need a premium model. Simple as that.
[ Catalogue / live ]
Featured models
From free reasoning models to flagship paid ones. Pick what fits your budget and task.
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
NVIDIA NIM-hosted yi-large (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.
NVIDIA NIM-hosted fuyu-8b (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.
NVIDIA NIM-hosted jamba-1.5-large-instruct (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.
LLM to represent and serve the linguistic and cultural diversity of Southeast Asia
NVIDIA NIM-hosted bge-m3 (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.
Web search built in
Perplexity-quality answers without Perplexity pricing. Real-time, cited responses from any model.
Web Search Tier
Use model: "web-search" and we route through Gemini Search and Perplexity Sonar, cheapest first.
Any Model + Web
Add web search to Claude, GPT, Llama, or DeepSeek via the Exa plugin. ~$0.003/request extra.
Research-Grade
Perplexity Sonar for deep, multi-source cited research. Including R1-powered reasoning + search.
Waterfall vs OpenRouter
They raised $40M to route to expensive models. We route to free ones first. Different philosophy.
| OpenRouter | Waterfall | |
|---|---|---|
| Funding | $40M raised, $500M valuation | Bootstrapped, profitable |
| Team | 20+ engineers in SF | Julie + Claude |
| Fees | 5.5% credit purchase fee | 4% transparent markup |
| Routing | Routes to paid models first | Routes to FREE models first |
| Free models | Some, as afterthought | 131+ free models, first-class |
| Web search | Plugin add-on | Built-in routing |
| Philosophy | "Which expensive model is best?" | "Can a free model handle this?" |
Ready to save on AI?
Drop-in OpenAI-compatible API. Change one line of code. Start with free models, scale to premium when you need to.
$7/month VPS. Not $40M in funding. We pass the savings to you.