[ Chaos ][ Clarity ]
No VC money

131+ free models. Smart routing. 4% markup.

The indie LLM router that tries free models first. OpenAI-compatible API, waterfall routing from free to cheap to premium. Built by one person + Claude, not 20 engineers in SF.

Live catalogue[ Route / Cost ]
[ Observed ]
131
$0
4%
1M

[ Flow / 01 ]

How waterfall routing works

Your request cascades down from free to paid until it gets a great answer. Most of the time, free is all you need.

01 / input

Send a request

Use our OpenAI-compatible API. Drop-in replacement — change one line of code and you are done.

02 / route

Smart routing

We try free models first (DeepSeek, Gemini, Llama), then cheap ones, then premium. Only escalate when needed.

03 / outcome

Save money

Most queries resolve on free models. You only pay when you actually need a premium model. Simple as that.

[ Catalogue / live ]

Featured models

From free reasoning models to flagship paid ones. Pick what fits your budget and task.

[ Capacity → choice ]
New
Z.ai: GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

ChatTool UseReasoning
01-ai/yi-large

NVIDIA NIM-hosted yi-large (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.

Free
Chat
adept/fuyu-8b

NVIDIA NIM-hosted fuyu-8b (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.

Free
VisionChat
ai21labs/jamba-1.5-large-instruct

NVIDIA NIM-hosted jamba-1.5-large-instruct (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.

Free
ChatTool Use
aisingapore/sea-lion-7b-instruct

LLM to represent and serve the linguistic and cultural diversity of Southeast Asia

Free
Chat
bge-m3

NVIDIA NIM-hosted bge-m3 (OpenAI-compatible endpoint at https://integrate.api.nvidia.com/v1). Added from live NIM model listing.

Free
Differentiator

Web search built in

Perplexity-quality answers without Perplexity pricing. Real-time, cited responses from any model.

Web Search Tier

Use model: "web-search" and we route through Gemini Search and Perplexity Sonar, cheapest first.

One line of code

Any Model + Web

Add web search to Claude, GPT, Llama, or DeepSeek via the Exa plugin. ~$0.003/request extra.

Works with 300+ models

Research-Grade

Perplexity Sonar for deep, multi-source cited research. Including R1-powered reasoning + search.

From $1/M tokens

Waterfall vs OpenRouter

They raised $40M to route to expensive models. We route to free ones first. Different philosophy.

OpenRouterWaterfall
Funding$40M raised, $500M valuationBootstrapped, profitable
Team20+ engineers in SFJulie + Claude
Fees5.5% credit purchase fee4% transparent markup
RoutingRoutes to paid models firstRoutes to FREE models first
Free modelsSome, as afterthought131+ free models, first-class
Web searchPlugin add-onBuilt-in routing
Philosophy"Which expensive model is best?""Can a free model handle this?"

Ready to save on AI?

Drop-in OpenAI-compatible API. Change one line of code. Start with free models, scale to premium when you need to.

$7/month VPS. Not $40M in funding. We pass the savings to you.