Why I Kept My Own LLM Gateway over LiteLLM and 9Router
I put four model profiles on LiteLLM, sent back two comments and two issues, read how 9Router handles the same cases, and still run 3,000 lines of my own. Here is what worked, what did not, and why.
I run a handful of services that call language models. For a long time each one named its own model, held its own provider key, picked its own reasoning effort and carried its own idea of what to do when the provider said no. Changing a provider meant editing every one of them.
This post is about the fix, which is small, and about a question the fix raises: there are good open source gateways already, so why run your own? I put the same setup on LiteLLM, read how 9Router handles the same cases, sent what I found back upstream, and kept my own gateway anyway. None of this is advice to avoid either project. It is a record of where they did not fit one narrow job.
Call a profile, not a model
The idea that did most of the work has nothing to do with which gateway serves it: application code never names a model. It names a profile.
I have four: free, lite, flash and pro. A profile is an ordered chain, and each link is a model at a reasoning effort:
| Profile | First choice | Then |
|---|---|---|
free | Gemma 4 31B on the Gemini API, effort high | nothing |
lite | Gemini 3.5 Flash-Lite on Vertex AI, effort low | nothing |
flash | GLM-5.3 Flash on Z.AI, effort low | Gemini 3.8 Flash, effort medium |
pro | GLM-5.3 on Z.AI, effort high | Gemini 3.1 Pro, effort high |
Two details matter more than they look.
The effort belongs to the profile. The same model at two efforts is two different products in price and in quality, so letting each caller choose one puts a pricing decision back into application code.
The rule is enforced, not agreed. A test fails the build if application code contains a provider's API hostname. Without that, a convention like this erodes one convenient exception at a time.
The idea is not new. LiteLLM calls it a model group, 9Router calls it a combo, OpenRouter has presets. What I add is only strictness: the profile is the only name a caller may use.
The same four profiles on LiteLLM
The proxy config for LiteLLM 1.103.0, with no database, is about fifty lines. This is the interesting half of it:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
model_list:
- model_name: lite
litellm_params:
model: vertex_ai/gemini-3.5-flash-lite
vertex_project: my-project
vertex_location: global
reasoning_effort: low
- model_name: flash
litellm_params:
model: zai/glm-5.3-flash
api_key: os.environ/ZAI_API_KEY
allowed_openai_params: [response_format]
extra_body: {reasoning_effort: low, thinking: {type: enabled, clear_thinking: false}}
- model_name: flash-backup
litellm_params:
model: vertex_ai/gemini-3.8-flash
vertex_project: my-project
vertex_location: global
reasoning_effort: medium
litellm_settings:
num_retries: 0
fallbacks: [{"flash": ["flash-backup"]}]
A lot worked, and worked quickly:
- All four profiles answered, and the cost LiteLLM reported per call matched the providers' published prices, GLM-5.3 included.
- The fallback from
flashto its backup worked, and the response said so in themodelfield and in anx-litellm-attempted-fallbacksheader. - The Anthropic-shaped
/v1/messagesendpoint served a Z.AI model to a client that only speaks that protocol. - Gemini's grounding metadata came back whole, sources and all.
- On the same question, the reasoning tokens spent were close to what I get calling the providers directly.
If you are starting today with mainstream providers, that list is most of what you need, from one YAML file.
Three things that did not fit
A fallback hid my own mistake
My first config put reasoning_effort: low directly on the Z.AI deployment. LiteLLM's Z.AI provider does not accept that parameter and raises UnsupportedParamsError before any request leaves the machine. That part is my mistake, and the error message says exactly how to fix it.
Nobody saw the message. The router treats that error like any other failure, so it fell back, and the proxy answered 200 from the backup model. The flash profile looked healthy while every call was served by Gemini. A one-word answer cost 0.000519 USD from the backup; once the config was fixed, the same call cost 0.0000044 USD from the model I had asked for. The only signs were one response header and the model field, and nothing was printed at the default log level.
It happened a second time that afternoon, with response_format.
You can send "disable_fallbacks": true on a request and the 400 appears at once, but that turns fallbacks off for real outages too. From reading async_function_with_fallbacks, every exception goes to the fallback list except the two kinds that have lists of their own (context window and content policy).
Effort is uneven across providers
On Gemini 3.x, reasoning_effort in the deployment's params is all it takes; LiteLLM translates it to the provider's own thinking level.
On Z.AI the parameter is rejected, so it has to travel as a raw body field, as in the config above. On Gemma 4 served by the Gemini API, LiteLLM turns the effort into a thinking budget, and Google answers 400: "Thinking budget is not supported for this model." Called directly, gemma-4-31b-it takes a thinking level, and only two of them: minimal and high. It refuses low, medium and any budget. I could not find that documented, so treat it as measured on 2026-10-06, not quoted.
JSON with a schema on Z.AI
By default LiteLLM does not send response_format to Z.AI at all. With a fallback configured, that is the silent failure above again.
Allowed through by hand, a json_schema response format comes back 200 with prose, because Z.AI does not take that type. What Z.AI does take is json_object. Send that and write the schema into the prompt, and the answer is valid JSON that matches. So it works, as long as every caller knows to do it.
What I sent back to LiteLLM
Two of these were already known, so I added evidence to the existing threads rather than opening duplicates:
- Effort on Z.AI: a reproduction on
1.103.0with GLM-5.3 and the workaround, on the open pull request. - JSON on Z.AI: three test results, including that
json_objectis honored natively, on the open issue. There is a pull request for that one too.
Two I could not find anywhere, so I opened them:
- Gemma 4 and
reasoning_effort, with the table of what the Gemini API accepts. - Do not fall back on
UnsupportedParamsError, as a feature request, with the config and the cost difference.
They are comments and issues, not pull requests, on purpose. Two already have someone else's patch. The fallback rule is a design decision that belongs to the maintainers. And for Gemma 4 there is no official documentation to build a mapping on, only my measurements.
What about 9Router?
9Router is built for a different job. Its README describes a local router that connects coding tools such as Claude Code, Codex, Cursor and Cline to more than forty providers, so that a subscription's quota is used first, then a cheap tier, then a free one. My calls come from background services with paid API keys, so most of what it is good at I would not use.
It does have the profile idea built in: a combo is a named chain of models, and the client asks for the combo by name.
I did not run it. I read its code for the three cases above, as of 2026-10-06:
- Fallback. A 4xx caused by the request itself is handed back to the caller instead of moving to the next model. The comment in
accountFallback.jsgives the same reason I would: such an error says nothing about the account, and hiding it hides the real cause. This is the behavior I wanted. - Effort on Z.AI. Handled explicitly in
thinkingUnified.js, down to the three levels GLM-5.3 accepts. - JSON with a schema. The schema-into-the-prompt translation exists in
default.js, but only for providers the user adds as generic OpenAI-compatible ones, not for the built-in GLM provider.
What kept me from it is shape, not quality. Its state, meaning providers, combos, keys and settings, lives in a SQLite database edited through a dashboard. I want the whole routing table in one file in git, changed by a commit.
What my own gateway does instead
It is about 3,000 lines of TypeScript with no runtime dependencies, and it does four things that map onto the sections above.
A wrong request never falls back. Provider failures are sorted into kinds. Rate limits, timeouts, quota and outages move on to the next link in the chain. A failure that says the request itself is wrong is returned to the caller, with the name of the model that refused it.
Every answer says who gave it. A successful response carries the provider, model and effort that served it, and the list of links that were skipped and why. "It returned 200" and "the model I chose answered" are different facts, and only the second one is worth an alert when it stops being true.
Effort is translated per provider, in one place. The profile says low; the adapter knows that this means a body field on Z.AI and a thinking level on Gemini, and that Gemma 4 takes only two levels.
A schema on Z.AI is rewritten inside the gateway, into json_object plus the schema in the prompt. Callers send a schema and get JSON, whichever model answers.
Why not just wait for the fixes
Because my need is narrow and theirs is not. I have three providers and a few hard rules. LiteLLM serves more than a hundred providers and has to stay compatible for all of them.
The two fixes I would need are in the queue. The pull request for JSON on Z.AI was opened on 2026-08-21, the one for effort on Z.AI on 2026-09-13, and neither is merged. For context, counted on 2026-10-06:
| LiteLLM repository | Count |
|---|---|
| Open issues | 1,796 |
| Open pull requests | 3,397 |
| Open pull requests older than three months | 574 |
| Pull requests merged in the last 30 days | 2,272 |
That is a project merging about seventy-five pull requests a day and still receiving them faster. For a less common provider like Z.AI, I cannot plan around the day a fix lands.
Reporting upstream and running something small of your own are not opposites. I did both on the same day. If those fixes are merged, the next person will not hit what I hit.
What it costs
Three thousand lines that I maintain alone, and three providers. Every new provider is an adapter I write by hand. LiteLLM also installed 118 Python packages where mine installs none, which mattered to me for a process that holds every provider key, and will not matter to most people.
If I needed a fifth provider, or per-user keys and budgets, LiteLLM would be the right choice and I would move.
Takeaways
- Put a profile between your code and the model names, with the effort inside it. This works on any gateway, including none.
- Check which model actually answered. A 200 from the fallback is a quiet way to pay many times more.
- When an existing tool does not fit, report what you found with a reproduction before you walk away.
