Post

Why I Kept My Own LLM Gateway over LiteLLM and 9Router

I put four model profiles on LiteLLM, sent back two comments and two issues, read how 9Router handles the same cases, and still run 3,000 lines of my own. Here is what worked, what did not, and why.

Why I Kept My Own LLM Gateway over LiteLLM and 9Router

I run a handful of services that call language models. For a long time each one named its own model, held its own provider key, picked its own reasoning effort and carried its own idea of what to do when the provider said no. Changing a provider meant editing every one of them.

This post is about the fix, which is small, and about a question the fix raises: there are good open source gateways already, so why run your own? I put the same setup on LiteLLM, read how 9Router handles the same cases, sent what I found back upstream, and kept my own gateway anyway. None of this is advice to avoid either project. It is a record of where they did not fit one narrow job.

Call a profile, not a model

The idea that did most of the work has nothing to do with which gateway serves it: application code never names a model. It names a profile.

I have four: free, lite, flash and pro. A profile is an ordered chain, and each link is a model at a reasoning effort:

ProfileFirst choiceThen
freeGemma 4 31B on the Gemini API, effort highnothing
liteGemini 3.5 Flash-Lite on Vertex AI, effort lownothing
flashGLM-5.3 Flash on Z.AI, effort lowGemini 3.8 Flash, effort medium
proGLM-5.3 on Z.AI, effort highGemini 3.1 Pro, effort high

Two details matter more than they look.

The effort belongs to the profile. The same model at two efforts is two different products in price and in quality, so letting each caller choose one puts a pricing decision back into application code.

The rule is enforced, not agreed. A test fails the build if application code contains a provider's API hostname. Without that, a convention like this erodes one convenient exception at a time.

The idea is not new. LiteLLM calls it a model group, 9Router calls it a combo, OpenRouter has presets. What I add is only strictness: the profile is the only name a caller may use.

The same four profiles on LiteLLM

The proxy config for LiteLLM 1.103.0, with no database, is about fifty lines. This is the interesting half of it:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
model_list:
  - model_name: lite
    litellm_params:
      model: vertex_ai/gemini-3.5-flash-lite
      vertex_project: my-project
      vertex_location: global
      reasoning_effort: low
  - model_name: flash
    litellm_params:
      model: zai/glm-5.3-flash
      api_key: os.environ/ZAI_API_KEY
      allowed_openai_params: [response_format]
      extra_body: {reasoning_effort: low, thinking: {type: enabled, clear_thinking: false}}
  - model_name: flash-backup
    litellm_params:
      model: vertex_ai/gemini-3.8-flash
      vertex_project: my-project
      vertex_location: global
      reasoning_effort: medium

litellm_settings:
  num_retries: 0
  fallbacks: [{"flash": ["flash-backup"]}]

A lot worked, and worked quickly:

  • All four profiles answered, and the cost LiteLLM reported per call matched the providers' published prices, GLM-5.3 included.
  • The fallback from flash to its backup worked, and the response said so in the model field and in an x-litellm-attempted-fallbacks header.
  • The Anthropic-shaped /v1/messages endpoint served a Z.AI model to a client that only speaks that protocol.
  • Gemini's grounding metadata came back whole, sources and all.
  • On the same question, the reasoning tokens spent were close to what I get calling the providers directly.

If you are starting today with mainstream providers, that list is most of what you need, from one YAML file.

Three things that did not fit

A fallback hid my own mistake

My first config put reasoning_effort: low directly on the Z.AI deployment. LiteLLM's Z.AI provider does not accept that parameter and raises UnsupportedParamsError before any request leaves the machine. That part is my mistake, and the error message says exactly how to fix it.

Nobody saw the message. The router treats that error like any other failure, so it fell back, and the proxy answered 200 from the backup model. The flash profile looked healthy while every call was served by Gemini. A one-word answer cost 0.000519 USD from the backup; once the config was fixed, the same call cost 0.0000044 USD from the model I had asked for. The only signs were one response header and the model field, and nothing was printed at the default log level.

It happened a second time that afternoon, with response_format.

You can send "disable_fallbacks": true on a request and the 400 appears at once, but that turns fallbacks off for real outages too. From reading async_function_with_fallbacks, every exception goes to the fallback list except the two kinds that have lists of their own (context window and content policy).

Effort is uneven across providers

On Gemini 3.x, reasoning_effort in the deployment's params is all it takes; LiteLLM translates it to the provider's own thinking level.

On Z.AI the parameter is rejected, so it has to travel as a raw body field, as in the config above. On Gemma 4 served by the Gemini API, LiteLLM turns the effort into a thinking budget, and Google answers 400: "Thinking budget is not supported for this model." Called directly, gemma-4-31b-it takes a thinking level, and only two of them: minimal and high. It refuses low, medium and any budget. I could not find that documented, so treat it as measured on 2026-10-06, not quoted.

JSON with a schema on Z.AI

By default LiteLLM does not send response_format to Z.AI at all. With a fallback configured, that is the silent failure above again.

Allowed through by hand, a json_schema response format comes back 200 with prose, because Z.AI does not take that type. What Z.AI does take is json_object. Send that and write the schema into the prompt, and the answer is valid JSON that matches. So it works, as long as every caller knows to do it.

What I sent back to LiteLLM

Two of these were already known, so I added evidence to the existing threads rather than opening duplicates:

Two I could not find anywhere, so I opened them:

They are comments and issues, not pull requests, on purpose. Two already have someone else's patch. The fallback rule is a design decision that belongs to the maintainers. And for Gemma 4 there is no official documentation to build a mapping on, only my measurements.

What about 9Router?

9Router is built for a different job. Its README describes a local router that connects coding tools such as Claude Code, Codex, Cursor and Cline to more than forty providers, so that a subscription's quota is used first, then a cheap tier, then a free one. My calls come from background services with paid API keys, so most of what it is good at I would not use.

It does have the profile idea built in: a combo is a named chain of models, and the client asks for the combo by name.

I did not run it. I read its code for the three cases above, as of 2026-10-06:

  • Fallback. A 4xx caused by the request itself is handed back to the caller instead of moving to the next model. The comment in accountFallback.js gives the same reason I would: such an error says nothing about the account, and hiding it hides the real cause. This is the behavior I wanted.
  • Effort on Z.AI. Handled explicitly in thinkingUnified.js, down to the three levels GLM-5.3 accepts.
  • JSON with a schema. The schema-into-the-prompt translation exists in default.js, but only for providers the user adds as generic OpenAI-compatible ones, not for the built-in GLM provider.

What kept me from it is shape, not quality. Its state, meaning providers, combos, keys and settings, lives in a SQLite database edited through a dashboard. I want the whole routing table in one file in git, changed by a commit.

What my own gateway does instead

It is about 3,000 lines of TypeScript with no runtime dependencies, and it does four things that map onto the sections above.

A wrong request never falls back. Provider failures are sorted into kinds. Rate limits, timeouts, quota and outages move on to the next link in the chain. A failure that says the request itself is wrong is returned to the caller, with the name of the model that refused it.

Every answer says who gave it. A successful response carries the provider, model and effort that served it, and the list of links that were skipped and why. "It returned 200" and "the model I chose answered" are different facts, and only the second one is worth an alert when it stops being true.

Effort is translated per provider, in one place. The profile says low; the adapter knows that this means a body field on Z.AI and a thinking level on Gemini, and that Gemma 4 takes only two levels.

A schema on Z.AI is rewritten inside the gateway, into json_object plus the schema in the prompt. Callers send a schema and get JSON, whichever model answers.

Why not just wait for the fixes

Because my need is narrow and theirs is not. I have three providers and a few hard rules. LiteLLM serves more than a hundred providers and has to stay compatible for all of them.

The two fixes I would need are in the queue. The pull request for JSON on Z.AI was opened on 2026-08-21, the one for effort on Z.AI on 2026-09-13, and neither is merged. For context, counted on 2026-10-06:

LiteLLM repositoryCount
Open issues1,796
Open pull requests3,397
Open pull requests older than three months574
Pull requests merged in the last 30 days2,272

That is a project merging about seventy-five pull requests a day and still receiving them faster. For a less common provider like Z.AI, I cannot plan around the day a fix lands.

Reporting upstream and running something small of your own are not opposites. I did both on the same day. If those fixes are merged, the next person will not hit what I hit.

What it costs

Three thousand lines that I maintain alone, and three providers. Every new provider is an adapter I write by hand. LiteLLM also installed 118 Python packages where mine installs none, which mattered to me for a process that holds every provider key, and will not matter to most people.

If I needed a fifth provider, or per-user keys and budgets, LiteLLM would be the right choice and I would move.

Takeaways

  • Put a profile between your code and the model names, with the effort inside it. This works on any gateway, including none.
  • Check which model actually answered. A 200 from the fallback is a quiet way to pay many times more.
  • When an existing tool does not fit, report what you found with a reproduction before you walk away.
This post is licensed under CC BY 4.0 by the author.