Count the places in your stack that call a model. The coding agent. The app that summarises tickets. The little script that cleans up a CSV. The background job that tags support mail. Each one has its own API key baked in, its own base URL, its own idea of which model to use. Each one bills to a different line you'll reconcile later, badly.

That sprawl is the actual problem. Not any single call — the fact that there's no single place where all of them happen.

So I built one. Everything I run — every project, every agent, every throwaway script — points at the same URL and lets a gateway in the middle do the deciding. It's called NautGate, and the pitch is boring on purpose: it speaks the OpenAI API, so nothing downstream has to change.

The one-line version

NautGate is an OpenAI-compatible proxy. Your code already knows how to talk to it — you just change the base URL:

https://api.openai.com/v1   →   http://localhost:8090/v1

Same request shape, same response shape, same SDKs. Point Claude Code, Codex, the OpenAI or Anthropic client, or your own fetch at :8090/v1, hand it a ng_… key, and it works. Nothing in your app learns that a gateway exists. That drop-in compatibility is the whole reason this is adoptable instead of aspirational — you don't rewrite anything to get every benefit below.

What moving to the middle buys you

Once every call goes through one door, things that were impossible when the logic lived in twelve scattered clients become one config change:

  • Cost, actually accounted. Every request and response is captured with its token counts, so spend is a real number per project, per model, per day — not a surprise at the end of the month across four vendor invoices. You see which job is expensive while it's being expensive.
  • Swap models without touching code. The model choice lives in the gateway, not hard-coded in each app. Move a workload from a frontier model to a cheaper one — or the other way — and every caller follows. Run a champion–challenger trial where a slice of traffic quietly hits a new model and you compare before committing.
  • Fallback that's already wired. When the upstream is down or rate-limited, the gateway can route the call to a local model instead — Ollama on your own box — so a provider outage degrades quality instead of taking you offline. Private inference for the calls that shouldn't leave, cloud for the ones where it doesn't matter.
  • Quality you can watch drift. Because the gateway sees the outputs, you can score them over time and treat a model like any other production dependency: track its behaviour, get alerted when it changes under you on a silent version bump.
The part that matters

None of these are features you bolt onto twelve clients. They're features of the middle. The moment there's one place every call passes through, cost accounting, model routing, fallback, and auditing stop being twelve small projects and become one small config. Centralising the boundary is the leverage — everything else is a consequence of it.

The receipts

There's a second reason I route everything through one place, and it's the one that matters more every month: evidence.

A captured request/response log is not just a debugging convenience. It's the answer to "which model made this decision, on what input, when?" — a question that stops being academic the moment you're doing anything a regulator, a client contract, or an incident review cares about. If you can't reconstruct what the model saw and said, you're taking that on faith. The gateway keeps the trail because it's the only component positioned to.

That governance angle is exactly where this connects to the compliance work I'll write about separately: you can't enforce "this class of data only goes to that class of model" if the routing decision is scattered across a dozen codebases. Put it in the gateway and the policy has one place to live — and one place to prove it was followed.

Isn't this just LiteLLM / a proxy?

Fair question, and partly yes — the drop-in-proxy idea isn't novel, and if a thin router is all you need, use one. NautGate leans harder on the observability and governance half: capturing every exchange for cost, quality, and audit as first-class, not as an afterthought bolted onto a router. I built my own because it's load-bearing across everything I ship and I wanted the analytics and policy layer to be the point, not a plugin. Your mileage depends on whether you care about the receipts as much as the routing. I do.

Where it fits

NautGate is the LLM boundary in a stack that's all about owning your boundaries — the same instinct as routing your agent's search through a box you run. Search you own so the query stream stays yours; a gateway you own so the inference stream stays yours — visible, swappable, accountable, and yours to keep the logs on.

It also isn't only mine to run. It's the per-node gateway inside the sandbox platform I'll cover next — one gateway per isolated agent, so even ephemeral, jurisdiction-pinned workers get the same single, accountable door. Same idea, one layer down.

The takeaway is smaller than the feature list makes it sound: stop letting every part of your system phone the model directly. Give them one number to call. Then all the things you wish you could see, you can — because they all happen in the same room.