Building / NautGate

NautGate

A memory-aware LLM gateway. Route Claude Code, Codex, Pi or your own app to any model — and get a provable record of which model really served each call, what it cost, and what left your machine.

Product demo

See NautGate in action.

NautGate walkthrough · 4:14 · EnglishOpen video ↗

01 / The build

Prove what the route did.

One gateway for every LLM call. Claude Code, Codex, Pi, Cline, OpenCode or your own app point at NautGate instead of at a provider, and it speaks the formats they already speak — OpenAI Chat, Anthropic Messages, OpenAI Responses.

What comes back is not just routing. For every call it keeps the requested model beside the model the provider reported as serving it, the route between them, the tokens, the timings, the cost and the policy decision — together, in one durable record. Plenty of gateways route; the difference is remembering enough to prove what the route did.

Four protocol surfaces accept the request — OpenAI Chat, OpenAI Responses, Anthropic Messages and a Pi adapter — and all of them enter the same routing and evidence pipeline. Protocol compatibility does not create four separate ledgers.

Why I built it

Ask a language model who it is and the answer reflects the harness around it. A client can label the assistant Claude, GPT or Gemini, and the model continues that fiction politely and confidently, because following the prompt is the job. The request is no better: routing layers can alias, override or substitute a model before a provider serves the call. Keep only the request and you have a record of intent, not a record of delivery. Once agents run on their own, that record is the only control surface left.

02 / Design decisions

The trade-offs are part of the story.

01

Read the response, not the request

The served model is parsed from the provider's own response and written beside the requested one. Where they differ, the routing-flow view highlights the exact hop where the substitution happened rather than burying it. Cost attribution then uses the model that actually produced the tokens, and quality comparisons group calls by the model that actually answered.

02

Observed evidence, not cryptographic proof

Provider-reported model identity is evidence, not proof. The record can show what NautGate observed and decided; it cannot independently prove which weights produced the tokens, and the receipt is an operational one rather than a signed attestation. That boundary is written into the product documentation on purpose, because a claim that quietly exceeds what the system can demonstrate is the failure mode this whole project exists to avoid. What it can do is let you check it without believing me: token counts and the served model line up against the provider's own billing dashboard, to the token, and a discrepancy stays visible instead of being normalised away.

03

Routing is a decision, not an invisible rewrite

A request is scored across task and payload dimensions, mapped to a tier, resolved through a routing table, and filtered by provider health — with explicit requests, key overrides and agent preferences applied in a defined order. Requested, selected and observed models are recorded separately, and a substitution flag goes back to the caller. Silent model switching is easy. Accountable switching is the rarer thing.

04

Separate durability contracts for decision and outcome

The decision is written before the upstream call, so it survives whatever happens next. The outcome — status, latency, tokens, cache tokens, cost, truncation, client disconnect — is written after, and if PostgreSQL cannot take it, it falls to a local SQLite WAL spool and is replayed on recovery rather than silently discarded. Many gateways emit logs. Fewer say out loud which parts are guaranteed and which are eventually durable.

05

Offline has to mean no egress

Pointing inference at localhost solves the largest data path and silences none of the rest. Provider heartbeats, health probes, catalogue refreshes and credit lookups all run on their own timers, so a 'local model' is not an offline system. Offline mode stands the whole gateway down together. In a 70-second socket sample it went from a connection to Anthropic every 60 seconds to 127.0.0.1 only — because configuration labels are not proof, and sampling the actual sockets is.

06

Classify before storing, not after

PII and secrets are detected before anything is written, and the sensitivity label feeds the routing decision as well as the capture policy. Regex and a second model can both miss context, so this reduces accidental retention rather than guaranteeing against data loss — but scanning after sensitive text is already in the database is the wrong place to put the check.

07

AGPL-3.0, deliberately

Copyleft keeps a hosted fork honest. It also means some enterprises cannot adopt it without a conversation, which is a cost I took on purpose rather than discovered late.