Skip to content

Architecture Article

Designing an AI Gateway

Follow one model request through identity, policy, routing, failover, streaming, cost accounting, caching, and observability.

Published 3 Sept 2026Updated 1 Oct 202614 min read
AI Gateway · LLM · Failover · Cost Metering · Streaming · Observability
One product request passing through authentication, policy, routing, metering, and one of several model providers.
On this page (15)

The first feature in a product that calls a language model usually calls the provider directly, and that is fine. By the fifth feature, the product has five retry policies, five ideas of which model to use, no single answer to "what did we spend yesterday, and on what", and no way to move traffic off a provider during an outage without changing five call sites. An AI gateway is the fix: one service, or one module, that every model call goes through, so that everything that should be true of every call is enforced in one place.

This chapter follows one illustrative request through a gateway, stage by stage, with the records each stage leaves behind. The design combines product experience with patterns larger companies have published; its records, identifiers, and implementation are illustrative.

How to read the diagrams

  • The gateway and its records: profiles, reservations, traces
  • Callers: product features and the people using them
  • Model providers, where every token costs money
  • Reuse: caches that avoid a provider call
  • Failure: provider errors, open circuits, refused requests

Read the numbered badges from the callers at the top to the records at the bottom. Steps 1–3 identify the caller and enforce policy before money is reserved. Steps 4–6 choose a cache or provider and adapt its stream. Step 7 settles the reservation and records the outcome.

Callers

Product features1

post rewrite · summarise · classify

Internal experiments

restricted pass-through

Gateway · ordered stages

Identity and idempotency2

who is calling · has this request run?

Policy3

use-case profile · safety · rate and spend limits

Routing4

model choice · breaker · retry · failover

Stream adapter6

provider events → one SSE contract

Provider path

Exact or semantic cache4

reuse only when the profile permits it

Provider A5

primary

Provider B5

failover

Records

Credit ledger37

reserve → settle or refund

Decision event7

profile · provider · usage · outcome

One controlled path from product features to model providers. The ledger and trace store observe the same request without storing its text.

A token is the provider's unit for pieces of input and output text; providers meter and rate-limit them separately. A rate limit caps requests or tokens in a short window, while a spend cap blocks more use after a money limit is reached. 429 is the HTTP status providers commonly use for both, so the response code alone cannot tell the gateway whether waiting will help. A retry repeats a failed call after a delay; backoff makes each delay longer, and a circuit breaker temporarily skips a provider that is already failing. Server-Sent Events, or SSE, is an HTTP stream in which the server sends named text events over one open response.

Why companies end up with one

The pattern shows up at almost every company that ships more than a couple of AI features, and the reasons they give are consistent.

Uber

GenAI Gateway

Why
Over 60 use cases were integrating with models in different ways
Shape
A Go service that mirrors the OpenAI API
Scale
About 30 teams, 16 million queries a month (2024)

Grab

AI Gateway

Why
Share provider quotas fairly across 300+ use cases
Shape
Reverse proxies to each provider, with its own rate limits on top
Scale
One gateway for every model call, billions of tokens a month (2026)

LinkedIn

GenAI Gateway

Why
Stop one feature starving the others
Shape
A common mid-tier in front of hosted models
Enforces
Outbound rate limits and per-feature quotas

Instacart

AI Gateway

Why
One abstraction over several providers
Shape
Real-time calls plus a batch path
Result
Batch jobs cost up to 50% less than real-time calls

Uber found that "the disparate integration strategies adopted by different teams have led to inefficiencies and redundant efforts", and built a gateway that also handles authentication, cost attribution, and audit logs. Grab enforces "its own rate limit on top of the global provider limits to make sure quotas are not consumed by a single service", and later described what the extra layer bought them: it lets a platform team "change which provider serves a model, configure fallback routing, set budgets, and manage cost attribution, without a single application touching its code" (Grab, 2026). LinkedIn's gateway exists partly "to prevent individual product features from overusing computing resources or starving other features."

The list of jobs is the same everywhere: one interface, one place to authenticate and authorise, one ledger of spend, one set of limits that sees total load, one failover policy, and one stream of telemetry. The rest of this chapter is about doing each of those correctly.

One request, end to end

Here is the request this chapter follows. A user of a social media product has written a post and clicks "Make it friendlier". The browser calls the gateway's streaming endpoint with the name of a use case, not a model. The ids, timings, token counts, credits, limits, and model names in this walkthrough are illustrative:

Sketch: request to one use-case route
POST /gateway/use-cases/{use_case}/stream
Authorization: Bearer eyJ…
Idempotency-Key: 5f1c7a2e-rewrite-18
Content-Type: application/json
 
{ "input": "We shipped the new scheduler. Update now.", "tone": "friendly" }

The gateway runs the same sequence of checks for every call. The order is not arbitrary, and several of the decisions below only make sense because of it:

post_rewrite · one request through the gateway

  1. 0 ms

    Authenticate, resolve the organisation

    Anonymous calls never reach a model. Every later stage is scoped to this organisation.

    org org_42 · user u_913

  2. 1 ms

    Replay or claim the idempotency key

    A concurrent retry joins or waits for the claimed call; a completed retry can replay a retained answer. The key prevents a second local start, but it cannot recover a provider response the gateway never recorded.

    key 5f1c7a2e-rewrite-18 · first time seen · claimed for 24 h

  3. 2 ms

    Scan the input for prompt injection and personal data

    Detection runs on the user's text only. Whether a high-risk result blocks the call is a policy switch.

    injection risk: none · PII: none

  4. 2 ms

    Check the organisation's rate limit

    org_42 · 3 of 60 this minute

  5. 3 ms

    Look up the use-case profile

    The profile holds the system prompt, the model tier, and the parameters. Only product-facing profiles are reachable over HTTP.

    post_rewrite · profile-v7

  6. 3 ms

    Choose the model

    no caller pin · profile tier primary-small

  7. 4 ms

    Reserve credits for the estimated cost

    The call cannot spend money the organisation doesn't have.

    estimate 2,044 tokens → reserve 3 credits

  8. 4 ms

    Check the cache

    Streaming calls skip the semantic cache in this design; the provider's own prompt cache still applies.

    bypassed: streamed response

  9. 5 ms

    Call the provider through the failover chain

    The first chunk commits the call to whichever provider produced it.

    openai · first token at 410 ms · 96 tokens streamed

  10. 1.9 s

    Settle on the model that actually answered

    The reservation is replaced by the real cost, and the difference is refunded.

    used 1,184 tokens → settle 2 credits, refund 1 → ledger row written

  11. 1.9 s

    Record the decision

    One structured line per call, with no prompt or output text in it.

    model_call.completed · outcome ok · requested = served

Two details in that order deserve attention.

The reservation happens before the cache lookup. The advantage is that cache hits and provider calls go through the same reserve-and-settle bookkeeping, and a cache hit settles at zero, which refunds the hold. The cost is less obvious: an organisation with no credits left is refused (or downgraded to a cheaper model) even when its request would have been served free from the cache. Putting the cache first avoids that. The miss path then reserves after the lookup, and the two paths start to diverge. Either choice is defensible. What matters is that someone chose it.

Model selection also happens before the reservation, because the price depends on the model. If the gateway later fails over to a different provider, the settlement uses the model that actually answered, not the one that was reserved against.

Use-case profiles, not model names

Callers ask for a use case. A use-case profile is the gateway's versioned recipe for one product action: its system instructions, allowed models, output shape, limits, and whether callers may reach it over HTTP. The gateway owns that recipe:

Sketch: a use-case profile and its version
export const postRewrite: UseCaseProfile = {
  key: "post_rewrite",
  exposure: "product",             // internal profiles are not HTTP-addressable
  systemPrompt: REWRITE_PROMPT,
  modelTier: "primary-small",
  creativity: "low",
  outputBudget: "standard",
  useCache: false,
};
 
// The version is a hash of everything that changes the output, so a prompt
// edit shows up in traces, caches, and evaluations without anyone bumping it.
export function profileVersion(p: UseCaseProfile): string {
  const material = JSON.stringify([
    p.key, p.systemPrompt, p.modelTier, p.creativity, p.outputBudget,
  ]);
  return "profile-" + sha256(material).slice(0, 8);
}

This moves every model decision into one file per feature. Changing the model for all "rewrite" traffic is a one-line change with one reviewer, and every trace carries the profile version that produced it.

There is a real trade-off against the other common design. Uber and Grab expose an OpenAI-compatible API, which means any team can adopt the gateway by changing a base URL, and every open-source client works unchanged. The price is that callers choose models, write prompts, and set parameters themselves, so the gateway sees traffic but owns few decisions. A profile-based API gives the platform team those decisions, at the cost of a registration step for every new use case. Most platforms end up with a mix: profiles for product features, and a controlled pass-through for internal experiments. Model pins are useful for evaluations and dangerous over public HTTP, so restrict them to authorised internal callers.

Unknown and internal-only profiles should return the same 404. If "exists but you may not call it" returns a different status from "doesn't exist", the endpoint becomes an oracle that lists your internal use cases.

Making the door the only door

A gateway only works if nothing goes around it. The easiest way around it is a feature that constructs its own vendor client because it needed something quickly, and that usually happens months after everyone has stopped thinking about the gateway.

The "one door" claim needs a test: can a feature reach a provider without crossing the gateway's authentication, policy, and meter? A source rule catches direct SDK construction during development, and an outbound network policy catches raw HTTP calls at runtime. Here is the source-rule shape:

Sketch: an invariant test that keeps model spend behind one door
const VENDOR_CLIENT =
  /\b(createOpenAI|createAnthropic|createGoogleGenerativeAI|new\s+OpenAI|new\s+Anthropic)\s*\(/;
 
it("constructs vendor clients only inside the gateway", () => {
  const offenders = applicationSourceFiles()      // comments and fixtures stripped
    .filter((file) => VENDOR_CLIENT.test(file.text))
    .filter((file) => file.owner !== "ai-gateway");
 
  expect(offenders).toEqual([]);
});

A test like this has blind spots. It catches client constructors but not a raw fetch to a provider's URL, so an outbound-host allowlist or network egress policy closes that gap. It should also assert that it actually scanned something: a refactor that moves the source directory will otherwise turn it into a test that passes on zero files. Every exception needs an owner and a written reason.

Failover by error class

Provider failover belongs in the normal request path. The gateway has to decide which errors deserve a retry, which deserve a different provider, and which should reach the caller unchanged:

error classifierdecides retry, fail over, or stop
error
transient
examples
connection failures, timeouts before a response, documented overloads, and retryable 5xx responses
action
retry within the remaining deadline, then fail over; retry 409 or other 4xx only when that provider says the specific operation is retryable
error
rate limited
examples
429 with a Retry-After header
action
honour Retry-After if it fits the budget; otherwise fail over
error
spend cap
examples
OpenAI spend-limit codes; Anthropic tier cap 429 or customer-set cap 400
action
use an authorised alternative or stop; waiting does not repair the configured limit
error
caller's fault
examples
other 4xx: invalid request, content policy, schema errors
action
stop and return the error; another provider won't fix the request
error
our fault
examples
401 or 403 from the provider: bad or revoked key
action
stop, alert, and don't count it against the provider's health
error
cancelled
examples
the caller aborted or disconnected
action
stop immediately; never retry

The classifier should inspect typed fields before text. These are the documented response fields checked in September 2026; the … stands for a provider-supplied message that must not drive the decision.

provider errors → gateway decision
provider response
OpenAI · 503
gateway decision
retry briefly, then fail over
typed fields to inspect
error.code = server_is_overloaded
provider response
OpenAI · 429
gateway decision
do not retry this project; use an authorised alternative or stop
typed fields to inspect
error.code = project_spend_limit_exceeded
provider response
Anthropic · 529
gateway decision
retry briefly, then fail over
typed fields to inspect
error.type = overloaded_error
provider response
Anthropic · 429 tier cap
gateway decision
do not retry until access resumes
typed fields to inspect
error.type = rate_limit_error · details.error_code = enforced_spend_limit_reached · no Retry-After
provider response
Anthropic · 400 customer-set cap
gateway decision
stop; the configured spend limit must be raised or reset
typed fields to inspect
error.type = invalid_request_error · message begins with the documented usage-limit text
provider response
Anthropic · 404
gateway decision
stop; routing or request is wrong
typed fields to inspect
error.type = not_found_error
Status narrows the class; a stable error code separates temporary throttling from a spend cap that waiting cannot fix.

Spend limits need two Anthropic branches. The usage tier's monthly cap returns 429 with details.error_code = enforced_spend_limit_reached and no retry-after. A lower limit set by the customer returns 400 invalid_request_error. Ordinary request and token rate limits return 429 with retry-after. The classifier must inspect the typed error fields before it decides that waiting will help.

The illustrative chain below tries OpenAI and then Anthropic. A real chain must include only providers authorised for that use case and data region. The first provider that can serve a pinned model gets it; later providers may answer with a different allowed model, which is a reminder that failover changes more than the vendor. A different model can produce a different format, a different length, and a different price, so structured-output schemas and evaluations have to hold across every model the chain can reach.

Here is a failover as it happens:

The primary is overloaded; the second provider answers

  1. 0 msGateway → OpenAI

    chat request, primary-small tier

  2. 2.1 sOpenAI → Gateway

    503 server_is_overloaded

  3. 2.1 sGateway → Breaker

    record a retryable failure

    4 in the last minute

  4. 2.3 sGateway → OpenAI

    retry after backoff with jitter

  5. 4.4 sOpenAI → Gateway

    503 again

  6. 4.4 sGateway → Breaker

    record a failure: threshold reached

    circuit opens; OpenAI skipped until the cooldown ends

  7. 4.4 sGateway → Anthropic

    fail over to the profile's allowed secondary model

  8. 5.2 sAnthropic → Gateway

    first token: the call is now committed to Anthropic

The request is settled at the price of the model that actually answered. The trace records both the requested and the served model, which is how failover shows up on a dashboard.

Each provider has its own circuit breaker. After a handful of retryable failures within a short window, the breaker opens and the provider is skipped entirely for a cooldown period. When the cooldown ends, one trial request is let through: if it succeeds the breaker closes, and if it fails the breaker opens again. Without breakers, every request during an outage pays the full timeout-and-retry cost against a provider that is known to be down.

Retries multiply

Retries are configured in more places than most teams realise. The Vercel AI SDK retries a failed call twice by default, and so do the official OpenAI and Anthropic SDKs. If the gateway's own loop makes two attempts per provider and each attempt goes through an SDK that retries twice more, one logical call can become six HTTP requests to a provider that is already struggling, before failover even starts:

HTTP requests to one provider for a single failing call

Gateway retries only

Two requests, each with its own timeout: a predictable worst case.

Gateway retries × SDK retries

Six requests, each with backoff: the outage gets more traffic, and the caller waits much longer to fail over.

y: requests before failover

Pick one layer to own retries (the gateway, which knows about failover and budgets) and set the SDK's retry count to zero. OpenAI's rate-limit guide makes a related point: "unsuccessful requests contribute to your per-minute limit, so continuously resending a request won't work."

Once a stream starts, it belongs to that provider

Failover only works before the first token. After the caller has seen half a sentence from one model, the gateway cannot switch to another without the user seeing a restart. So the first chunk commits the call: from then on, an error ends the stream with an error event, and the question becomes what to bill, which is covered below.

Paying for it: reserve, then settle

Metering has to be correct when calls fail, not only when they succeed. The pattern is two steps: reserve the estimated cost before the call, then replace the reservation with the real cost afterwards.

credit ledger · org_42reserve before the call, settle after
momentmodeltokenscreditsbalance used
before the call
not chosennot chosennot startednot charged418 / 1,000
reserve: input ≈ 1,020 tokens (system prompt and post) + max output 1,024
reserveprimary-small≈2,044 (estimate)3421 / 1,000
settle: the provider reported real usage, 1 credit returned
settleprimary-small1,1842420 / 1,000
another call fails before any spend: settle to zero
settleclaude-sonnet00 (hold refunded)420 / 1,000
Illustrative values. Production reservation should use the provider's tokenizer or token-count endpoint plus the maximum output. A character estimate can undercount code and many languages, so it needs conservative headroom and is not a hard spend guarantee.

Now take the failure branch the normal timeline did not take. The provider disconnects after some text has reached the browser. The model work already happened, so a policy that bills delivered usage settles part of the reservation instead of refunding all of it.

credit ledger · partial stream branch
momentdelivered outputledger statebalance used
before the call
ready0 tokensnot charged418 / 1,000
reserve the illustrative worst-case estimate
provider not called0 tokens3 credits held421 / 1,000
stream fails after the browser received 96 output tokens
provider connection lost96 tokens3 credits still held421 / 1,000
settle the measured delivered usage under this example policy
closed96 tokenscharge 1 · refund 2419 / 1,000
Illustrative policy and values. Another product may refund partial streams; the gateway must make the rule explicit and apply it consistently.

The reservation and the balance check must be one atomic operation, or two concurrent calls can both see enough balance and both spend it. When the balance can't cover the estimate, there are two sensible responses: downgrade to a cheaper model, or refuse with 402 Payment Required. That choice belongs to the use case, not the plan. A post rewrite can quietly use a smaller model; a contract summary probably should stop and say so.

Keep provider cost and customer credits as two related ledgers. The provider cost record settles from the provider's reported usage or later billing data, including work generated but not delivered to the browser. The customer-credit record follows the product's published policy and may refund a partial stream. Using “tokens delivered to the browser” as the only cost record hides spend the provider still charged.

Several things break this ledger in practice:

  • A worker that crashes between reserve and settle leaves the reservation in place. Give reservations an expiry and a sweeper that settles them, or a crash becomes a permanent charge.
  • A stream that fails after tokens were sent has already cost something. Decide explicitly whether the customer pays for the tokens they received or nothing at all, write the rule down, and test it. Refunding everything is generous; it also means a flaky provider costs you money on every partial stream.
  • Side calls are easy to forget: a judge model scoring an answer, an embedding for a cache lookup, a moderation check. If they don't go through the same meter, the ledger undercounts real spend.
  • When several identical requests are coalesced into one provider call, decide who pays. Charging every caller the full cost of a call that happened once is a quiet overcharge.

Grab and LinkedIn both put quotas in the gateway for a different reason: fairness. Grab found that "batch usage interfered with the uptime of online services", and moved batch traffic to asynchronous APIs. Instacart routes large jobs to provider batch endpoints and reports savings of "up to 50% on LLM costs compared to standard real-time calls". A deployment-wide daily ceiling on spend is a useful last line of defence too: a counter that refuses new calls once the day's total crosses a limit turns a runaway loop into an alert instead of an invoice.

Streaming

Streaming changes the error model. Once the gateway has written 200 OK and the text/event-stream headers, every later failure has to be expressed inside the stream. So the gateway should run every check that can refuse the request (limits, balance, profile, the first provider round-trip) before it writes the headers. Then a 402 or 429 arrives as a real status code, which clients and proxies understand, instead of as an error event inside a successful response.

Sketch: the event stream for the post_rewrite call
HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache, no-transform
 
data: {"delta":"We shipped"}
 
data: {"delta":" the new scheduler! "}
 
event: done
data: {"useCase":"post_rewrite","profileVersion":"profile-v7","modelTier":"primary-small","provider":"openai","totalTokens":1184,"credits":2}

Three more details matter in production. When the client disconnects, request cancellation of the upstream call; this reduces wasted generation but does not prove the provider stopped or will not bill it. A proxy or load balancer may close an idle connection, and a model that thinks before its first token can look idle, so configure the actual timeout and send SSE comment lines (lines starting with a colon) at a shorter interval. And the final usage figures arrive after the last token, so give settlement a short deadline and a fallback: hold the conservative reservation, flag the record, and reconcile with provider usage data rather than treating the estimate as measured cost.

Caching: exact prefixes first

Two different caches get called "caching" in AI systems, and they solve different problems.

Provider prompt caching

What matches
An identical prefix: tools, system prompt, then earlier messages
What it saves
Input work; read and write multipliers depend on provider, model, and cache duration
Risk
The model still generates a fresh answer, but stale shared prefixes and workspace scope still need review
Scope
Provider workspace or project boundary; verify the provider's current rule

Gateway semantic cache

What matches
A different request whose embedding is close enough
What it saves
The whole call, output included
Risk
Serving the answer to a different question
Scope
Whatever your cache key says, so it must include the tenant

Provider prompt caching is usually the first cache to consider because it does not reuse an old answer: the model still generates a fresh response from an identical prefix. It still needs a privacy review of the provider's cache scope, retention, and workspace boundaries. Anthropic requires "100% identical prompt segments" in the order tools, system, messages; OpenAI enables it by default above a minimum prompt length and requires "the entire rendered prefix to match". The design consequence is ordering: put stable content (tool definitions, the system prompt, long reference documents) first, and anything that varies per request (timestamps, user names, the user's input) last. A timestamp at the top of a system prompt disables the cache for every call. There is also a capacity effect: for most Claude models, cached input tokens don't count toward the input-tokens-per-minute limit, so a high hit rate raises the effective rate limit.

A semantic cache is a much sharper tool, with real risks for tenant isolation and for agent traffic. It has its own chapter: Semantic caching for LLM calls.

Privacy, safety, and what gets logged

A gateway sees every prompt, which makes it the right place for privacy controls and the wrong place for careless logging.

Uber's gateway redacts personal data before a prompt leaves the company and restores it in the answer, replacing names and other identifiers with stand-in tokens such as ANONYMIZED_NAME_0. Their post is candid that the redactor "introduces challenges to both latency and result quality": a model that sees stand-in tokens instead of names writes a slightly worse answer. Redaction is worth it for some use cases and not others, which is another argument for making it a profile setting rather than a global switch.

Prompt-injection scanning belongs in the gateway for the same reason, with a separate switch for detecting and for blocking, so a new detector can run and be measured before it is allowed to refuse anyone.

Telemetry should record usage, not content. Model-call spans can omit inputs and outputs while the gateway's decision log carries identifiers, counts, and outcomes. If an evaluation needs content, use a separately authorised sample with retention and access controls instead of turning prompt logging on for every call.

Seeing every decision

Every call produces one structured line. This is what lets you answer "what did we spend on rewrites last week, on which models, and how often did we fail over" with a query instead of an investigation:

Sketch: one decision record
{
  "event": "model_call.completed",
  "useCase": "post_rewrite",
  "profileVersion": "profile-v7",
  "outcome": "ok",
  "streamed": true,
  "requestedTier": "primary-small",
  "servedTier": "primary-small",
  "provider": "openai",
  "cacheHit": false,
  "inputTokens": 1088,
  "outputTokens": 96,
  "chargedCredits": 2,
  "downgradedTo": null,
  "injectionRisk": "none",
  "piiDetected": false,
  "orgId": "org_42",
  "durationMs": 1912
}

Metrics are built from the same events, with one rule: the organisation id goes in the log line, never in a metric label. Labels with thousands of distinct values make a time-series database slow and expensive; use case, model, provider, and outcome are enough dimensions for dashboards, and the logs answer per-customer questions.

Quality needs watching too, because a model that starts giving worse answers doesn't throw an error. An LLM judge can score an authorised sample against a rubric, or check each answer inline for use cases where a bad answer is expensive: generate one or more candidates, score them, keep the best, and allow one extra attempt if nothing passes, flagging the result as low-confidence rather than looping. Every judge call is a model call, so it goes through the same meter.

Build it or buy it

Managed and open-source gateways now cover a lot of this, and their documentation is worth reading even if you build your own, because their defaults show where the sharp edges are.

documented behaviour of off-the-shelf gateways
gateway
Cloudflare AI Gateway
caching
Exact-match only; semantic caching is on the roadmap
retries and failover
Up to 5 retries, at most 60 s delay; fallback on any error by default
operational detail
Cache is 'volatile': two identical concurrent requests can both miss
gateway
LiteLLM
caching
Exact and semantic
retries and failover
Cooldown after 3 failures; immediate 5 s cooldown on 429
operational detail
Its docs warn semantic caching 'goes badly wrong on agentic traffic'
gateway
Portkey
caching
Exact and semantic (default threshold 0.95)
retries and failover
Up to 5 attempts on 429, 500, 502, 503, 504, 529
operational detail
Semantic matching ignores the system prompt
gateway
Kong AI Proxy
caching
Exact and semantic plugins
retries and failover
Client errors don't fail over unless configured
operational detail
Semantic plugin can answer near-matches even with exact caching on
From each project's documentation as of September 2026. Defaults change; check the current docs before relying on one.

Buying makes sense when the main need is a single interface, provider failover, and basic spend visibility. Building makes sense when the gateway has to understand your product: per-customer credits and plans, use-case profiles owned by feature teams, tenant-scoped caches, and a ledger that billing can trust. Many teams end up doing both: a managed gateway for provider routing and a thin layer of their own for metering and policy.

A checklist

The door

  • Every model call goes through the gateway, enforced by a test or an egress policy
  • Callers name a use case; model pins are restricted to internal callers
  • Unknown and internal-only use cases return the same 404

Order of checks

  • Auth and idempotency before anything that costs money
  • Model chosen before the reservation, settled on the model that answered
  • Refusals happen before streaming headers are written

Failover

  • Errors classified into retry, fail over, and stop
  • Spend-cap 429s are not treated as transient
  • One layer owns retries; SDK retries set to zero
  • A circuit breaker per provider
  • Schemas and evals hold across every model in the chain

Money

  • Reserve and balance check are atomic
  • Reservations expire and get swept
  • A rule for billing partial streams, tested
  • Judge, embedding, and moderation calls are metered
  • Quotas per feature and a daily ceiling per deployment

Caching

  • Stable prompt content first, variable content last
  • Semantic caching only where scoped per tenant and measured

Privacy and visibility

  • Redaction and injection checks as per-use-case settings
  • No prompt or output text in logs or spans by default
  • One decision record per call; no tenant ids in metric labels
  • Quality sampled by a judge, and the judge metered

Sources