The first feature in a product that calls a language model usually calls the provider directly, and that is fine. By the fifth feature, the product has five retry policies, five ideas of which model to use, no single answer to "what did we spend yesterday, and on what", and no way to move traffic off a provider during an outage without changing five call sites. An AI gateway is the fix: one service, or one module, that every model call goes through, so that everything that should be true of every call is enforced in one place.
This chapter follows one illustrative request through a gateway, stage by stage, with the records each stage leaves behind. The design combines product experience with patterns larger companies have published; its records, identifiers, and implementation are illustrative.
How to read the diagrams
- The gateway and its records: profiles, reservations, traces
- Callers: product features and the people using them
- Model providers, where every token costs money
- Reuse: caches that avoid a provider call
- Failure: provider errors, open circuits, refused requests
Read the numbered badges from the callers at the top to the records at the bottom. Steps 1–3 identify the caller and enforce policy before money is reserved. Steps 4–6 choose a cache or provider and adapt its stream. Step 7 settles the reservation and records the outcome.
Callers
Product features1
post rewrite · summarise · classify
Internal experiments
restricted pass-through
Gateway · ordered stages
Identity and idempotency2
who is calling · has this request run?
Policy3
use-case profile · safety · rate and spend limits
Routing4
model choice · breaker · retry · failover
Stream adapter6
provider events → one SSE contract
Provider path
Exact or semantic cache4
reuse only when the profile permits it
Provider A5
primary
Provider B5
failover
Records
Credit ledger37
reserve → settle or refund
Decision event7
profile · provider · usage · outcome
A token is the provider's unit for pieces of input and output text; providers meter and rate-limit them separately. A rate limit caps requests or tokens in a short window, while a spend cap blocks more use after a money limit is reached. 429 is the HTTP status providers commonly use for both, so the response code alone cannot tell the gateway whether waiting will help. A retry repeats a failed call after a delay; backoff makes each delay longer, and a circuit breaker temporarily skips a provider that is already failing. Server-Sent Events, or SSE, is an HTTP stream in which the server sends named text events over one open response.
Why companies end up with one
The pattern shows up at almost every company that ships more than a couple of AI features, and the reasons they give are consistent.
Uber
GenAI Gateway
- Why
- Over 60 use cases were integrating with models in different ways
- Shape
- A Go service that mirrors the OpenAI API
- Scale
- About 30 teams, 16 million queries a month (2024)
Grab
AI Gateway
- Why
- Share provider quotas fairly across 300+ use cases
- Shape
- Reverse proxies to each provider, with its own rate limits on top
- Scale
- One gateway for every model call, billions of tokens a month (2026)
GenAI Gateway
- Why
- Stop one feature starving the others
- Shape
- A common mid-tier in front of hosted models
- Enforces
- Outbound rate limits and per-feature quotas
Instacart
AI Gateway
- Why
- One abstraction over several providers
- Shape
- Real-time calls plus a batch path
- Result
- Batch jobs cost up to 50% less than real-time calls
Uber found that "the disparate integration strategies adopted by different teams have led to inefficiencies and redundant efforts", and built a gateway that also handles authentication, cost attribution, and audit logs. Grab enforces "its own rate limit on top of the global provider limits to make sure quotas are not consumed by a single service", and later described what the extra layer bought them: it lets a platform team "change which provider serves a model, configure fallback routing, set budgets, and manage cost attribution, without a single application touching its code" (Grab, 2026). LinkedIn's gateway exists partly "to prevent individual product features from overusing computing resources or starving other features."
The list of jobs is the same everywhere: one interface, one place to authenticate and authorise, one ledger of spend, one set of limits that sees total load, one failover policy, and one stream of telemetry. The rest of this chapter is about doing each of those correctly.
One request, end to end
Here is the request this chapter follows. A user of a social media product has written a post and clicks "Make it friendlier". The browser calls the gateway's streaming endpoint with the name of a use case, not a model. The ids, timings, token counts, credits, limits, and model names in this walkthrough are illustrative:
POST /gateway/use-cases/{use_case}/stream
Authorization: Bearer eyJ…
Idempotency-Key: 5f1c7a2e-rewrite-18
Content-Type: application/json
{ "input": "We shipped the new scheduler. Update now.", "tone": "friendly" }The gateway runs the same sequence of checks for every call. The order is not arbitrary, and several of the decisions below only make sense because of it:
post_rewrite · one request through the gateway
- 0 ms
Authenticate, resolve the organisation
Anonymous calls never reach a model. Every later stage is scoped to this organisation.
org org_42 · user u_913
- 1 ms
Replay or claim the idempotency key
A concurrent retry joins or waits for the claimed call; a completed retry can replay a retained answer. The key prevents a second local start, but it cannot recover a provider response the gateway never recorded.
key 5f1c7a2e-rewrite-18 · first time seen · claimed for 24 h
- 2 ms
Scan the input for prompt injection and personal data
Detection runs on the user's text only. Whether a high-risk result blocks the call is a policy switch.
injection risk: none · PII: none
- 2 ms
Check the organisation's rate limit
org_42 · 3 of 60 this minute
- 3 ms
Look up the use-case profile
The profile holds the system prompt, the model tier, and the parameters. Only product-facing profiles are reachable over HTTP.
post_rewrite · profile-v7
- 3 ms
Choose the model
no caller pin · profile tier primary-small
- 4 ms
Reserve credits for the estimated cost
The call cannot spend money the organisation doesn't have.
estimate 2,044 tokens → reserve 3 credits
- 4 ms
Check the cache
Streaming calls skip the semantic cache in this design; the provider's own prompt cache still applies.
bypassed: streamed response
- 5 ms
Call the provider through the failover chain
The first chunk commits the call to whichever provider produced it.
openai · first token at 410 ms · 96 tokens streamed
- 1.9 s
Settle on the model that actually answered
The reservation is replaced by the real cost, and the difference is refunded.
used 1,184 tokens → settle 2 credits, refund 1 → ledger row written
- 1.9 s
Record the decision
One structured line per call, with no prompt or output text in it.
model_call.completed · outcome ok · requested = served
Two details in that order deserve attention.
The reservation happens before the cache lookup. The advantage is that cache hits and provider calls go through the same reserve-and-settle bookkeeping, and a cache hit settles at zero, which refunds the hold. The cost is less obvious: an organisation with no credits left is refused (or downgraded to a cheaper model) even when its request would have been served free from the cache. Putting the cache first avoids that. The miss path then reserves after the lookup, and the two paths start to diverge. Either choice is defensible. What matters is that someone chose it.
Model selection also happens before the reservation, because the price depends on the model. If the gateway later fails over to a different provider, the settlement uses the model that actually answered, not the one that was reserved against.
Use-case profiles, not model names
Callers ask for a use case. A use-case profile is the gateway's versioned recipe for one product action: its system instructions, allowed models, output shape, limits, and whether callers may reach it over HTTP. The gateway owns that recipe:
export const postRewrite: UseCaseProfile = {
key: "post_rewrite",
exposure: "product", // internal profiles are not HTTP-addressable
systemPrompt: REWRITE_PROMPT,
modelTier: "primary-small",
creativity: "low",
outputBudget: "standard",
useCache: false,
};
// The version is a hash of everything that changes the output, so a prompt
// edit shows up in traces, caches, and evaluations without anyone bumping it.
export function profileVersion(p: UseCaseProfile): string {
const material = JSON.stringify([
p.key, p.systemPrompt, p.modelTier, p.creativity, p.outputBudget,
]);
return "profile-" + sha256(material).slice(0, 8);
}This moves every model decision into one file per feature. Changing the model for all "rewrite" traffic is a one-line change with one reviewer, and every trace carries the profile version that produced it.
There is a real trade-off against the other common design. Uber and Grab expose an OpenAI-compatible API, which means any team can adopt the gateway by changing a base URL, and every open-source client works unchanged. The price is that callers choose models, write prompts, and set parameters themselves, so the gateway sees traffic but owns few decisions. A profile-based API gives the platform team those decisions, at the cost of a registration step for every new use case. Most platforms end up with a mix: profiles for product features, and a controlled pass-through for internal experiments. Model pins are useful for evaluations and dangerous over public HTTP, so restrict them to authorised internal callers.
Unknown and internal-only profiles should return the same 404. If "exists but you may not call it" returns a different status from "doesn't exist", the endpoint becomes an oracle that lists your internal use cases.
Making the door the only door
A gateway only works if nothing goes around it. The easiest way around it is a feature that constructs its own vendor client because it needed something quickly, and that usually happens months after everyone has stopped thinking about the gateway.
The "one door" claim needs a test: can a feature reach a provider without crossing the gateway's authentication, policy, and meter? A source rule catches direct SDK construction during development, and an outbound network policy catches raw HTTP calls at runtime. Here is the source-rule shape:
const VENDOR_CLIENT =
/\b(createOpenAI|createAnthropic|createGoogleGenerativeAI|new\s+OpenAI|new\s+Anthropic)\s*\(/;
it("constructs vendor clients only inside the gateway", () => {
const offenders = applicationSourceFiles() // comments and fixtures stripped
.filter((file) => VENDOR_CLIENT.test(file.text))
.filter((file) => file.owner !== "ai-gateway");
expect(offenders).toEqual([]);
});A test like this has blind spots. It catches client constructors but not a raw fetch to a provider's URL, so an outbound-host allowlist or network egress policy closes that gap. It should also assert that it actually scanned something: a refactor that moves the source directory will otherwise turn it into a test that passes on zero files. Every exception needs an owner and a written reason.
Failover by error class
Provider failover belongs in the normal request path. The gateway has to decide which errors deserve a retry, which deserve a different provider, and which should reach the caller unchanged:
- error
- transient
- examples
- connection failures, timeouts before a response, documented overloads, and retryable 5xx responses
- action
- retry within the remaining deadline, then fail over; retry 409 or other 4xx only when that provider says the specific operation is retryable
- error
- rate limited
- examples
- 429 with a Retry-After header
- action
- honour Retry-After if it fits the budget; otherwise fail over
- error
- spend cap
- examples
- OpenAI spend-limit codes; Anthropic tier cap 429 or customer-set cap 400
- action
- use an authorised alternative or stop; waiting does not repair the configured limit
- error
- caller's fault
- examples
- other 4xx: invalid request, content policy, schema errors
- action
- stop and return the error; another provider won't fix the request
- error
- our fault
- examples
- 401 or 403 from the provider: bad or revoked key
- action
- stop, alert, and don't count it against the provider's health
- error
- cancelled
- examples
- the caller aborted or disconnected
- action
- stop immediately; never retry
The classifier should inspect typed fields before text. These are the documented response fields checked in September 2026; the … stands for a provider-supplied message that must not drive the decision.
- provider response
- OpenAI · 503
- gateway decision
- retry briefly, then fail over
- typed fields to inspect
- error.code = server_is_overloaded
- provider response
- OpenAI · 429
- gateway decision
- do not retry this project; use an authorised alternative or stop
- typed fields to inspect
- error.code = project_spend_limit_exceeded
- provider response
- Anthropic · 529
- gateway decision
- retry briefly, then fail over
- typed fields to inspect
- error.type = overloaded_error
- provider response
- Anthropic · 429 tier cap
- gateway decision
- do not retry until access resumes
- typed fields to inspect
- error.type = rate_limit_error · details.error_code = enforced_spend_limit_reached · no Retry-After
- provider response
- Anthropic · 400 customer-set cap
- gateway decision
- stop; the configured spend limit must be raised or reset
- typed fields to inspect
- error.type = invalid_request_error · message begins with the documented usage-limit text
- provider response
- Anthropic · 404
- gateway decision
- stop; routing or request is wrong
- typed fields to inspect
- error.type = not_found_error
Spend limits need two Anthropic branches. The usage tier's monthly cap returns
429 with details.error_code = enforced_spend_limit_reached and no
retry-after. A lower limit set by the customer returns 400
invalid_request_error. Ordinary request and token rate limits return 429 with
retry-after. The classifier must inspect the typed error fields before it
decides that waiting will help.
The illustrative chain below tries OpenAI and then Anthropic. A real chain must include only providers authorised for that use case and data region. The first provider that can serve a pinned model gets it; later providers may answer with a different allowed model, which is a reminder that failover changes more than the vendor. A different model can produce a different format, a different length, and a different price, so structured-output schemas and evaluations have to hold across every model the chain can reach.
Here is a failover as it happens:
The primary is overloaded; the second provider answers
- 0 msGateway → OpenAI
chat request, primary-small tier
- 2.1 sOpenAI → Gateway
503 server_is_overloaded
- 2.1 sGateway → Breaker
record a retryable failure
4 in the last minute
- 2.3 sGateway → OpenAI
retry after backoff with jitter
- 4.4 sOpenAI → Gateway
503 again
- 4.4 sGateway → Breaker
record a failure: threshold reached
circuit opens; OpenAI skipped until the cooldown ends
- 4.4 sGateway → Anthropic
fail over to the profile's allowed secondary model
- 5.2 sAnthropic → Gateway
first token: the call is now committed to Anthropic
Each provider has its own circuit breaker. After a handful of retryable failures within a short window, the breaker opens and the provider is skipped entirely for a cooldown period. When the cooldown ends, one trial request is let through: if it succeeds the breaker closes, and if it fails the breaker opens again. Without breakers, every request during an outage pays the full timeout-and-retry cost against a provider that is known to be down.
Retries multiply
Retries are configured in more places than most teams realise. The Vercel AI SDK retries a failed call twice by default, and so do the official OpenAI and Anthropic SDKs. If the gateway's own loop makes two attempts per provider and each attempt goes through an SDK that retries twice more, one logical call can become six HTTP requests to a provider that is already struggling, before failover even starts:
HTTP requests to one provider for a single failing call
Gateway retries only
Two requests, each with its own timeout: a predictable worst case.
Gateway retries × SDK retries
Six requests, each with backoff: the outage gets more traffic, and the caller waits much longer to fail over.
y: requests before failover
Pick one layer to own retries (the gateway, which knows about failover and budgets) and set the SDK's retry count to zero. OpenAI's rate-limit guide makes a related point: "unsuccessful requests contribute to your per-minute limit, so continuously resending a request won't work."
Once a stream starts, it belongs to that provider
Failover only works before the first token. After the caller has seen half a sentence from one model, the gateway cannot switch to another without the user seeing a restart. So the first chunk commits the call: from then on, an error ends the stream with an error event, and the question becomes what to bill, which is covered below.
Paying for it: reserve, then settle
Metering has to be correct when calls fail, not only when they succeed. The pattern is two steps: reserve the estimated cost before the call, then replace the reservation with the real cost afterwards.
| moment | model | tokens | credits | balance used |
|---|---|---|---|---|
| before the call | ||||
| not chosen | not chosen | not started | not charged | 418 / 1,000 |
| reserve: input ≈ 1,020 tokens (system prompt and post) + max output 1,024 | ||||
| reserve | primary-small | ≈2,044 (estimate) | 3 | 421 / 1,000 |
| settle: the provider reported real usage, 1 credit returned | ||||
| settle | primary-small | 1,184 | 2 | 420 / 1,000 |
| another call fails before any spend: settle to zero | ||||
| settle | claude-sonnet | 0 | 0 (hold refunded) | 420 / 1,000 |
Now take the failure branch the normal timeline did not take. The provider disconnects after some text has reached the browser. The model work already happened, so a policy that bills delivered usage settles part of the reservation instead of refunding all of it.
| moment | delivered output | ledger state | balance used |
|---|---|---|---|
| before the call | |||
| ready | 0 tokens | not charged | 418 / 1,000 |
| reserve the illustrative worst-case estimate | |||
| provider not called | 0 tokens | 3 credits held | 421 / 1,000 |
| stream fails after the browser received 96 output tokens | |||
| provider connection lost | 96 tokens | 3 credits still held | 421 / 1,000 |
| settle the measured delivered usage under this example policy | |||
| closed | 96 tokens | charge 1 · refund 2 | 419 / 1,000 |
The reservation and the balance check must be one atomic operation, or two concurrent calls can both see enough balance and both spend it. When the balance can't cover the estimate, there are two sensible responses: downgrade to a cheaper model, or refuse with 402 Payment Required. That choice belongs to the use case, not the plan. A post rewrite can quietly use a smaller model; a contract summary probably should stop and say so.
Keep provider cost and customer credits as two related ledgers. The provider cost record settles from the provider's reported usage or later billing data, including work generated but not delivered to the browser. The customer-credit record follows the product's published policy and may refund a partial stream. Using “tokens delivered to the browser” as the only cost record hides spend the provider still charged.
Several things break this ledger in practice:
- A worker that crashes between reserve and settle leaves the reservation in place. Give reservations an expiry and a sweeper that settles them, or a crash becomes a permanent charge.
- A stream that fails after tokens were sent has already cost something. Decide explicitly whether the customer pays for the tokens they received or nothing at all, write the rule down, and test it. Refunding everything is generous; it also means a flaky provider costs you money on every partial stream.
- Side calls are easy to forget: a judge model scoring an answer, an embedding for a cache lookup, a moderation check. If they don't go through the same meter, the ledger undercounts real spend.
- When several identical requests are coalesced into one provider call, decide who pays. Charging every caller the full cost of a call that happened once is a quiet overcharge.
Grab and LinkedIn both put quotas in the gateway for a different reason: fairness. Grab found that "batch usage interfered with the uptime of online services", and moved batch traffic to asynchronous APIs. Instacart routes large jobs to provider batch endpoints and reports savings of "up to 50% on LLM costs compared to standard real-time calls". A deployment-wide daily ceiling on spend is a useful last line of defence too: a counter that refuses new calls once the day's total crosses a limit turns a runaway loop into an alert instead of an invoice.
Streaming
Streaming changes the error model. Once the gateway has written 200 OK and the text/event-stream headers, every later failure has to be expressed inside the stream. So the gateway should run every check that can refuse the request (limits, balance, profile, the first provider round-trip) before it writes the headers. Then a 402 or 429 arrives as a real status code, which clients and proxies understand, instead of as an error event inside a successful response.
HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache, no-transform
data: {"delta":"We shipped"}
data: {"delta":" the new scheduler! "}
event: done
data: {"useCase":"post_rewrite","profileVersion":"profile-v7","modelTier":"primary-small","provider":"openai","totalTokens":1184,"credits":2}Three more details matter in production. When the client disconnects, request cancellation of the upstream call; this reduces wasted generation but does not prove the provider stopped or will not bill it. A proxy or load balancer may close an idle connection, and a model that thinks before its first token can look idle, so configure the actual timeout and send SSE comment lines (lines starting with a colon) at a shorter interval. And the final usage figures arrive after the last token, so give settlement a short deadline and a fallback: hold the conservative reservation, flag the record, and reconcile with provider usage data rather than treating the estimate as measured cost.
Caching: exact prefixes first
Two different caches get called "caching" in AI systems, and they solve different problems.
Provider prompt caching
- What matches
- An identical prefix: tools, system prompt, then earlier messages
- What it saves
- Input work; read and write multipliers depend on provider, model, and cache duration
- Risk
- The model still generates a fresh answer, but stale shared prefixes and workspace scope still need review
- Scope
- Provider workspace or project boundary; verify the provider's current rule
Gateway semantic cache
- What matches
- A different request whose embedding is close enough
- What it saves
- The whole call, output included
- Risk
- Serving the answer to a different question
- Scope
- Whatever your cache key says, so it must include the tenant
Provider prompt caching is usually the first cache to consider because it does not reuse an old answer: the model still generates a fresh response from an identical prefix. It still needs a privacy review of the provider's cache scope, retention, and workspace boundaries. Anthropic requires "100% identical prompt segments" in the order tools, system, messages; OpenAI enables it by default above a minimum prompt length and requires "the entire rendered prefix to match". The design consequence is ordering: put stable content (tool definitions, the system prompt, long reference documents) first, and anything that varies per request (timestamps, user names, the user's input) last. A timestamp at the top of a system prompt disables the cache for every call. There is also a capacity effect: for most Claude models, cached input tokens don't count toward the input-tokens-per-minute limit, so a high hit rate raises the effective rate limit.
A semantic cache is a much sharper tool, with real risks for tenant isolation and for agent traffic. It has its own chapter: Semantic caching for LLM calls.
Privacy, safety, and what gets logged
A gateway sees every prompt, which makes it the right place for privacy controls and the wrong place for careless logging.
Uber's gateway redacts personal data before a prompt leaves the company and restores it in the answer, replacing names and other identifiers with stand-in tokens such as ANONYMIZED_NAME_0. Their post is candid that the redactor "introduces challenges to both latency and result quality": a model that sees stand-in tokens instead of names writes a slightly worse answer. Redaction is worth it for some use cases and not others, which is another argument for making it a profile setting rather than a global switch.
Prompt-injection scanning belongs in the gateway for the same reason, with a separate switch for detecting and for blocking, so a new detector can run and be measured before it is allowed to refuse anyone.
Telemetry should record usage, not content. Model-call spans can omit inputs and outputs while the gateway's decision log carries identifiers, counts, and outcomes. If an evaluation needs content, use a separately authorised sample with retention and access controls instead of turning prompt logging on for every call.
Seeing every decision
Every call produces one structured line. This is what lets you answer "what did we spend on rewrites last week, on which models, and how often did we fail over" with a query instead of an investigation:
{
"event": "model_call.completed",
"useCase": "post_rewrite",
"profileVersion": "profile-v7",
"outcome": "ok",
"streamed": true,
"requestedTier": "primary-small",
"servedTier": "primary-small",
"provider": "openai",
"cacheHit": false,
"inputTokens": 1088,
"outputTokens": 96,
"chargedCredits": 2,
"downgradedTo": null,
"injectionRisk": "none",
"piiDetected": false,
"orgId": "org_42",
"durationMs": 1912
}Metrics are built from the same events, with one rule: the organisation id goes in the log line, never in a metric label. Labels with thousands of distinct values make a time-series database slow and expensive; use case, model, provider, and outcome are enough dimensions for dashboards, and the logs answer per-customer questions.
Quality needs watching too, because a model that starts giving worse answers doesn't throw an error. An LLM judge can score an authorised sample against a rubric, or check each answer inline for use cases where a bad answer is expensive: generate one or more candidates, score them, keep the best, and allow one extra attempt if nothing passes, flagging the result as low-confidence rather than looping. Every judge call is a model call, so it goes through the same meter.
Build it or buy it
Managed and open-source gateways now cover a lot of this, and their documentation is worth reading even if you build your own, because their defaults show where the sharp edges are.
- gateway
- Cloudflare AI Gateway
- caching
- Exact-match only; semantic caching is on the roadmap
- retries and failover
- Up to 5 retries, at most 60 s delay; fallback on any error by default
- operational detail
- Cache is 'volatile': two identical concurrent requests can both miss
- gateway
- LiteLLM
- caching
- Exact and semantic
- retries and failover
- Cooldown after 3 failures; immediate 5 s cooldown on 429
- operational detail
- Its docs warn semantic caching 'goes badly wrong on agentic traffic'
- gateway
- Portkey
- caching
- Exact and semantic (default threshold 0.95)
- retries and failover
- Up to 5 attempts on 429, 500, 502, 503, 504, 529
- operational detail
- Semantic matching ignores the system prompt
- gateway
- Kong AI Proxy
- caching
- Exact and semantic plugins
- retries and failover
- Client errors don't fail over unless configured
- operational detail
- Semantic plugin can answer near-matches even with exact caching on
Buying makes sense when the main need is a single interface, provider failover, and basic spend visibility. Building makes sense when the gateway has to understand your product: per-customer credits and plans, use-case profiles owned by feature teams, tenant-scoped caches, and a ledger that billing can trust. Many teams end up doing both: a managed gateway for provider routing and a thin layer of their own for metering and policy.
A checklist
The door
- Every model call goes through the gateway, enforced by a test or an egress policy
- Callers name a use case; model pins are restricted to internal callers
- Unknown and internal-only use cases return the same 404
Order of checks
- Auth and idempotency before anything that costs money
- Model chosen before the reservation, settled on the model that answered
- Refusals happen before streaming headers are written
Failover
- Errors classified into retry, fail over, and stop
- Spend-cap 429s are not treated as transient
- One layer owns retries; SDK retries set to zero
- A circuit breaker per provider
- Schemas and evals hold across every model in the chain
Money
- Reserve and balance check are atomic
- Reservations expire and get swept
- A rule for billing partial streams, tested
- Judge, embedding, and moderation calls are metered
- Quotas per feature and a daily ceiling per deployment
Caching
- Stable prompt content first, variable content last
- Semantic caching only where scoped per tenant and measured
Privacy and visibility
- Redaction and injection checks as per-use-case settings
- No prompt or output text in logs or spans by default
- One decision record per call; no tenant ids in metric labels
- Quality sampled by a judge, and the judge metered
Sources
- Uber, Navigating the LLM Landscape: Uber’s Innovation with GenAI Gateway (July 2024)
- Grab, Grab AI Gateway (February 2025) and How Grab builds and runs AI agents at scale (July 2026)
- LinkedIn, How LinkedIn's platform strategy helps us innovate at speed (February 2024)
- Instacart, Simplifying large-scale LLM processing across Instacart with Maple (August 2025)
- Anthropic, Errors, Rate limits, and Prompt caching
- OpenAI, Error codes, Rate limits, and Prompt caching
- Cloudflare, AI Gateway fallbacks, request handling, and caching
- LiteLLM, Routing and Caching; Portkey, Automatic retries; Kong, AI Proxy Advanced