An AI gateway is a control point between your applications or AI agents and the model providers they call. Code points its existing SDK at the gateway instead of the provider; the gateway holds the provider credential, decides whether the call may go out, and applies spend and content rules in one place rather than in every service. "LLM gateway" and "model gateway" usually mean the same thing.
This guide is for platform, security and FinOps leads who are deciding whether they need one and what to ask of it. It covers what an AI gateway does, how it differs from an MCP gateway and an API gateway, what bring-your-own-key means in practice, the requirements worth putting in an evaluation, the questions to ask any vendor, and, in the last section, how Praesidia's model gateway works and where it stops.
What is an AI gateway?
An AI gateway is a proxy for model traffic that adds policy. A request for a completion or an embedding goes to the gateway, which checks who is calling, attaches the right provider credential, applies the rules that cover that caller, forwards the request to the provider and returns the response, streamed or not.
Teams usually adopt one for four reasons:
- Credential sprawl. Without a gateway, every service and every agent carries a provider API key. Each copy is something to rotate, and each one leaked is a bill you did not plan for.
- Spend that compounds. Agents loop, retry and call models many times per task. A cost limit is only useful if it is checked before the call goes out, not reconciled from the invoice.
- Data leaving the building. Prompts carry customer data, secrets and internal documents. A single control point is where content rules can be applied consistently.
- One place to stop. When something goes wrong, you want one switch for model traffic, not a hunt through every service that holds a key.
Model traffic is not the only traffic an agent produces. An agent also calls tools, often through MCP servers, and talks to other agents. A model gateway sees none of that, which is why the comparison below matters.
AI gateway vs MCP gateway vs API gateway
The three names overlap in marketing, but each sits in front of something different and makes a different decision.
| AI gateway (LLM gateway) | MCP gateway | API gateway | |
|---|---|---|---|
| Sits in front of | Model providers: hosted APIs and self-hosted models | MCP servers that expose tools and data to agents | Your own HTTP APIs and services |
| Typical caller | An application or agent asking a model for a completion or embedding | An agent calling a tool on behalf of a task | An application, partner or user client |
| The question it answers | May this caller send this request to this model, with which credential, within which budget, and does the content pass? | May this agent call this tool, with these arguments, on this server? | May this client reach this route, at this rate? |
| Credentials it handles | Provider API keys, kept away from the caller | The server's credential or the agent's identity | Client API keys or tokens for your own APIs |
| Typical controls | Spend limits, content guardrails, provider choice, a stop switch for model traffic | Tool allow and deny lists, argument constraints, approval before a risky call runs | Authentication, rate limits, routing, request transformation |
| What it cannot see | Tool calls and agent-to-agent traffic that do not go through it | Model calls and anything that skips the gateway | What the model or agent intended; traffic outside its routes |
An API gateway you already run does not become an AI gateway by adding a plugin: it can route a request to a model provider, but it rarely understands tokens, model pricing or prompt content. The AI control plane vs API gateway post covers that boundary in more depth, and what an MCP gateway is is defined in the glossary.
What bring-your-own-key means for an AI gateway
Bring your own key (BYOK) means you keep your own accounts with model providers and give the gateway the provider keys, rather than buying model access through the gateway vendor. Three consequences follow.
- The provider relationship stays yours. Your provider contract, data-processing terms, model access and rate limits apply. Provider charges are billed to your provider account, so a free gateway plan does not make model calls free.
- Keys leave your code. Applications and agents hold a gateway credential instead of the provider key. Revoking that credential cuts off one application without rotating the provider key that every other application relies on.
- Rotation happens in one place. When you rotate a provider key, you change it in the gateway configuration, not in every deployment that used to carry it.
BYOK does not move data residency or retention decisions to the gateway. Where a prompt is processed and how long the provider keeps it are still set by the provider and the account you chose. The BYOK and LLM configuration post covers managing provider keys across teams.
What to require from an AI gateway
These are the properties worth writing into an evaluation. They apply to any product, including one you build.
| Requirement | Why it matters | How to test it |
|---|---|---|
| Provider keys held by the gateway, never by the caller | A key in application code or an agent's environment can be copied and used outside every control you set | Confirm the calling code only ever sees a gateway credential, and that the provider key cannot be read back after it is saved |
| Works with the SDKs you already use | A gateway that needs a rewrite of every client will be bypassed | Point an existing OpenAI-compatible or Anthropic client at it by changing only the base URL and key |
| Spend checked before the call leaves | A limit enforced after the response only tells you how much you overspent | Set a small budget, send calls until it is reached, and confirm the next one is refused rather than billed |
| Content rules on the prompt and the response | Sensitive data goes out in prompts and can come back in responses | Send a test prompt containing a fake card number or secret and check it is blocked or redacted as configured |
| Control over outside domains | Agents' tools can reach destinations nobody approved | Allow one domain, deny another, and check which calls are refused |
| One switch to stop traffic | An incident should not depend on finding every key | Freeze the organization in a test workspace and confirm model calls stop |
| Defined behaviour when its policy service is unreachable | A gateway that lets calls through unchecked during an outage is a gap at the worst moment | Ask whether it fails closed or open, and get the answer in writing |
| A stated recording scope | Auditors will ask what the record covers, and what it leaves out | Ask which model calls are written to the audit record and which are not |
The cost side of these requirements, from attribution to forecasting, is covered in AI FinOps: controlling AI agent costs. The attack that makes pre-call spend checks matter is described in denial of wallet, and the OWASP Top 10 for LLM Applications lists it as LLM10:2025 Unbounded Consumption.
Questions to ask any AI gateway vendor
- Where are provider keys stored, who can read them, and can they ever be returned to a caller?
- Which providers can I reach through one gateway address, and for each, which API surfaces: chat, embeddings, streaming?
- Do I keep my existing SDKs, or do I adopt yours?
- Is a budget checked before a call is sent? At which scopes: application, agent, organization?
- Which content rules run on the prompt, which on the response, and what can a match do: block, redact, warn?
- What happens to calls when the gateway cannot reach the service that holds its policies?
- Can I stop all model traffic for the organization in one step, and what does that not stop?
- Which model calls end up in an audit record, which do not, and can the record be checked without your platform?
- Does the gateway also govern tool calls and MCP traffic, or only model calls?
- Which of these controls depend on my plan?
A vendor that answers question 8 with "everything" has either not checked or is not telling you. A useful answer names what is recorded and what is not.
How Praesidia's model gateway works
Praesidia's model gateway gives your applications and agents one gateway address in front of the model providers you use, with your own provider keys. What it does, and where it stops:
- Your keys, held by the gateway. You add a provider key once in your workspace's LLM configuration; bring-your-own-key model configuration is in every plan. Your code sends a Praesidia application key, and the gateway resolves the configured provider credential server-side, so the provider key is not in your application. Real provider calls incur provider charges, including when you use a free Praesidia plan.
- The providers you can reach. Through the gateway's one base URL or as an agent's configured model, from the same list on both paths: OpenAI, Anthropic, Azure OpenAI, Google Gemini, Mistral AI, Cohere, DeepSeek, Qwen, Moonshot, a self-hosted Ollama, or your own OpenAI-compatible endpoint.
- Your existing SDK. OpenAI-compatible clients change only the base URL and key. Anthropic clients use Anthropic's own Messages API through the gateway rather than a translated one. The address for your workspace comes from your administrator or Praesidia support.
- Spend checked before the call. Spend caps are checked before a call leaves. Budget policies, on the Team and Enterprise plans, refuse a model call through the gateway that would exceed an agent's or the organization's budget; no budget is checked on an MCP tool call. Usage appears in the workspace by application and organization, after a processing delay.
- Content and destinations. Guardrails apply to model traffic through the gateway and can block or redact what they match. The outside domains your agents' tools reach can be allowed or denied.
- One switch, and a defined failure mode. An organization freeze stops model traffic too. If the enforcement point cannot reach the control plane, governed calls are refused rather than let through unchecked.
The limits matter as much as the features:
- Recording. A model call through the gateway reaches the trail only when the gateway blocks it or requires a step-up; an allowed gateway call is not written there. The gateway's redactions and warnings are not written to the audit trail either. Praesidia records the decisions it makes in an append-only, hash-chained audit trail, signed under the default configuration; an operator can turn signing off. A signed record shows it was not changed afterwards. It does not show that every action was captured. Calls that do not pass through Praesidia are not recorded.
- Stopping one agent. Once an agent is quarantined or revoked, model calls made through the gateway with its application key continue to be accepted for up to a minute, until the gateway's next policy refresh refuses the key; requests already running then finish, and a streamed response can last up to 10 minutes. Its budget and egress policy still apply in that window.
- Residency. Requests that carry no organization, and the model gateway's own calls, are not checked against the region. Where a prompt is processed is set by the provider and the account you configured.
- Tools are a separate path. The model gateway carries model calls, not MCP tool calls. Tool calls are governed on their own path by the MCP gateway, which allows or denies each tool per connection for the MCP servers you register.
To send a first request, follow your first model request. How the gateway fits with the other runtime controls is on AI agent runtime security, and budgets are covered in AI FinOps.
Common questions
What is an AI gateway?
An AI gateway is a control point between applications or AI agents and the model providers they call. It holds the provider credentials, decides whether a model call may go out, and applies spend and content rules in one place instead of in every service.
Is an AI gateway the same as an LLM gateway?
In most usage, yes. "AI gateway", "LLM gateway" and "model gateway" all describe a proxy for model traffic that adds credentials, cost controls and content rules. Some vendors use "AI gateway" more broadly to include tool or MCP traffic, so ask which traffic a product actually sees.
What is the difference between an AI gateway and an MCP gateway?
An AI gateway sits in front of model providers and decides whether a model call may go out. An MCP gateway sits in front of MCP servers and decides whether an agent may call a given tool with given arguments. An agent that both reasons with a model and acts through tools needs a control on each path.
Do I need an AI gateway if I already have an API gateway?
An API gateway can route requests to a model provider, but it usually does not understand tokens, model pricing, budgets or prompt content, and it does not keep provider keys away from callers. If you only proxy a single model for a single application, it may be enough. Once several teams or agents share providers and budgets, a gateway built for model traffic is easier to govern.
What does BYOK mean for an AI gateway?
Bring your own key means you keep your own provider accounts and give the gateway the provider keys. Your contract, data terms and charges stay with the provider, and your code holds only a gateway credential that you can revoke without rotating the provider key.
Does an AI gateway add latency?
Any gateway adds a network hop and the time its checks take. Whether that matters depends on your models and traffic, so measure it on your own workload, with streaming, before and after you route through the gateway.
Which providers can Praesidia's model gateway reach?
OpenAI, Anthropic, Azure OpenAI, Google Gemini, Mistral AI, Cohere, DeepSeek, Qwen, Moonshot, a self-hosted Ollama, or your own OpenAI-compatible endpoint, using your own provider keys. The model request quickstart walks through the first call.
Can a budget stop a model call before it is sent?
With Praesidia, budget policies, on the Team and Enterprise plans, refuse a model call through the gateway that would exceed an agent's or the organization's budget. No budget is checked on an MCP tool call. See AI FinOps for how to set budgets and scopes.
Does Praesidia record model calls in its audit trail?
A model call through the gateway reaches the trail only when the gateway blocks it or requires a step-up; an allowed gateway call is not written there. A signed record shows it was not changed afterwards. It does not show that every action was captured.
Further reading
- Your first model request: configure a provider and send one request through the gateway.
- AI FinOps: the complete guide: attributing, budgeting and forecasting agent spend.
- MCP gateway and security: the control on the tool-call path.
- LLM gateway comparison 2026: how gateway products differ on routing, spend control and security.
- Bring your own key: managing LLM configurations: keeping provider keys in one place.
- AI control plane vs API gateway and the AI control plane guide: where the gateway sits in a wider control plane.
Primary source: OWASP Top 10 for LLM Applications 2025.