7 Best AI Gateways in 2026: Kong, LiteLLM, TrueFoundry, Portkey, Cloudflare, Helicone and OpenRouter Compared
We compared Kong, LiteLLM, TrueFoundry, Portkey, Cloudflare AI Gateway, Helicone, and OpenRouter on deployment model, pricing, and licensing — with real pros and cons for each.
The 7 platforms enterprises actually use to route, govern, and monitor traffic to large language models — compared on deployment model, pricing, and where each one breaks down.
If your engineering team is calling five different LLM providers from a dozen different services, you already have the problem an AI gateway solves: no single place to see spend, no consistent way to enforce rate limits or PII policies, and no fallback when a provider has an outage. An AI gateway sits between your applications and model providers — OpenAI, Anthropic, Bedrock, Vertex, self-hosted models — and gives you one API, one place for auth and budgets, and one dashboard for observability. This roundup compares the seven gateways worth evaluating in 2026, based only on what each vendor publishes about deployment model, pricing, and licensing.
How we picked these
We looked for platforms that are purpose-built to sit in the request path between an application and one or more LLM providers — not general observability wrappers, not agent frameworks. Each entry below is verified against the company's own pricing page, documentation, or GitHub repository as of August 2026. Where a company doesn't publish a number, we say so rather than estimate one. Companies are ordered by a mix of deployment flexibility, breadth of governance features, and how verifiable their claims are — not by revenue or funding.
Quick comparison
| Company | Best for | Deployment | Pricing model |
|---|---|---|---|
| Kong AI Gateway | Teams already running Kong for API management | Self-hosted OSS or Konnect (managed/hybrid) | OSS free; Konnect enterprise pricing on request |
| LiteLLM | Platform teams who want to self-host and own their data | Self-hosted (Docker/Helm) or LiteLLM Cloud | Free OSS; Enterprise pricing on request |
| TrueFoundry | Regulated enterprises needing VPC/air-gapped deployment | SaaS, VPC, on-prem, air-gapped | Free up to 50k requests/month; paid plans on request |
| Portkey | Teams wanting a hosted gateway with prompt management built in | Hosted SaaS or self-hosted (open-source core) | Free; $49/mo Production; Enterprise on request |
| Cloudflare AI Gateway | Teams already on Cloudflare's edge network | Cloudflare's global network (SaaS only) | Free with usage-based add-ons |
| Helicone | Startups wanting gateway + observability in one open-source tool | Hosted SaaS or self-hosted (OSS core) | Free; $79/mo Pro; $799/mo Team; Enterprise on request |
| OpenRouter | Developers who want one API key across hundreds of models | Hosted SaaS only | Pass-through model pricing + 5.5% fee on credit purchases |
1. Kong AI Gateway

Kong built its name as an open-source API gateway before extending the same Gateway runtime to handle LLM traffic. Kong AI Gateway plugs AI-specific plugins — token rate limiting, semantic caching, prompt guardrails, multi-LLM load balancing — into the same Kong Gateway data plane that already handles regular API traffic. It also now covers MCP server governance and agent-to-agent traffic, which is a newer and less proven part of the product than its core LLM proxying.
Best for: Organizations that already run Kong Gateway for API management and want AI traffic on the same control plane.
Pros
- Core Kong Gateway is open source (Apache-2.0) with a large existing plugin ecosystem
- Single control plane for both conventional API traffic and LLM traffic, useful if Kong is already in the stack
- Broad protocol coverage: LLM proxying, MCP registry/governance, and agent-to-agent observability
Cons
- Kong doesn't publish self-serve pricing for Konnect (the managed control plane) or for AI Gateway specifically — you need to talk to sales for anything beyond the open-source core
- The MCP and agent-to-agent governance features are recent additions; less battle-tested than the core API gateway
- Running Kong OSS yourself still means operating Postgres/etcd and the Kong data plane — meaningfully more infrastructure than a pure SaaS gateway
2. LiteLLM

LiteLLM is the most widely deployed open-source AI gateway by GitHub activity — over 53,000 stars and 240 million+ Docker pulls, per its own site. It puts an OpenAI-compatible API in front of more than 140 providers and 1,800+ models, with virtual keys, per-team budgets, load balancing, and semantic caching. The company published a benchmark in 2026 claiming its newer Rust-based proxy adds 0.66ms of overhead at p99, versus 2.29ms for Portkey and 4.54ms for Bifrost in the same test — a self-reported comparison worth treating as directional rather than independently audited.
Best for: Platform teams that want to self-host a gateway and keep full control of their data and infrastructure.
Pros
- MIT-licensed open-source core with no feature paywall on the basics — budgets, teams, guardrails, and load balancing are all in the free tier
- Very broad provider coverage (140+) and fast day-zero support for newly released models, per customer testimonials on its site
- Self-hosts anywhere, including fully air-gapped environments, since there's no telemetry phone-home
Cons
- Enterprise pricing (SSO, RBAC, audit logs, support SLAs) isn't published — it's a "talk to sales" conversation
- Self-hosting means you own the operational burden: Postgres, Redis, and the proxy itself all need to be run and monitored by your team
- The headline latency benchmark is self-published by LiteLLM, not run by a neutral third party
3. TrueFoundry

TrueFoundry's AI Gateway is built around enterprise deployment flexibility: it runs as SaaS, inside a customer's VPC, on-prem, or fully air-gapped, with the same policy and observability layer in every mode. It supports more than 1,600 models behind one API, plus MCP tool governance, model routing/fallbacks, and RBAC-based quota management. TrueFoundry also runs a broader AI platform (model deployment, fine-tuning, an agent harness called TrueForge) — the gateway is one product in that suite, not a standalone company.
Best for: Regulated enterprises (healthcare, financial services, government) that need the gateway to run entirely inside their own cloud or on-prem environment.
Pros
- Genuinely flexible deployment — VPC, on-prem, and air-gapped are first-class options, not enterprise-only afterthoughts
- Free tier covers 50,000 requests/month across all 1,600+ models with observability and cost tracking included, useful for evaluation before committing
- Serves self-hosted open-source models (Llama, Mistral) through the same gateway and policy layer as hosted providers, via vLLM/SGLang/Triton support
Cons
- Paid-tier and enterprise pricing beyond the free 50k-request tier isn't published on the site — like most competitors here, that requires a sales conversation
- The company's own "30% average cost optimization" and "10B+ requests processed/month" figures are self-reported marketing stats without an independent audit trail
- Because the gateway is bundled inside a larger AI platform, teams that only want a lightweight LLM proxy may find the product surface (deployment, fine-tuning, agent harness) larger than they need
4. Portkey

Portkey pairs a gateway (universal API, fallbacks, load balancing, caching) with prompt management and evaluation tooling. Notably, in mid-2026 Portkey's enterprise product was rebranded and folded into Palo Alto Networks' Prisma AIRS AI Gateway line, following what its site describes as a partnership/acquisition-style integration — worth checking directly with Portkey/Palo Alto Networks before signing an enterprise contract, since the commercial terms may now route through Prisma AIRS rather than Portkey directly.
Best for: Teams that want prompt versioning and evaluation tooling bundled with the gateway rather than as a separate product.
Pros
- Published, self-serve pricing with a genuine free tier (10k logs/month) and a clear $49/month step to Production
- Open-source gateway core available for self-hosting, in addition to the hosted SaaS
- Prompt management (templates, versioning, playground) is bundled in rather than a bolt-on
Cons
- The Production tier's log limits are tight for high-volume production traffic (100k logs/month before $9-per-100k overage kicks in)
- Its recent enterprise-tier integration with Palo Alto Networks' Prisma AIRS adds a layer of commercial complexity for buyers evaluating long-term vendor stability
- Retention windows on lower tiers are short (3 days on Free, 30 days on Production) compared to some competitors
5. Cloudflare AI Gateway
Cloudflare AI Gateway adds analytics, logging, caching, rate limiting, and request retry/fallback in front of AI providers (OpenAI, Anthropic, Google, Workers AI, and others), available on all Cloudflare plans including the free tier. It's positioned as a lightweight add-on to Cloudflare's existing edge network rather than a standalone governance platform — there's no separate pricing page because it's bundled into Cloudflare's broader developer platform pricing.
Best for: Teams already building on Cloudflare Workers who want basic AI observability and caching without adding a new vendor.
Pros
- Available on Cloudflare's free plan — no separate contract or minimum spend to start
- Runs on Cloudflare's global edge network, so requests are routed from wherever the caller is closest to
- One line of code to get started, per Cloudflare's own documentation, with tight integration into Workers AI and Vectorize
Cons
- Feature set is narrower than dedicated AI gateways — no built-in PII/prompt guardrails, no RBAC-based team budgets comparable to Portkey, LiteLLM, or TrueFoundry
- Effectively locks you into evaluating it as part of a broader Cloudflare commitment rather than as a portable, standalone product
- Documentation is thinner on enterprise governance features (audit logs, SSO) compared to the dedicated AI gateway vendors in this list
6. Helicone
Helicone started as an open-source LLM observability tool and has grown gateway-style features (caching, rate limiting, automatic fallbacks) on top of its logging core. Pricing is transparent and usage-based: a real free tier (10,000 requests/month), a $79/month Pro tier, and a $799/month Team tier with SOC 2/HIPAA compliance and a dedicated Slack channel.
Best for: Startups that want observability-first tooling with gateway features layered on, rather than a governance-first platform.
Pros
- Fully transparent, calculator-based pricing — you can estimate your bill before signing up, unlike most gateways on this list
- Open-source core (5,800+ GitHub stars) with a genuinely useful free tier for small projects
- Discounts published for startups (50% off first year), non-profits, open-source projects, and students
Cons
- Gateway-specific features (caching, rate limits, automatic fallbacks) are less mature than its observability tooling — it reads as observability-first, gateway-second
- SOC 2 and HIPAA compliance are gated behind the $799/month Team tier, pricing out smaller regulated teams
- No stated support for self-hosted/air-gapped enterprise deployment comparable to Kong, LiteLLM, or TrueFoundry
7. OpenRouter
OpenRouter is closer to a model marketplace than a governance-focused gateway: one API key, one credit balance, automatic fallback across providers, and pass-through pricing with no markup on inference — OpenRouter's own FAQ states it charges a 5.5% (minimum $0.80) fee only when you purchase credits via Stripe, or 5% via crypto. It doesn't publish enterprise governance features like RBAC or audit logging in the same depth as Kong, Portkey, or TrueFoundry.
Best for: Individual developers and small teams who want quick access to many models without picking a single provider or negotiating enterprise contracts.
Pros
- No markup on model pricing — you pay the provider's list price, only the credit-purchase fee is added on top
- Extremely broad model catalog spanning many providers, with automatic failover if one provider errors out
- Drop-in compatible with the OpenAI SDK, so migration from a direct OpenAI integration is close to a one-line change
Cons
- No volume discounts, per its own FAQ — pricing doesn't improve as usage scales, unlike enterprise-negotiated gateway contracts
- Limited enterprise governance tooling (no published RBAC, SSO, or audit-log feature set) compared to the other six entries here
- Credits can expire after one year of purchase per OpenRouter's terms, and unused-credit refunds are only honored within 24 hours of purchase
How to choose
If you're already standardized on an infrastructure vendor — Kong for API management, Cloudflare for edge/CDN — start there before adding a new vendor relationship. If deployment flexibility is the hard requirement (VPC, on-prem, air-gapped, for compliance reasons), TrueFoundry and self-hosted LiteLLM or Kong OSS are the entries that explicitly support that. If you want the lowest-friction way to try many models without picking a governance platform, OpenRouter is built for exactly that and nothing more. If prompt management and evaluation matter as much as routing, Portkey folds those in; if observability and transparent, calculator-based pricing matter more, Helicone leans that direction. None of the platforms here publish independently audited latency or cost-savings numbers, so treat any single vendor's benchmark as a starting point for your own testing, not a final verdict.
Frequently Asked Questions
What is an AI gateway, exactly?
An AI gateway is middleware that sits between your applications and LLM providers (OpenAI, Anthropic, Bedrock, self-hosted models). It gives you one API and one set of credentials to manage, plus centralized logging, rate limiting, caching, and fallback routing across providers.
Do I need an AI gateway if I only use one model provider?
Probably not yet. Gateways earn their keep once you're calling multiple providers, need centralized cost/usage tracking across teams, or want automatic fallback if a provider has an outage. A single-provider setup can usually get by with that provider's own SDK and dashboard.
Is it safe to self-host an open-source gateway like LiteLLM or Kong instead of using a hosted SaaS?
Yes, and it's a common choice for compliance-sensitive teams, but it shifts operational responsibility to you — you're running the proxy, its database, and its scaling instead of a vendor doing it. Weigh that against the value of not sending traffic metadata to a third party.
How much do these gateways typically cost?
It ranges widely. Cloudflare AI Gateway and the open-source cores of LiteLLM, Kong, and Portkey are free to start. Hosted plans with more logging, retention, and support run from roughly $49/month (Portkey Production) to $799/month (Helicone Team), with most vendors reserving true enterprise pricing — SSO, VPC deployment, dedicated support — for custom quotes.
Does using a gateway add noticeable latency to my LLM calls?
Vendors report gateway-added overhead in the single-digit milliseconds (LiteLLM's self-published benchmark claims 0.66ms at p99 for its Rust proxy), which is small relative to typical LLM response times of hundreds of milliseconds to several seconds. Treat vendor-published latency numbers as a starting point and benchmark your own workload before committing.
Editor's note — sources: Company pricing and product pages fetched directly (konghq.com, litellm.ai, truefoundry.com, portkey.ai, developers.cloudflare.com, helicone.ai, openrouter.ai) as of August 2026. GitHub star counts and license terms are as published by each project. Where a company does not publish a specific number — enterprise pricing, SLA terms, audited benchmark data — this article says so rather than estimating.