Roundups

Top 7 AI Gateways for Production LLM Workloads in 2026

The best AI gateways in 2026 compared: Portkey, LiteLLM, Kong AI Gateway, Cloudflare, TrueFoundry, OpenRouter and Helicone, with features, real pricing and honest pros and cons.

Illustration of an AI gateway routing LLM traffic between an application and many model providers

An AI gateway is a proxy that sits between your application and large language model providers, giving you one API for many models plus routing, fallbacks, cost controls, guardrails, and observability. The strongest options in 2026 are Portkey, LiteLLM, and Kong AI Gateway, with Cloudflare AI Gateway, TrueFoundry, OpenRouter, and Helicone each fitting specific needs. Here is how they compare.

TL;DR

  • Portkey is the most feature-complete dedicated gateway (routing, guardrails, prompt management, MCP) — now owned by Palo Alto Networks as of May 2026.
  • LiteLLM is the open-source standard: MIT-licensed, 58,000+ GitHub stars, and the widest model coverage (140+ providers).
  • Cloudflare AI Gateway is effectively free with any Cloudflare account; OpenRouter passes through inference costs with no markup and was acquired by Stripe in August 2026.
  • Published entry pricing (as of October 2026): Portkey $49/mo, Helicone $79/mo Pro, TrueFoundry free to 30,000 requests/mo. Most enterprise tiers are quote-only.

What an AI gateway does

An AI gateway centralizes LLM traffic. Instead of each service calling OpenAI, Anthropic, or an open model directly, calls route through the gateway, which handles authentication, provider failover, caching, rate limiting, spend tracking, and logging in one place.

That matters once you run more than one model in production. You get a single audit trail, the ability to swap providers without touching application code, and a choke point to enforce budgets and content policies. Several gateways now also proxy Model Context Protocol traffic for agent tool calls, which is becoming a distinct requirement of its own.

The category overlaps with — but is separate from — the AI inference providers that actually run the models, and the AI guardrails and LLM security platforms that specialize in content filtering.

How we picked these

We weighted seven criteria: breadth of model and provider coverage; routing and reliability features (fallbacks, load balancing, caching); governance (guardrails, RBAC, budgets); observability depth; deployment flexibility (SaaS, self-hosted, open source); pricing transparency; and verifiable adoption. We made no judgment on sponsorship or placement — none exists. Rankings reflect how complete each product is for running LLM traffic in production, not alphabetical order or popularity alone.

Quick comparison

Company Best for Deployment Pricing model
Portkey All-in-one governance SaaS + self-hosted Free / $49 mo / custom
LiteLLM Open-source self-hosters Self-hosted (MIT) / managed Free / enterprise quote
Kong AI Gateway Teams already on Kong OSS core + Konnect SaaS Consumption / quote
Cloudflare AI Gateway Existing Cloudflare users SaaS only Bundled with Workers
TrueFoundry Gateway + model serving SaaS + self-hosted VPC Free to 30K req / demo
OpenRouter One key, many models SaaS only Pass-through + fees
Helicone Observability first SaaS + self-hosted Free / $79 mo / custom

1. Portkey

Portkey is a full-stack AI gateway offering a unified API to more than 1,600 models, with routing, fallbacks, load balancing, semantic caching, guardrails for PII and prompt security, prompt management, and an MCP gateway for agent traffic. It is the most feature-complete dedicated gateway on this list.

In May 2026, Palo Alto Networks completed its acquisition of Portkey for a stated $117 million, and in July 2026 the product was folded into the vendor's Prisma AIRS security platform. The Portkey brand continues under that ownership.

Best for: teams that want routing, observability, and guardrails in a single product.

Pros

  • Combines gateway, observability, guardrails, prompt management, and an MCP gateway in one tool.
  • Transparent entry pricing: a free tier with 10,000 logs/month and a Production plan at $49/month (as of October 2026).
  • Now backed by a large security vendor, which helps enterprise procurement.

Cons

  • Post-acquisition roadmap and pricing carry uncertainty; the public pricing page may lag the Prisma AIRS reality.
  • Log caps on lower tiers push teams toward custom pricing quickly.
  • The product is now subordinate to a security platform's priorities rather than an independent roadmap.
Portkey open-source AI gateway GitHub repository
Image: Portkey

2. LiteLLM

LiteLLM is the open-source AI gateway most teams reach for first. It exposes one OpenAI-compatible API across 140+ providers and more than 1,800 models, with spend tracking, budgets, lowest-cost routing, semantic caching, guardrails, and virtual keys. The core is MIT-licensed; an Enterprise edition sits under a separate commercial license.

Adoption is the broadest here: 58,800 GitHub stars and over 240 million Docker pulls (vendor-stated, October 2026). It can run fully self-hosted, including air-gapped, via Docker, Helm, or Terraform.

Best for: platform teams standardizing LLM access org-wide on open source.

Pros

  • Free and MIT-licensed, with a fully self-hostable core including air-gapped deployment.
  • The widest provider and model coverage of any gateway here.
  • A large, active community with a public changelog and security disclosures.

Cons

  • Enterprise pricing is sales-only, with no published figure.
  • Self-hosting means operating your own Postgres, Redis, and Kubernetes.
  • Its headline latency benchmark against competitors is vendor-run, not independently audited.
LiteLLM open-source LLM gateway and proxy homepage
Image: LiteLLM

3. Kong AI Gateway

Kong AI Gateway is an AI module layered on Kong's mature API gateway. It adds multi-LLM routing, semantic caching, semantic prompt guards, token-cost governance, agent-to-agent traffic handling, and MCP support. The underlying Kong Gateway core is Apache 2.0; the advanced AI features run through Kong's paid Konnect platform.

If your organization already runs Kong for API management, extending into LLM governance is incremental rather than a new system. Kong cites customers including Rabobank and Siemens across its broader platform.

Best for: enterprises already standardized on Kong for API management.

Pros

  • Built on a battle-tested, widely deployed open-source gateway core.
  • A broad plugin and governance ecosystem that reaches well beyond LLM routing.
  • Proven operation at large-enterprise scale.

Cons

  • AI Gateway enterprise pricing is quote-only, with no public figure.
  • Advanced features like semantic guards are gated behind paid Konnect or Enterprise tiers.
  • Headline adoption stats are platform-wide, so isolating AI Gateway traction is difficult.
Kong API and AI gateway GitHub repository
Image: Kong

4. Cloudflare AI Gateway

Cloudflare AI Gateway is an edge-network proxy in front of LLM providers, offering a unified API, caching, rate limiting, logging, and cost tracking. It is built into Cloudflare's Workers AI platform and runs on the company's global edge.

Its main appeal is price and reach: it is free with any Cloudflare account, with no separate gateway fee beyond your Workers plan. The Workers free tier includes 100,000 logs; paid includes 1 million, with a Logpush add-on at $0.05 per additional million requests over 10 million/month (as of October 2026).

Best for: teams already on Cloudflare wanting a lightweight gateway with no new infrastructure.

Pros

  • Effectively free and bundled with an existing Cloudflare account.
  • Runs on Cloudflare's global edge for low default latency.
  • Backed by a large, stable public infrastructure company.

Cons

  • Locked into the Cloudflare ecosystem with no self-hosted option.
  • A narrower feature set than dedicated gateways — less depth on guardrails, prompt management, and RBAC.
  • Log-volume caps mean serious observability needs the separate Logpush add-on.
Cloudflare AI Gateway developer documentation
Image: Cloudflare

5. TrueFoundry

TrueFoundry offers an AI gateway within a broader AI-infrastructure platform. The gateway provides a unified API to more than 1,600 models with routing, fallbacks, an observability dashboard, guardrails, and an MCP gateway. What distinguishes it from pure-routing peers is that the same platform also runs self-hosted model serving on vLLM, SGLang, KServe, and Triton.

The free tier covers 30,000 requests/month across all models with up to three seats; paid tiers are available by demo, so effectively pricing on request. TrueFoundry last disclosed a $19 million Series A, closed in December 2024, led by Intel Capital and Peak XV Partners.

Best for: platform teams wanting a gateway alongside model-deployment infrastructure.

Pros

  • Pairs gateway routing with real model-serving infrastructure, a genuine differentiator versus routing-only tools.
  • A generous free tier: 30,000 requests/month across 1,600+ models.
  • Named enterprise customers across varied industries, including Siemens Healthineers and ResMed.

Cons

  • All paid pricing is gated behind a sales demo, with no self-serve paid tiers.
  • Open-source and community adoption signals for the gateway specifically are weak.
  • No funding disclosed since December 2024, versus more recently capitalized competitors.
TrueFoundry AI Gateway observability dashboard screenshot
Image: TrueFoundry

6. OpenRouter

OpenRouter is a unified API and marketplace that routes across 500+ models from 80+ providers, with provider fallback, price and performance routing, and bring-your-own-key support. It is the most developer-popular way to hit many models through one key.

Its pricing model is pass-through: no markup on inference. OpenRouter charges a 5.5% fee ($0.80 minimum) on credit purchases via card, 5% via crypto, and a 5% fee on BYOK usage above $25,000/month of list-price inference (as of October 2026). Stripe announced its acquisition of OpenRouter in August 2026; the company continues to operate under its own brand per its founder.

Best for: developers wanting one key across many providers with pass-through pricing.

Pros

  • True pass-through inference pricing with no per-token markup.
  • Very broad model and provider coverage with uptime-optimized routing.
  • Now backed by Stripe's infrastructure and balance sheet.

Cons

  • Fully dependent on OpenRouter's hosted infrastructure, with no self-hosted option.
  • The acquisition introduces integration and roadmap uncertainty despite stated continuity.
  • Real added costs via the credit-purchase fee and the BYOK overage fee at scale.
OpenRouter unified interface for AI models
Image: OpenRouter

7. Helicone

Helicone is an observability-first gateway. Its primary strength is request logging, tracing, caching, prompt management, and cost tracking, with a gateway proxy layered on. If your main need is understanding and debugging LLM traffic rather than multi-provider routing, it fits.

Pricing is transparent: a free Hobby tier with 10,000 requests/month, Pro at $79/month, Team at $799/month, and custom Enterprise (as of October 2026). It offers both a managed service and an open-source self-hosted option, and is Y Combinator-backed.

Best for: teams that prioritize LLM observability and debugging over routing.

Pros

  • Observability-first design with more detailed tracing than bolt-on logging elsewhere.
  • Both a free self-serve tier and an open-source self-hosted option.
  • Y Combinator-backed with named enterprise customers.

Cons

  • The smallest GitHub community of the seven, at 6,100 stars.
  • Routing and governance features are lighter than dedicated gateways.
  • Publicly reported seed-funding figures conflict across sources.
Helicone open-source LLM observability gateway GitHub repository
Image: Helicone

How to choose

Match the tool to your constraint, not the hype.

  • You want open source and full control: LiteLLM, self-hosted. Budget for running the supporting infrastructure.
  • You want the most features in one product: Portkey — but weigh the Palo Alto Networks roadmap question.
  • You already run Kong or Cloudflare: extend what you have rather than add a system.
  • You need a gateway and model serving together: TrueFoundry.
  • You just want one key across every model: OpenRouter.
  • Your real problem is observability: Helicone.

For most teams running multiple models in production, the practical shortlist is LiteLLM if you self-host and Portkey if you buy. Validate the current pricing yourself — several vendors here changed ownership in 2026.

Frequently Asked Questions

What is an AI gateway?

An AI gateway is a proxy between your application and large language model providers. It gives you one API for many models and centralizes routing, provider failover, caching, rate limiting, spend tracking, guardrails, and logging, so you can switch providers and enforce policy without changing application code.

How much does an AI gateway cost?

It ranges from free to enterprise quotes. Cloudflare AI Gateway is bundled free with a Cloudflare account, and LiteLLM's core is free and open source. Paid entry tiers run roughly $49–$79/month (Portkey, Helicone) as of October 2026, while Kong and TrueFoundry price enterprise use by quote.

How do AI gateways control costs and token usage?

Gateways track spend per key, team, or model and enforce budgets that cut off or throttle traffic at a limit. Semantic caching returns stored answers for repeat queries, and lowest-cost routing sends each request to the cheapest provider that meets your latency and quality bar.

Is Cloudflare AI Gateway free?

Yes, with limits. It is included with any Cloudflare account at no separate charge. The Workers free tier covers 100,000 logs and the paid tier 1 million; heavier logging requires the Logpush add-on, billed at $0.05 per additional million requests above 10 million/month (as of October 2026).

What is an MCP gateway?

An MCP gateway proxies Model Context Protocol traffic — the tool and data calls that AI agents make. It applies the same routing, authentication, and observability to agent tool use that a standard AI gateway applies to model calls. Portkey, LiteLLM, Kong, and TrueFoundry all offer MCP handling.

Editor's note — sources: Pricing, model counts, and feature claims are drawn from each vendor's official site, documentation, and pricing pages as of October 2026, and are attributed as vendor-stated where not independently verified. Acquisition facts (Palo Alto Networks/Portkey, Stripe/OpenRouter) come from the acquirers' announcements and the companies' own statements. Deal prices reported only in the press are omitted. GitHub stars and adoption figures are vendor-stated.

Get Edgewisely in your inbox

Business stories that matter, free. Enter your email — no password, no account to set up.
jamie@example.com
Subscribe