Best LLM Gateways in 2026: Portkey vs LiteLLM vs Cloudflare vs Helicone — Pricing & ROI

2026-10-06 · AI Developer Tools · · 📖 32 min read
⚡ TL;DR
Portkey, LiteLLM, Cloudflare AI Gateway, and Helicone compared on routing, failover, pricing, operational burden, and the ROI test teams should run before production.

A one-million-request workload averaging 1,000 input tokens and 200 output tokens per call sends 1.2 billion model tokens through the stack: 1.0 billion in, 200 million out, before retries. That is a workload calculation, not a market statistic. An LLM gateway will not shrink that bill just by sitting between your app and a model provider.

It can make model access more reliable, observable, and governable. Whether those controls repay their price depends on your traffic, failure rate, engineering capacity, and ability to route or cache requests without degrading answers. Portkey, LiteLLM, Cloudflare AI Gateway, and Helicone all sit in this middle layer, but they sell different operating models. This comparison uses their public pricing and product documentation checked October 6, 2026; it is not a controlled performance benchmark.

What a model routing layer should solve before you buy one

A gateway receives an application request, applies policies or routing logic, then sends it to a model provider. The useful features are mundane: one API shape across providers, credential control, per-team budgets, request logs, rate limits, retries, fallback destinations, and a way to see which model produced a costly or failed response. Some products also add prompt management, caching, guardrails, or provider credits.

The business case is not “we have many models.” It is a specific recurring cost or risk. Perhaps a provider outage interrupts checkout. Perhaps every product team has copied API keys into a different secret store. Perhaps no one can explain why inference spend rose 38% last month. If none of these causes actual lost revenue, excess labor, or compliance exposure, adding a proxy can create one more service to maintain without fixing a material problem.

Treat an LLM gateway as a control plane, not a magic model optimizer. A fallback to a cheaper or more available model can change tone, accuracy, tool-call behavior, or structured output. A retry can duplicate a request that was processed upstream but timed out on the return path. A cache can serve a stale answer or the wrong user’s response if keys are not scoped correctly. Those are design decisions, not checkboxes.

LLM gateway pricing comparison: costs that deserve a line item

ProductPublic gateway price checked Oct. 6, 2026Routing and controlLikely ROI fit
PortkeyDeveloper free; Production listed at $49/month; Enterprise customConfig-based fallback, load balancing, retries, caching, and observability; open-source gateway optionTeams that want a managed control plane and a clear monthly entry price
LiteLLMOpen-source gateway is $0 to self-host; enterprise is quote-based100+ provider connections, virtual keys, budgets, rate limits, retries, and fallbacksPlatform teams that value control and can own deployment and on-call work
Cloudflare AI GatewayCore features on all plans; inference passes through at provider rates; Unified Billing credits carry a 5% feeAnalytics, caching, rate limits, request retries, and model fallbackApps already on Cloudflare that want a low-cost first control layer
HeliconeHobby free for 10,000 requests and 1 GB storage; Pro $79/month; Team $799/month; usage-based charges applyRequest analytics plus caching, limits, automatic fallbacks, and provider routingProduct teams wanting gateway controls tied closely to debugging and usage review

These figures are list-page snapshots, not quotes. Portkey's plan table includes request and retention limits that should be confirmed against the current checkout page. Cloudflare log pricing depends partly on when the first gateway was created; its documentation now routes new customers to Workers Logs pricing. Helicone lists paid tiers alongside usage charges, so the plan price is not the full observability bill. LiteLLM’s zero license cost does not include the staff time or infrastructure needed to operate it.

The categories also differ. LiteLLM is an open-source proxy you can run yourself. Cloudflare’s gateway is part of a broader edge platform. Portkey sells a managed gateway layer and publishes an open-source option. Helicone combines gateway functions with request analysis. Compare the full workflow you need, not just the row with the smallest dollar figure.

How each option earns its place

Portkey: managed controls without building the control plane

Portkey is a practical candidate when an application team needs routing policies, request-level visibility, and centralized settings but does not want to assemble those pieces from separate services. Its documentation describes configs for fallbacks, load balancing, retries, and caching; a config can be attached to requests or API keys. The gateway also has an open-source deployment path, though self-hosting it does not automatically give you the hosted product’s support and operations.

The current pricing page lists a free Developer plan and Production at $49 per month, with Enterprise priced by sales. That gives a small team a visible starting number, but verify what the request allowance, log retention, overage, and policy limits mean for your account. The fine print matters more than the headline once logs become evidence for an incident review or compliance audit.

Portkey is a fit when teams need one place to configure behavior across several apps. It is less compelling for a single low-volume service using one model and one provider key: a few environment variables and direct provider logs may be easier to support. For procurement, compare managed and self-hosted costs separately rather than assuming they are interchangeable.

LiteLLM: flexible self-hosting with an operations bill

LiteLLM’s official pricing page lists its open-source gateway at $0 to self-host, with enterprise governance, security, and support sold on an annual quote. Its documentation highlights 100+ providers behind an OpenAI-style API, virtual keys, team budgets, rate limits, spend tracking, retries, and fallback handling. This makes it attractive when a platform group wants the proxy inside its own network and can manage credentials, deployment, upgrades, and alerts.

The phrase “free gateway” answers only the license question. It does not price a production deployment, a database for logs, monitoring, backups, incident response, version testing, or the engineer who gets paged when the proxy is down. Put those items in the LiteLLM proxy vs managed gateway calculation. A small team may pay more in labor than it saves on software; a mature platform team may consider that control worth the work.

Fallback quality also needs testing. LiteLLM documents retries and model or provider fallback, while its router can track deployment health and cooldowns. A fallback chain should preserve the response contract your app expects. Test tool calls, JSON schemas, context limits, latency, and refusal behavior on each destination before allowing automatic switching in production.

Cloudflare AI Gateway: low-friction controls for Cloudflare stacks

Cloudflare’s pricing documentation says core gateway features are available across plans, including dashboard analytics, caching, and rate limiting. Its product overview also lists request retries and model fallback, with support for several external providers. For a team already using Cloudflare, this can be a low-friction way to put visibility and basic controls around outbound model traffic.

There are two billing details worth separating. First, provider inference is passed through at provider rates without a markup under the documented Unified Billing arrangement, but purchasing credits through that route adds a 5% fee. Second, logs are not one timeless free bucket: the pricing page, updated September 24, 2026, distinguishes new customers from accounts using legacy logs and refers new accounts to Workers Logs pricing and retention.

The best Cloudflare AI Gateway pricing for new accounts is therefore not a single number. Core gateway access may be free, while persistent logging, Workers AI guardrail inference, purchased credits, and the rest of your Cloudflare footprint can have separate charges. This is appealing if Cloudflare already operates your app’s edge and identity controls. It is not a reason by itself to move an application or accept a new data path.

Helicone: gateway decisions with usage evidence nearby

Helicone’s public page lists a Hobby tier with 10,000 requests and 1 GB of storage at no charge, Pro at $79 per month, and Team at $799 per month; usage-based charges apply to paid plans. The same page lists gateway features such as caching, rate limits, and automatic fallbacks, while its routing documentation describes provider selection and failover. It emphasizes request analysis as much as routing.

That combination can shorten the feedback loop: inspect a failure, see its provider and cost, then adjust routing or retry rules. The ROI case is strongest when the team already spends time reconstructing request histories across application logs and provider dashboards. Retention is tiered on the pricing page, from seven days on Hobby to longer windows on paid plans, so check whether incident and audit needs exceed the included period.

Helicone’s documentation says routing can select among providers offering the requested model, with the gateway applying user keys first when configured and managed keys as a fallback. Automatic provider selection is convenient, but some organizations require a named region, deployment, or contract. In those cases, pin the route deliberately and test that failover cannot escape the approved boundary.

Model the return before comparing small monthly fees

Start with three costs: gateway subscription or usage charges, model inference, and internal operations. Then list the benefits you can measure: fewer failed requests, less engineering time spent tracing provider issues, lower inference spend from safe caching or routing, or reduced credential-management work. Do not count a retry as a saved request unless it actually avoids a user-visible failure; often it creates another model call.

A simple cache example shows why measurement beats a vendor demo. Assume 100,000 repeated requests per month, each with $0.002 of avoidable model cost, and a cache that safely serves half of them. The theoretical inference reduction is $100 monthly: 100,000 × 50% × $0.002. If the gateway, extra logging, and engineering support cost $300 per month, caching alone does not pay for it. Add independently measured outage savings or operator hours only if they are real and not already counted elsewhere.

For the earlier one-million-request workload, even a 1% failed-call rate means 10,000 requests need a recovery policy. Before adding automatic failover, decide whether the backup should use the same model via another provider or a different, cheaper model. Record the expected extra call rate, latency ceiling, and quality threshold. The cheapest successful response is not necessarily the correct answer if the application depends on a precise extraction or tool invocation.

Use a small production canary to compare a direct provider call with the gateway path. Track p50 and p95 latency, timeout and retry rates, token counts, model/provider, fallback frequency, cache hit rate, and output-quality checks. A change that reduces spend but raises invalid JSON from 0.4% to 2% can erase the apparent savings through support work or downstream failures. Keep model behavior and transport reliability as separate measurements.

Teams weighing an observability-first approach can also read the LLM observability tools comparison. For retrieval-heavy applications, the vector database comparison covers a separate layer of the same architecture; a gateway will not repair poor retrieval or bad source documents.

Choose by team shape, not feature count

Choose Portkey if a managed configuration layer and visible entry pricing matter more than owning the runtime. Choose LiteLLM if you have a platform owner, need network-level control, and accept the ongoing operations burden. Choose Cloudflare if your app already runs there and the core gateway controls cover your use case; read the log and credit terms before calling the deployment free. Choose Helicone if request debugging, analytics, and provider routing belong in one workflow and its data-retention terms fit your needs.

For a single product, compare the best AI model routing gateway for production against a direct integration plus a thin internal adapter. The internal adapter is often enough when you use one provider, have modest traffic, and need only a fallback. A full gateway earns its complexity when multiple teams share providers, policies, keys, or usage budgets. If those conditions are not present, do not buy future flexibility at the cost of a current dependency.

Frequently Asked Questions

Which model gateway is cheapest for a small team?

Cloudflare lists core features across plans, LiteLLM’s self-hosted open-source gateway has no license fee, and Portkey lists a free Developer plan. The cheapest total depends on logging, hosting, support, traffic, and operator time. Start with one month of measured usage rather than comparing license prices alone.

Is LiteLLM better than a managed AI gateway?

It is better when self-hosting, provider flexibility, and direct control outweigh deployment and on-call costs. A managed gateway is often easier for a small team that needs a working control plane without taking ownership of upgrades and availability. Compare annual labor and infrastructure, not only software fees.

Does a model gateway reduce API costs automatically?

No. It can support caching, routing, and spend controls, but savings depend on repeated traffic, model choice, provider prices, and whether the response remains acceptable. Measure cache hits and model quality together. Retries and verbose logs can raise costs rather than lower them.

Can a gateway fail over between model providers?

Portkey, LiteLLM, Cloudflare AI Gateway, and Helicone document fallback or retry capabilities. Configure the order, status codes, timeout policy, and allowed regions; test the behavior during controlled failures. Failover should never silently send sensitive prompts to an unapproved provider.

Should I use an LLM API gateway if I only call one model?

Not by default. A direct SDK call with secure secrets and provider-side usage controls is simpler. Add a gateway when centralized budgets, reliable failover, cross-app logs, or multiple teams solve a measured operational problem—not because a diagram looks more mature with another box.

The buying rule

The right choice is the least expensive system that meets your reliability, governance, and debugging requirements without adding unowned infrastructure. Compare the monthly subscription with model spend, logs, hosting, and staff time; verify routing behavior with your real prompts; and keep a direct-call fallback plan for the gateway itself. An LLM gateway is justified when its measured savings or avoided failures exceed those costs, not when its feature list is long.

Pricing and product details were checked October 6, 2026. Confirm current terms before purchase: Portkey pricing and gateway configs; LiteLLM pricing and reliability docs; Cloudflare AI Gateway pricing and overview; Helicone pricing and provider routing.

About the author: This article was written by the AI Tool Lab Editorial Team, with 5+ years of paid AI tool testing experience and $200+ monthly subscription spend. All reviews are based on real paid long-term use.

Data statement: All data in this article cites its source and is verifiable. Found an error? Report it via our contact page, we verify within 48 hours.