All articles
forge ai-agents multi-provider 5 min read

Stop asking for the expensive model. Let the gateway decide.

Forge's AI gateway now resolves a tier name — top, mid, or low — into the best model that connector can actually serve, at the moment of the request. When a cheaper model is good enough, it gets picked automatically. When the vendor ships a new flagship, every caller upgrades without a code change.

by Vitaly Nikitin

Here is the pattern every team running agents at scale eventually falls into: you configure the most capable model you can access, because capable is what you need for the hard tasks. Then you watch your usage bills arrive and realise that about eighty percent of your requests — the summarisations, the reformatting, the quick lookups — didn’t need the expensive model at all. They needed a model.

The problem isn’t that the expensive model is bad. It’s that the caller picked it once, up front, for the hardest thing it might ever do — and that choice stuck everywhere.

This week, Forge shipped a different approach.

Tiers, not names

Instead of naming a model in your workflow configuration, you now name a tier: forge/anthropic/top, forge/openai/mid, forge/google/low. The gateway keeps a live, auto-reconciled catalogue of every provider’s models and classifies each one into a slot — top, mid, or low — within its family. When your request arrives, the gateway resolves the tier into whichever model that connector can actually serve right now.

The practical consequence: your agent asks for forge/anthropic/mid and gets the best mid-tier Anthropic model available at that moment. You don’t hardcode a model name. You don’t maintain a list. You describe the capability level you need, and the gateway handles the rest.

There’s also @prev — a modifier that pins to the previous generation of a tier slot. Useful when you want cost predictability and aren’t chasing the absolute frontier.

Why this matters right now

Runaway AI spend became a mainstream worry in September and October 2026. Writing on October 3rd, Simon Willison argued that hard budget caps need to be the default on pretty much every pay-by-usage service, because the agents we now point at those services make it trivially easy to spin up something expensive: “Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.”

A cap is the right backstop, but it only tells you when to stop — not how to spend less in the first place. The Weave Router launch on Hacker News (119 points, September 30th) attacked that half of the problem with concrete numbers: an open-source routing model that matched GPT-6 Astra’s pass rates on Terminal Bench 4.0 and SWE Atlas while costing 52% and 54% as much, simply by matching the capability tier to the task difficulty.

Both projects are pointing at the same root cause: when callers hardcode model names, the platform can’t help them make better tradeoffs. The solution isn’t a smarter caller. It’s a smarter gateway.

What the gateway resolves

When a request arrives with a tier alias, the gateway does three things:

Classify. Every model in the connected provider catalogues is classified into a family (anthropic, openai, google, and others) and a slot (top, mid, low) according to a live mapping table. The classification is automatic when a new model joins the catalogue; operators can override individual models and switch any model off entirely so it never gets served on any path.

Resolve. The tier alias maps to the best available model in that slot for that connector, right now. The decision is logged in the response headers — X-Forge-Requested-Model, X-Forge-Served-Model, and X-Forge-Model-Resolution — so you can audit what the gateway actually served and why.

Degrade gracefully. If the requested slot has no available model, degradation goes downward within the family. A top request that can’t be served doesn’t jump to a different family; it steps to mid in the same family. Family aliases never cross families, which matters because different families have genuinely different capability profiles and you don’t want your forge/anthropic/top workload silently rerouting to a different vendor’s model.

Operators get a control surface

This isn’t just a routing rule in a config file. The gateway ships with a Model mapping console — a family × slot grid that shows you exactly how every model in your connected catalogues is classified, where the overrides are, and which models have been switched off. Unclassified models (new arrivals that haven’t been slotted yet) appear in a review queue so nothing disappears silently.

This is where the operational story gets interesting. When GPT-7 ships and OpenAI updates their catalogue, the gateway reconciles automatically. Every request pointing at forge/openai/top moves to the new model without any action from you. The console shows you the change happened; your agents don’t need to know about it.

The broader picture

This tier system extends Forge’s existing multi-provider routing : the earlier capability made sure your agents could reach multiple providers and fail over automatically. The tier system adds a second axis — within a provider, across a price and capability spectrum — and puts the routing decision where it belongs: in the gateway, with visibility, rather than scattered across every workflow configuration file.

If you’re running agent pipelines today and haven’t revisited how you’re picking models, this is the moment. The pricing landscape has changed fast enough that decisions made six months ago are probably leaving real money on the table.

Take a look at Forge’s multi-provider routing features , or talk to us if you want to walk through what tier-based routing would look like for your workload.