How Running an AI Hub Can Pay for Itself
Joe Poyner is a security engineer based in Oklahoma City with more than 15 years in the EDR world across Dell, Palo Alto Networks, Endgame, and SentinelOne, where he founded the ThreatOps threat-hunting championship and pitted human analysts against AI on stage before it was cool. He’s now a Senior Solutions Engineer at JetStream Security, working on governing enterprise AI.
A few weeks back my wife and I sat down at the kitchen table for our monthly budget review. You know the drill. We got to the streaming line and she said what she always says, that we could drop HBO (MAX) and just do Paramount+ because it’s cheaper.
And I said what I always say. It’s cheaper because it’s half the catalog.
We went back and forth on it for a while, and somewhere in there it hit me. This is the exact same conversation every engineering team in America is having about AI right now. Claude is HBO (MAX). It’s the best and it’s priced like it knows it. Alternatives are out there, but nobody wants to give up the extra juice for a lower price.
Meet the AI Hub
An AI hub, also known as an LLM proxy or a gateway, sits between your tools and your models. Think LiteLLM, or Portkey, or one of those Kong-style AI gateways. It parks itself in front of all your providers and routes each request off to whatever model you picked for that job. Keep Claude Fable 5 and Opus 4.8 up on the shelf for the work that actually needs them. The everyday stuff, your boilerplate and your tests and your docs and the fortieth summary of the day, that goes to something like Z.ai’s GLM-5.2 for a fraction of the price. One config file gives you the best of both worlds. That’s one control point sitting in front of every model you call. Full warning, I build governance for these at JetStream, so I’ll keep the how-to vendor-neutral and tell you where we fit at the end.
People are already doing this
Don’t take my word for it. Go look at OpenRouter, where a whole crowd of people route Claude Code (that’s Anthropic’s own coding agent) through a proxy. The second most-used model over there is GLM 5.2, at 811 billion tokens in the 30 days to July 7, sitting right behind Claude Opus 4.8 (OpenRouter, Claude Code app page, 30-day token volume, retrieved July 7, 2026). That only counts the folks who went and deliberately pointed Claude Code away from Anthropic, so it’s a cost-conscious bunch to start with. But that’s kind of the whole point. These are people voting with their wallets every single day, and they keep splitting the work up: cheap model for most of it and the frontier model for the rest. If the swap fell apart on real work they’d have quit by now. Demand got hot enough that Zhipu, the outfit that makes GLM, had to throttle new signups down to a fifth of normal back in January of 2026 because they flat ran out of compute (Bloomberg, January 21, 2026, corroborated by SCMP). The cheap lane is popular enough that they had to ration it.
Is GLM-5.2 as good as Claude? Nope, and I’m not going to sit here and pretend it is. On the Artificial Analysis coding index it comes in at 51.1, against Opus 4.8’s 55.7 and Fable 5’s 59.9 (Artificial Analysis Intelligence Index, Programming, retrieved July 7, 2026) Often you won’t even notice the difference, and when you would, the security-sensitive stuff, that’s exactly the quarter the router keeps up on the frontier model anyway.
Let’s put real money on it
Say you’ve got a company. We’ll call it Blank Label Software, with 40 engineers, made up but the bill is real enough. Everybody’s coding with an agent all day long, and between the lot of them they chew through something like 2 billion input tokens and 400 million output tokens a month. At list price, per million tokens, Fable 5 runs you $10 and $50, Opus 4.8 is $5 and $25, and GLM-5.2 is $1.40 and $4.40 (Anthropic and Z.ai pricing pages, retrieved July 7, 2026).
Run every bit of that on Opus and you’re looking at (2B × $5/M) + (400M × $25/M) = $10,000 + $10,000 = $20,000 a month. Call it $240,000 a year just to keep the lights on. Now split it 75/25. Three-quarters of the work goes over to GLM-5.2 and the hard quarter stays on Opus. The GLM side runs about $3,420 and the Opus side about $5,000, which puts you at $8,420 a month. That’s better than $11,500 a month landing back in your budget, pushing $140,000 over a year.
The 25/75 Swing
$20k down to ~$8.5k/mo
$20k down to ~$8.5k/mo
About $11,600 a month back in the budget, near $140,000 a year. Hub cost to run it: roughly $100 a month plus a week to stand up.
Now hold on a second before you go quoting that number around the office. It’s assuming you buy Claude by the token. If your team’s on flat-rate Max seats, or you cut yourself a volume deal on Bedrock, that gap gets smaller, so run the math against your own bill. The hub itself costs you about a $100 a month to host, plus a week of somebody’s time to stand it up, so it pays off its own build cost before the first invoice even clears.
A couple honest caveats while I’m at it. The cheaper models tend to retry more and ramble more, so they’ll burn extra tokens getting to the same place a frontier model would. Measure what a finished task actually costs you, not just the per-token sticker. And caching cuts both directions. Anthropic knocks around 90% off cached input (Anthropic, prompt caching pricing), so on an agent loop that’s replaying the same context over and over, cached Opus can actually come in under uncached GLM. The output tokens still lean GLM’s way, just not by as much as that little table makes it look. The direction’s right. The exact multiple, that’s on you to measure.
Okay, but foreign models? You’re the security guy.
I am. And here’s the way I think it through. Whose servers the model runs on, and how the model actually behaves are two completely different questions. Only the first one is really about where the model originated.
Start with the servers, because that’s the real one. It’s what the regulators actually went after, the government-device bans, Italy pulling the app off the shelf. And it all comes down to the data path. Your prompts land on servers sitting inside another country (DeepSeek Privacy Policy; Al Jazeera February 6, 2025), where its National Intelligence Law can lean on the company to hand your stuff over (National Intelligence Law of the PRC, 2017, Article 7). But take that same open-weight file and run it on Western infrastructure and the whole problem just evaporates. All three of the big clouds will host GLM for you managed right now. Bedrock and Google Vertex have GLM-5 (AWS Bedrock and Google Vertex AI documentation, retrieved July 7, 2026), and Azure’s got GLM-5.1 running through its Fireworks setup (Microsoft Learn, Fireworks models on Microsoft Foundry, updated June 30, 2026). Sort out the data path and you’ve handled the biggest source of security red flags. One caveat if you are regulated. A managed route can still carry its own compliance gaps. Azure’s GLM through Fireworks, for instance, sits outside the EU Data Boundary and has not cleared FedRAMP, so measure the specific route against your own bar.
Then there’s the thing everybody actually loses sleep over, which is whether you can trust the model itself. Honest answer? Not all the way. But here’s the kicker, you can’t fully trust any of them, and that’s the whole point. Those scary stories about foreign models are true. They write insecure code, they can be jailbroken, they’ll happily cheat a benchmark if cheating’s easier than doing the work. Thing is, all of that is just as true of Claude and GPT. Veracode ran a scan across more than a hundred models and found that close to half of all AI-written code ships with a security hole sitting in it, and it did not much matter whose logo was on the model (Veracode, GenAI Code Security Report, 2025 and Spring 2026 update) Anthropic and OpenAI have both put out research catching their own models fudging tests to fake a passing grade (Anthropic, “Sycophancy to Subterfuge,” June 2024; OpenAI, arXiv:2503.11926, March 2025), which is the exact same thing Z.ai owned up to about GLM-5.2 (Z.ai, “GLM-5.2: Built for Long-Horizon Tasks”)
And a company grading its own homework is worth about what you’d figure, so your real control is the test suite, not the vendor’s changelog. On the one honest head-to-head cheating test somebody actually bothered to run, it was OpenAI’s models that cheated the most, with Claude and Qwen coming out about even, and the whole thing tracked how capable the model was and not where it was born (ImpossibleBench, arXiv:2510.20270, 2025).
I’ll give you one fair exception. The open foreign weights have tended to ship with thinner safety training than Claude does. Cisco managed to jailbreak DeepSeek-R1 on every single prompt they threw at it, where Claude held the line about two-thirds of the time (Cisco and University of Pennsylvania, January 2025). That’s a real gap, though it’s been closing. Bottom line, “the model might do something dumb” is a reason to keep a human and a test suite between any model and production. It is not a reason to be scared of the cheap lane.
And the security work here isn’t anything exotic. It’s the same agent hygiene you ought to be running on Claude already. Give the tools the least access they can get away with, sandbox anything that’s touching input you don’t control, and never once hand a model a shell, and your secrets, and the open internet all at the same time. A booby-trapped README can hijack an agent into running whatever the attacker wants, sure, but it’ll pull that same trick on Claude Code running Claude just as fast. That’s an agent problem, not a State-owned-model problem (OWASP Top 10 for LLM Applications, 2025; Simon Willison, “The Lethal Trifecta for AI Agents,” June 2025) . Do the work one time and it covers every model you send through. And notice where that work happens. At the same control point you stood up to save money. One place in front of your models, pulling double duty.
Where to actually start
Now, the analogy I’ve been leaning on hides the hard part. Figuring out which requests are the easy three-quarters and which are the hard quarter. That’s the actual engineering, and it’s on you. Get it wrong and you’re either paying Opus prices to write boilerplate or you’re shoving the cheap model at something it’s going to botch. So start rough. Route by repo and by the kind of task. Docs and tests and scaffolding go to GLM, and the payments service and anything security-sensitive stays on Opus. Then you tighten it up over time, once your telemetry starts showing you where the cheap lane is falling down. And pilot the thing before you commit to it. Mirror a slice of your real traffic over to the cheap model, diff the output against what you’d normally get, and keep a close eye on the tool calls especially, because the hub is quietly translating Anthropic’s format over to OpenAI’s behind the scenes and tool-calling is the first place that starts to fray. A week of those tests will tell you your real split, and your real savings, a whole lot better than any table I can print here.
One last thing, and here’s where I show my cards. That control point, the one place you route for cost and lock down for safety, is what we build at JetStream. Ours is JetStream AI Hub™, a governed take on the same gateway this whole piece is about. This week we shipped a JetStream Verified MCP™ catalog that runs alongside it. I kept the how-to vendor-neutral on purpose, so stand it up with LiteLLM and be happy. But if you’d rather it come governed, now you know where we sit.
Still thinking “hell no”?
Fair enough. My own gut said the same thing. “Run a Chinese model in production” sounds like a flat no, and for some shops it genuinely is. But hang on a second before you close the tab. Airbnb, a company whose entire business runs on strangers trusting strangers, is running its AI customer-service agent on thirteen different models, and one of them is Alibaba’s Qwen. Their CEO’s reason? He said it’s “very good, fast and cheap” (Forbes, May 21, 2026; Chesky via Bloomberg, October 2025). Then Congress went and opened an inquiry, and Brian Chesky’s answer was basically this whole article boiled down to one line. The weights are open-source, and “we are not providing data to any Chinese companies” (Bloomberg, May 20, 2026). A Fortune 500 CEO with a House committee breathing down his neck landed on the exact same split I did at my kitchen table. The data path is the risk, and you can engineer your way around it.
The cherry on top
The best argument for running a hub isn’t even in that pricing table. It’s that the pricing table has a shelf life of about a month.
DeepSeek slashed its flagship price by 75% in May and then made the cut permanent (Engadget, May 23, 2026). There’s a new model family landing every few weeks now. And don’t go kidding yourself that Anthropic is going to sit there and take it. Fable 6 could show up with a steep discount to win everybody back, and just like that the “cheap” models aren’t the cheap ones anymore.
If your tooling is welded to one vendor, every single one of those moves turns into a migration project. Run a hub and it’s a config change instead. Swap the model’s name, watch out for the gotchas (that Bedrock GLM-5 route is 200k of context and not the 1M you got used to) and get on with your day. And keep one more thing in your back pocket while you’re at it. A managed foreign model is one export rule away from getting yanked off Bedrock overnight, so know how you’d self-host the open weights before you ever actually need to.
Don’t go marrying a vendor in a market moving this fast. Keep the remote in your own hand.
My wife was right that Paramount+ is cheaper. I was right that HBO (MAX) is better. The hub is the thing streaming never lets us have. Keep the good stuff for movie night, run everything else on the cheap plan, and switch the second a better deal comes along.
LATE BREAKING NEWS Postscript: July 9
Well, it happened. I pulled those prices for this article on July 7. On July 9, two days later, OpenAI went and shipped GPT-5.6, three sizes of it no less (Sol, Terra, and Luna), and the big one, Sol, runs $5 in and $30 out per million tokens (OpenAI, “GPT-5.6,” July 9, 2026). That’s half of Fable 5’s input sticker on a model OpenAI swears is frontier class, and they’ve got it 2.8 points higher than Fable 5 on the coding agent index (a different Artificial Analysis benchmark from the coding index cited earlier), 80 to 77.2, using less than half the output tokens (OpenAI, “GPT-5.6,” July 9, 2026). Now, a company grading its own homework, we already covered what that’s worth, and fair’s fair, OpenAI’s own table shows it trailing Fable 5 by better than 15 points on SWE-Bench Pro, and their fine print admits the cost and latency numbers are offline simulations, their words, “real-world results may vary substantially”(OpenAI, “GPT-5.6,” July 9, 2026) But the independent folks over at Artificial Analysis put it one single point behind Fable 5 on their intelligence index at about a third of the cost, and measured a real coding task coming in around 40% cheaper than Fable 5 and 10% cheaper than Opus (Artificial Analysis, “GPT-5.6 has landed,” July 2026). One point shy of the best model on the market. A third of the cost. That’s not a budget model, that’s a price war.
And notice something while you’re at it. Sol’s output sticker is actually higher than Opus’s, $30 against $25, but it finishes the task cheaper because it doesn’t ramble getting there. Which is the exact thing I told you earlier, measure the finished task and not the per-token sticker. So here we are. My whole pricing table up there was stale before this thing even got published, just like I said it would be. If your tooling is welded to one vendor, GPT-5.6 is a migration project you have to go justify to somebody. If you’re running a hub, it’s a config line and a week of mirrored traffic to see if the hype holds up on your code. Keep the remote in your own hand.