Back to blog
Sep 28, 2026
11 min read

One Control Point

At my day job, I was asked to be more efficient with AI spend. The honest answer was a semantic router in front of every model, so I pointed my own coding agent at one. It broke. Five bugs later, the fixes are upstream, which means everybody gets them.

Every CIO I work with has a chart.

The chart every CIO has
AI usage climbing, and token spend climbing with it.
?next quartertimetoken spendAI usage
Schematic. No numbers on purpose: every CIO's version has its own.

It shows AI usage climbing, and token spend climbing with it, on infrastructure they are accountable for. The line goes up and to the right, and the people drawing it are the people being asked what it costs next quarter.

At my day job, where I am a Field CTO at AHEAD, that question arrived as a simple request: be more efficient. It was a reasonable request, and it is where this story starts.

My own question was narrower. Can the tools pick the right model and reasoning level for each request, automatically? Not a human choosing from a dropdown, and not a config file somebody updates quarterly. A decision made per request, by something that has read the request.

One control point

That something is a semantic router. It sits in front of your models and reads each prompt before deciding where it goes. One place, three jobs.

Cost. Send each request to the cheapest model that can actually handle it. Many requests do not need your most expensive model.

Availability. Move models in and out behind a stable name. Swap a backend without downtime, because callers address the router, not the model.

Security. Screen prompts before a model sees them. Jailbreak and PII signals at the front door rather than after the fact.

In most enterprises I work with, those are three projects owned by three teams. Putting them at one control point is the interesting part, not any one of them alone.

Cost, availability and security in one place
One router in front of every model, doing three jobs.
Claude CodePOST /v1/messagesApps and agentsPOST /v1/chat/completionsgatewaySemantic Routerone control pointCostcheapest model that can handle the promptAvailabilityswap models behind a stable nameSecurityscreen prompts before a model sees themthis promptModels behind one stable nameSmall modelLarge modelHosted provider

Scroll sideways to see the whole figure.

Schematic. Shows the routing path, not measured latency.

The router is not the only routing decision a request meets. There are three, and it helps to keep them apart.

The stack a request moves through
Three layers, three decisions: the gateway picks the service, Semantic Router picks the model, the inference platform picks the replica.
ClientClaude Code, apps, agents/v1/messages or /v1/chat/completionsGatewayEnvoy 1.35.3 or agentgateway v1.3.1ext_proc gRPC streamSemantic Routerreads the prompt,sets routing headers#4266 changed thisInference platformllm-d, AIBrix, NVIDIA Dynamo,vLLM Production StackModel replicasvLLM model servers or hosted providersWhich service?moves traffic, owns auth and limitsWhich model?from the prompt's meaning and your policyWhich replica?from capacity, locality and health
Source: vLLM Semantic Router docs. Tested on Envoy 1.35.3 and agentgateway v1.3.1.

The project’s own documentation draws the same three-way split, and it is worth borrowing even if you never use this software. Its advice on the first layer is the practical kind: use the gateway your platform already operates. You are not supposed to adopt a new data plane to get smarter routing.

I routed my own work through it, and it broke

I did not evaluate this by reading about it. I put my own work behind it.

My coding agent is Claude Code. I pointed it at vLLM Semantic Router, and the Router at vLLM running on a GH200 machine in my lab. Personal hardware, because my work laptop was not powerful enough for the models I wanted behind it. That detail matters later.

Then I used it for real work, on the path that gets less testing. Claude Code speaks the Anthropic Messages API, POST /v1/messages, not OpenAI’s Chat Completions. Most routing examples and tests use Chat Completions. Two of the bugs below, and the open one at the end of this post, are on it.

The path my own work took
Where each bug sat, from Claude Code to vLLM, and the fix for each. One is still open.
the Anthropic Messages path: less traffic than Chat CompletionsPOST /v1/messagesClaude CodeGatewayEnvoy or agentgatewaySemantic RoutervLLM on the GH200arm641stop_sequence missing instreaming /v1/messagesvLLM #55324 → #55325 (my fix)12rejected responses withoutstop_sequence#3500 → #394423buffered response diagnosticsnever reported#4118 → #412534arm64 builds failed,caught on my Mac#4132 → #412545full duplex droppedthe routing decision#3892 → #4266 (my fix)5open: null stop fields innon-streaming /v1/messagesvLLM #56376 (my patch)my codefixedopen
REAL. Five fixed bugs and one open, hit running my own work on this path, Sep 4 to Sep 28, 2026.
Show the data
pinbugissuefixfix author
1stop_sequence missing in streaming /v1/messagesvLLM #55324vLLM #55325colinmcnamara
2Router rejected responses without stop_sequence#3500#3944Xunzhuo
3buffered response diagnostics never reported#4118#4125Xunzhuo
4arm64 builds failed on libgomp.so.1#4132#4125Xunzhuo
5full duplex dropped the routing decision#3892#4266colinmcnamara
opennull stop fields in non-streaming /v1/messagesvLLM #56371vLLM #56376colinmcnamara (open)
BugIssue and fixWho fixed it
stop_sequence missing from message_delta in streaming /v1/messagesvLLM #55324 → #55325My fix, merged by Robert Shaw, a vLLM Lead Maintainer
Router rejected responses that lacked stop_sequence#3500 → #3944Xunzhuo, using the lenient-with-diagnostic design I recommended
Buffered response diagnostics never reported#4118 → #4125Xunzhuo
arm64 builds failed on libgomp.so.1#4132 → #4125Xunzhuo. Caught on my arm64 Mac; the GH200 is arm64 too
Full duplex dropped the routing decision#3892 → #4266My fix, merged 2026-09-28

I wrote up the first of those in The Router Was Right. Short version: a field was missing from a stream, I wrote the wrong fix first, and the schema settled it.

The three in the middle were fixed by the maintainers and shipped in the v0.4 “Hermes” release on 2026-09-27. To be precise about my part: three bugs I reported were fixed in v0.4, and the maintainers wrote that code. My first merged Semantic Router code is #4266, which landed after the release.

The fix I am proudest of

The plain-English version of #4266: the routing decision was made, and then it was thrown away.

What broke

In full duplex mode (FULL_DUPLEX_STREAMED in Envoy, FullDuplexStreamed in agentgateway), the gateway streams the request body to the Router without waiting for replies. That is the point of the mode, and it is efficient. Here is what main did with it.

What main did: answer first, route later
Before #4266: the Router answered the headers at once, and its routing decision was dropped.
ClientOpenAI-style appGatewayEnvoy / agentgatewaySemantic Routerpkg/extprocDefaultbackendSelectedbackendPOST /v1/chat/completionsAuthorization: client credentialrequest headersheader reply, sent at oncenothing routed yetbody chunks, end_streamclassify the prompt,pick a modelbody replyrouting headers droppedroute-cache clear ignored/v1/chat/completions, client credentialwrong backend
REAL. Observed on Envoy 1.35.3 and agentgateway v1.3.1 before the fix.

The Router replied to the request headers immediately, before it had read a byte of the prompt. It then read the body, picked the model at the end of it, and set the routing headers: x-selected-model, the provider path, the provider credential.

But a body reply cannot change headers in this mode. The header stage was already over. Envoy said so in its own log line:

NOT clearing route cache. No header mutation in the response

The result on main was not a subtle degradation. Requests went to the default backend carrying the client’s own credential, even when the body was small enough to arrive in a single write. A request with trailers got a 500 from Envoy with Spurious response message. Through agentgateway, the backend rejected the body as invalid JSON and returned a 400.

What fixed it

The design that fixed it was not mine. YAMAMOTO Takashi, who reported #3892, proposed it in the issue: hold the header reply until the body has been routed. Then send it, carrying the routing headers and the route cache clear, then the body reply, then a trailers reply if the request has trailers.

What the fix does: route first, then answer
With #4266: the Router holds its header reply until the body is routed.
ClientOpenAI-style appGatewayEnvoy / agentgatewaySemantic Routerpkg/extprocDefaultbackendSelectedbackendPOST /v1/chat/completionsAuthorization: client credentialrequest headersheader reply heldbody chunks, end_streamclassify the prompt,pick a modelheader reply+ routing headers + route-cache clearbody reply, then trailers replytrailers reply only when the request has trailers/provider/v1/chat/completions, provider credentialrouted
REAL. Observed on Envoy 1.35.3 and agentgateway v1.3.1 with the fix.

What makes that legitimate rather than a hack is that Envoy’s specification explicitly permits it: a server may wait for the body before answering the headers. The mode was never the problem. The Router answered too early to act on what it learned from the body.

With the fix, the selected backend, the provider path and the provider credential all arrive correctly, on both gateways. The diagrams show a Chat Completions request from the project’s E2E tests; the fix covers any request streamed in full duplex, /v1/messages included.

My favorite detail is in main’s own source: a comment noted that the header changes “have no effect,” and the code then threw them away. It was a known limitation, written down in the code. The fix moves those changes onto the held reply, where they do have effect.

Choosing a mode

Here is the whole trade-off:

Four ways to hand the request body to the Router
Every mode gave up memory or correctness until #4266 moved full duplex into the corner that gives up neither.
yesnoCorrect: the Router's choice sets the routenoyesEfficient: the gateway does not buffer the whole bodyBUFFEREDcorrect, but the gateway holds every prompt in memoryFull duplex with #4266efficient and correctslow and wrongFull duplex, before #4266the routing decision was always droppedSTREAMEDsometimesthe decision can arrive after the headers go upstreamHeaders onlynever reads the prompt#4266
Schematic. Positions are qualitative, from the Envoy ext_proc processing modes.

Two operational notes came out of review and are now in the documentation:

  • On agentgateway, routing on Router-set headers requires the ExtProc policy in traffic.phase: PreRouting, because the default PostRouting picks the route before the Router speaks.
  • Keep failure_mode_allow: false for full duplex. If the stream fails while the header reply is held, failing open would forward the client’s original headers, which is the situation you were trying to prevent.

The people who fixed it

This is the part I want to be specific about, because “the community fixed it” is not a credit, it is a shrug.

Xunzhuo Liu (@Xunzhuo, AMD) is a Semantic Router maintainer and, in practice, its technical lead: 463 merged PRs, and a co-lead of Data Plane & Networking. Xunzhuo accepted #3500 within minutes and built the design I had recommended, then fixed #4118 and #4132 for v0.4. Xunzhuo also asked for the end-to-end (E2E) coverage, and reopened #3892 and marked it in progress for me, which is why I got to write the fix at all.

Stefan Wang (@1fanwang, LinkedIn) trimmed my PR description to the project’s style, specified the E2E test exactly, and gave the first approval: “LGTM. Thanks for the quick turnaround.” Stefan is a committer, one of four Data Plane & Networking co-leads and a lead of Agentic & Context.

Abhinav Mahajan (@abhinav-m22, Barclays), a committer, asked whether a held reply survives a stream error. It was the best question on the PR, and the answer found a real gap. I fixed it, posted the evidence, and Abhinav approved.

Wilson Wu (@wilsonwu, DaoCloud), a committer and a lead of Developer Experience & Ecosystem, approved.

YAMAMOTO Takashi (@yamt) reported #3892, diagnosed it, proposed the hold-the-header-reply design, and approved the implementation with “lgtm thank you.”

Binbin Zhang (@Bevisy, China Telecom) updated the branch so it could merge. Binbin is a committer, a lead in three workgroups and a Kata Containers maintainer.

On the vLLM side, #55325 was my own patch, 29 lines across two files. Robert Shaw (@robertgshaw2-redhat, Red Hat), one of vLLM’s five Lead Maintainers, merged it, and Sylvester Kaczmarek (@sylvesterkaczmarek) reviewed and approved it after checking CI.

Every bug, from filing to fix
Twenty-four days, five fixes, drawn to scale in UTC.
v0.4.0 Hermes releasedvLLM #55324 → #55325 (my fix)6.5 days#3500 → #394415 daysaccepted in 6 minutes#4118 → #4125about a day#4132 → #4125about 12 hours#3892 → #4266 (my fix)1.9 daysSep 4Sep 8Sep 12Sep 16Sep 20Sep 24Sep 282026, UTC
REAL. Timestamps read from the GitHub API on 2026-09-28. Durations are computed from those timestamps, never typed. Blue bars are my own code.
Show the data
bugfixclock starts (UTC)fixed (UTC)elapsedfix author
vLLM #55324vLLM #553252026-09-04 11:15:242026-09-10 22:21:226.5 dayscolinmcnamara
#3500#39442026-09-04 23:56:142026-09-20 02:35:5015 daysXunzhuo
#4118#41252026-09-24 08:42:302026-09-25 08:59:32about a dayXunzhuo
#4132#41252026-09-24 21:19:442026-09-25 08:59:32about 12 hoursXunzhuo
#3892#42662026-09-26 13:01:00 (claimed)2026-09-28 11:18:381.9 dayscolinmcnamara

There is a lesson in that timeline, and it is a lesson rather than a grievance. This team moves on reports, which is the whole reason to file them here. It is also why you should claim an issue first if you want to write the code yourself. I claimed #3892 at 13:01Z and opened #4266 at 23:55Z the same day. Claim to merge was 1.9 days.

How to do this yourself

Six things, in the order they mattered.

1. Prove the bug on real software before you claim it. #4266 was demonstrated against a real Envoy 1.35.3 and a real agentgateway v1.3.1, not a mock.

2. Write tests that fail on main and pass with the fix. The routing tests in #4266 did. Then plant deliberately wrong fixes and confirm the tests catch them.

About six lines of tests for every line of fix
Drawn to scale: 187 lines of fix, 1,157 lines of tests.
The fixRouter code: 187 linesThe tests1,157 linesunit tests 407E2E tests 707test backend 4323 of 23 planted wrong fixes caught#4266: 24 files, +1,359 −27 lines
REAL. Line counts from git diff --numstat on PR #4266.
Show the data
partkindlines added
Router codefix187
unit teststest407
E2E teststest707
test backendtest43
docsother11
toolingother4
total1,359

3. Have several AI models review the work independently, then check their findings yourself. They disagree with each other, which is useful. They are also confidently wrong on occasion, which is why the checking is not optional.

4. Own the reviewer’s catch in public. Prove exactly where the reviewer was right, fix it, and post the evidence in the thread. Do not quietly patch it.

5. Build exactly what the maintainers ask for. Stefan specified the E2E test. The E2E test matched the specification. That is most of what “easy to review” means.

6. Credit everyone by name.

Full circle

I needed this capability for my own work. Stated plainly, the capability is this: optimize token spend in real time, and move models in and out transparently without taking services down. Because the fixes went upstream instead of into a private fork, it is not mine alone.

Leaders watching that chart have a lever on spend worth measuring. AHEAD’s engineers, sellers and project managers can use the same tools I use. Everyone else who runs this stack gets it for free, including the people who will hit that full duplex bug next month and never know it was there.

That is the argument for working upstream rather than around it.

A private fix helps one company, and that company maintains it forever. An upstream fix helps everyone, and somebody else maintains it with you.

From a business question to an upstream fix
Upstream, not a private fork. Measuring the savings is the next turn of the loop.
The questionAI spend is rising.Can we be more efficient?One control pointroute each promptto the right modelMy own work through itClaude Code, the Router,vLLM on a GH200It brokefive bugs along the wayFixed upstreamwith the vLLM and SemanticRouter maintainersEveryone can use itAHEAD’s teams andthe whole communitynext: measure the savingsUpstream, not a private fork
Schematic.

What is next

Measure it. I have not measured the savings. This is where the work laptop comes back in. The before-and-after token number, taken there, is the actual answer to the question that started this, and I have not run it yet. I would rather publish the mechanism now and the measurement when it exists than imply a number I do not have.

Bugs in flight. My patch for null stop fields in non-streaming /v1/messages, vLLM #56376, is open and waiting on review. And the Router’s finish-reason allowlist rejects values that vLLM itself can return. That one is next, once I have captured it on my own hardware.

Research directions, and I mean research rather than product. The control point is the natural place to screen for prompt and logic injection, because it is the one place every request passes through. AI honeypots are the other thread I keep pulling on. Neither is something I am shipping. Both are things I think the control point makes newly possible.

The chart still goes up. The difference is that there is now somewhere to stand.

Let's Build AI That Works

Ready to implement these ideas in your organization?