Every CIO I work with has a chart.
It shows AI usage climbing, and token spend climbing with it, on infrastructure they are accountable for. The line goes up and to the right, and the people drawing it are the people being asked what it costs next quarter.
At my day job, where I am a Field CTO at AHEAD, that question arrived as a simple request: be more efficient. It was a reasonable request, and it is where this story starts.
My own question was narrower. Can the tools pick the right model and reasoning level for each request, automatically? Not a human choosing from a dropdown, and not a config file somebody updates quarterly. A decision made per request, by something that has read the request.
One control point
That something is a semantic router. It sits in front of your models and reads each prompt before deciding where it goes. One place, three jobs.
Cost. Send each request to the cheapest model that can actually handle it. Many requests do not need your most expensive model.
Availability. Move models in and out behind a stable name. Swap a backend without downtime, because callers address the router, not the model.
Security. Screen prompts before a model sees them. Jailbreak and PII signals at the front door rather than after the fact.
In most enterprises I work with, those are three projects owned by three teams. Putting them at one control point is the interesting part, not any one of them alone.
Scroll sideways to see the whole figure.
The router is not the only routing decision a request meets. There are three, and it helps to keep them apart.
The project’s own documentation draws the same three-way split, and it is worth borrowing even if you never use this software. Its advice on the first layer is the practical kind: use the gateway your platform already operates. You are not supposed to adopt a new data plane to get smarter routing.
I routed my own work through it, and it broke
I did not evaluate this by reading about it. I put my own work behind it.
My coding agent is Claude Code. I pointed it at vLLM Semantic Router, and the Router at vLLM running on a GH200 machine in my lab. Personal hardware, because my work laptop was not powerful enough for the models I wanted behind it. That detail matters later.
Then I used it for real work, on the path that gets less testing. Claude Code speaks the Anthropic Messages API, POST /v1/messages, not OpenAI’s Chat Completions. Most routing examples and tests use Chat Completions. Two of the bugs below, and the open one at the end of this post, are on it.
Show the data
| pin | bug | issue | fix | fix author |
|---|---|---|---|---|
| 1 | stop_sequence missing in streaming /v1/messages | vLLM #55324 | vLLM #55325 | colinmcnamara |
| 2 | Router rejected responses without stop_sequence | #3500 | #3944 | Xunzhuo |
| 3 | buffered response diagnostics never reported | #4118 | #4125 | Xunzhuo |
| 4 | arm64 builds failed on libgomp.so.1 | #4132 | #4125 | Xunzhuo |
| 5 | full duplex dropped the routing decision | #3892 | #4266 | colinmcnamara |
| open | null stop fields in non-streaming /v1/messages | vLLM #56371 | vLLM #56376 | colinmcnamara (open) |
| Bug | Issue and fix | Who fixed it |
|---|---|---|
stop_sequence missing from message_delta in streaming /v1/messages | vLLM #55324 → #55325 | My fix, merged by Robert Shaw, a vLLM Lead Maintainer |
Router rejected responses that lacked stop_sequence | #3500 → #3944 | Xunzhuo, using the lenient-with-diagnostic design I recommended |
| Buffered response diagnostics never reported | #4118 → #4125 | Xunzhuo |
arm64 builds failed on libgomp.so.1 | #4132 → #4125 | Xunzhuo. Caught on my arm64 Mac; the GH200 is arm64 too |
| Full duplex dropped the routing decision | #3892 → #4266 | My fix, merged 2026-09-28 |
I wrote up the first of those in The Router Was Right. Short version: a field was missing from a stream, I wrote the wrong fix first, and the schema settled it.
The three in the middle were fixed by the maintainers and shipped in the v0.4 “Hermes” release on 2026-09-27. To be precise about my part: three bugs I reported were fixed in v0.4, and the maintainers wrote that code. My first merged Semantic Router code is #4266, which landed after the release.
The fix I am proudest of
The plain-English version of #4266: the routing decision was made, and then it was thrown away.
What broke
In full duplex mode (FULL_DUPLEX_STREAMED in Envoy, FullDuplexStreamed in agentgateway), the gateway streams the request body to the Router without waiting for replies. That is the point of the mode, and it is efficient. Here is what main did with it.
The Router replied to the request headers immediately, before it had read a byte of the prompt. It then read the body, picked the model at the end of it, and set the routing headers: x-selected-model, the provider path, the provider credential.
But a body reply cannot change headers in this mode. The header stage was already over. Envoy said so in its own log line:
NOT clearing route cache. No header mutation in the response
The result on main was not a subtle degradation. Requests went to the default backend carrying the client’s own credential, even when the body was small enough to arrive in a single write. A request with trailers got a 500 from Envoy with Spurious response message. Through agentgateway, the backend rejected the body as invalid JSON and returned a 400.
What fixed it
The design that fixed it was not mine. YAMAMOTO Takashi, who reported #3892, proposed it in the issue: hold the header reply until the body has been routed. Then send it, carrying the routing headers and the route cache clear, then the body reply, then a trailers reply if the request has trailers.
What makes that legitimate rather than a hack is that Envoy’s specification explicitly permits it: a server may wait for the body before answering the headers. The mode was never the problem. The Router answered too early to act on what it learned from the body.
With the fix, the selected backend, the provider path and the provider credential all arrive correctly, on both gateways. The diagrams show a Chat Completions request from the project’s E2E tests; the fix covers any request streamed in full duplex, /v1/messages included.
My favorite detail is in main’s own source: a comment noted that the header changes “have no effect,” and the code then threw them away. It was a known limitation, written down in the code. The fix moves those changes onto the held reply, where they do have effect.
Choosing a mode
Here is the whole trade-off:
Two operational notes came out of review and are now in the documentation:
- On agentgateway, routing on Router-set headers requires the ExtProc policy in
traffic.phase: PreRouting, because the defaultPostRoutingpicks the route before the Router speaks. - Keep
failure_mode_allow: falsefor full duplex. If the stream fails while the header reply is held, failing open would forward the client’s original headers, which is the situation you were trying to prevent.
The people who fixed it
This is the part I want to be specific about, because “the community fixed it” is not a credit, it is a shrug.
Xunzhuo Liu (@Xunzhuo, AMD) is a Semantic Router maintainer and, in practice, its technical lead: 463 merged PRs, and a co-lead of Data Plane & Networking. Xunzhuo accepted #3500 within minutes and built the design I had recommended, then fixed #4118 and #4132 for v0.4. Xunzhuo also asked for the end-to-end (E2E) coverage, and reopened #3892 and marked it in progress for me, which is why I got to write the fix at all.
Stefan Wang (@1fanwang, LinkedIn) trimmed my PR description to the project’s style, specified the E2E test exactly, and gave the first approval: “LGTM. Thanks for the quick turnaround.” Stefan is a committer, one of four Data Plane & Networking co-leads and a lead of Agentic & Context.
Abhinav Mahajan (@abhinav-m22, Barclays), a committer, asked whether a held reply survives a stream error. It was the best question on the PR, and the answer found a real gap. I fixed it, posted the evidence, and Abhinav approved.
Wilson Wu (@wilsonwu, DaoCloud), a committer and a lead of Developer Experience & Ecosystem, approved.
YAMAMOTO Takashi (@yamt) reported #3892, diagnosed it, proposed the hold-the-header-reply design, and approved the implementation with “lgtm thank you.”
Binbin Zhang (@Bevisy, China Telecom) updated the branch so it could merge. Binbin is a committer, a lead in three workgroups and a Kata Containers maintainer.
On the vLLM side, #55325 was my own patch, 29 lines across two files. Robert Shaw (@robertgshaw2-redhat, Red Hat), one of vLLM’s five Lead Maintainers, merged it, and Sylvester Kaczmarek (@sylvesterkaczmarek) reviewed and approved it after checking CI.
Show the data
| bug | fix | clock starts (UTC) | fixed (UTC) | elapsed | fix author |
|---|---|---|---|---|---|
| vLLM #55324 | vLLM #55325 | 2026-09-04 11:15:24 | 2026-09-10 22:21:22 | 6.5 days | colinmcnamara |
| #3500 | #3944 | 2026-09-04 23:56:14 | 2026-09-20 02:35:50 | 15 days | Xunzhuo |
| #4118 | #4125 | 2026-09-24 08:42:30 | 2026-09-25 08:59:32 | about a day | Xunzhuo |
| #4132 | #4125 | 2026-09-24 21:19:44 | 2026-09-25 08:59:32 | about 12 hours | Xunzhuo |
| #3892 | #4266 | 2026-09-26 13:01:00 (claimed) | 2026-09-28 11:18:38 | 1.9 days | colinmcnamara |
There is a lesson in that timeline, and it is a lesson rather than a grievance. This team moves on reports, which is the whole reason to file them here. It is also why you should claim an issue first if you want to write the code yourself. I claimed #3892 at 13:01Z and opened #4266 at 23:55Z the same day. Claim to merge was 1.9 days.
How to do this yourself
Six things, in the order they mattered.
1. Prove the bug on real software before you claim it. #4266 was demonstrated against a real Envoy 1.35.3 and a real agentgateway v1.3.1, not a mock.
2. Write tests that fail on main and pass with the fix. The routing tests in #4266 did. Then plant deliberately wrong fixes and confirm the tests catch them.
Show the data
| part | kind | lines added |
|---|---|---|
| Router code | fix | 187 |
| unit tests | test | 407 |
| E2E tests | test | 707 |
| test backend | test | 43 |
| docs | other | 11 |
| tooling | other | 4 |
| total | 1,359 |
3. Have several AI models review the work independently, then check their findings yourself. They disagree with each other, which is useful. They are also confidently wrong on occasion, which is why the checking is not optional.
4. Own the reviewer’s catch in public. Prove exactly where the reviewer was right, fix it, and post the evidence in the thread. Do not quietly patch it.
5. Build exactly what the maintainers ask for. Stefan specified the E2E test. The E2E test matched the specification. That is most of what “easy to review” means.
6. Credit everyone by name.
Full circle
I needed this capability for my own work. Stated plainly, the capability is this: optimize token spend in real time, and move models in and out transparently without taking services down. Because the fixes went upstream instead of into a private fork, it is not mine alone.
Leaders watching that chart have a lever on spend worth measuring. AHEAD’s engineers, sellers and project managers can use the same tools I use. Everyone else who runs this stack gets it for free, including the people who will hit that full duplex bug next month and never know it was there.
That is the argument for working upstream rather than around it.
A private fix helps one company, and that company maintains it forever. An upstream fix helps everyone, and somebody else maintains it with you.
What is next
Measure it. I have not measured the savings. This is where the work laptop comes back in. The before-and-after token number, taken there, is the actual answer to the question that started this, and I have not run it yet. I would rather publish the mechanism now and the measurement when it exists than imply a number I do not have.
Bugs in flight. My patch for null stop fields in non-streaming /v1/messages, vLLM #56376, is open and waiting on review. And the Router’s finish-reason allowlist rejects values that vLLM itself can return. That one is next, once I have captured it on my own hardware.
Research directions, and I mean research rather than product. The control point is the natural place to screen for prompt and logic injection, because it is the one place every request passes through. AI honeypots are the other thread I keep pulling on. Neither is something I am shipping. Both are things I think the control point makes newly possible.
The chart still goes up. The difference is that there is now somewhere to stand.