Out of all the smart ways to optimize an AI system, the smartest is model routing.
Model routing is intelligently picking the optimal model based on the task complexity. A tiny model can handle most questions. But larger models will jump in if the task demands reasoning. A router LLM decides which model should handle it.
It doesn’t compromise quality or the app’s capability. Everything the app should do, it will do. The end user barely notices the difference.
Engineers used a powerful model for the routing task. They make accurate decisions. But it comes with some serious drawbacks.
First, frontier models are too slow even for simple decision-making. Roughly 3- 329 seconds for classification tasks. Then, they cost a lot. Sure, the routing architecture saves cost in output tokens by leveraging smaller models. But the routing part itself was expensive. Lastly, even larger frontier models sometimes pick the correct choice, but in the wrong way. It makes them unreliable. For instance, a spam classifier is expected to return either ‘spam’ or ‘safe’. But once in a while, the model tries to be a bit extra smart and returns ‘spamy’ with an extra ‘y’. Your app breaks because you never thought this would happen.
This makes model routing unreliable and the cost savings achieved minuscule. For this reason, routing was thought of as an enhancement rather than a design choice. It rarely appears on prototypes.
There has been a lot of buzz around Jev lately. Because it solves a fundamental problem the other frontier models overlooked. Type-safe decision-making. It doesn’t converse with you the way GPT or Claude does. It doesn’t think like these highly intelligent models. It does one thing and does it well—faster, better, cheaper.
Jev could make choices in 70-500 milliseconds (compared to 3-329 seconds of frontier models). And it can do it for as low as $0.042 per million input tokens. No extra cost for output or reasoning tokens. Besides, it doesn’t occasionally try to outsmart the prompt. You define what you get, and Jev sticks to it. These qualities make Jev perfect for model routing.
How Jev can route the incoming request to different models.
In the rest of this post, I’ll take you through how we can use Jev for model routing using a worked example.
Routing with Jev
Jev is a proprietary model. You can access it through their client SDK. You have to set up the account, add credits (minimum $5), and create an API key. You can do it on TypeSafe AI’s portal.
Once you have created the API key, you can set an environment variable. There are many ways to do it. It all depends on where and how you run your code. I prefer to run my Python code with UV and set my environment variables in a .env file.
Create a .env file at the root of your project folder with the following content.
In addition to the Typesafe API key, I’ve also set the anthropic api key. That’s because I’m going to let Claude do the real task. Jev will route the query to the correct Claude model: haiku, sonnet, or opus.
The following code illustrates simple model routing.
The code above uses the Choice question type. This question type helps us choose an option from a given list. Other types include noul, which returns true or false, and score, which returns a numeric score for every option.
Inside the Choice object, we’ve laid out our instructions and the criteria to help Jev pick the class. We’ve also set a minimum confidence level. If Jev couldn’t pick a class with enough confidence, we can use this score to handle it separately. In my code, I’m routing it to the most powerful model. But it’s entirely up to the application.
For a trivial but potentially reasoning-requiring question, this is how the output looks.
Jev picked Sonnet to handle this question. But it has given a very low confidence score of 0.38. Because of this, my application code routes it to Opus instead of Sonnet.
Here’s the response for a simpler question:
This is a trivial question that doesn’t require reasoning. But not so trivial to answer without enough knowledge. So Jev picked sonnet and assigned a high confidence score of 0.89. My app accordingly used Claude Sonnet to respond.
Final Thoughts
Despite model routing bringing immense benefit to AI systems, its adoption is weak. The cost-benefit doesn’t seem to offset the unreliability and increased latency. In most systems, it was thought of as an optional enhancement.
But all this time, engineering teams were using frontier models for routing, which is overkill. Jev turned the table. Now, model routing is fast and cheap. This helps us build apps without compromising quality or adding extra waiting time.
This post shows a worked example of how to implement model routing using Jev. Hope you find it helpful.

