
/Community Topics4 min read
Testing Jev: Speed, Cost, and Classification Compared to LLMs
Jev has taken the world by storm and claims to be up to 200x faster than a normal LLM. In this blog I discuss my own findings from a test that I ran using Tavily, Jev, and GPT-6 Luna from OpenAI.
TypeSafe says Jev is "up to 200x faster." So I decided to test it out myself.
Jev launched two weeks ago and immediately took over everyone's timeline. It is a "System One" decision model, which means it doesn't generate text like a traditional LLM. Instead, it is designed to deliver fast, cheap, structured answers for making decisions.
The phrase that has caught everyone's attention is "up to 200x faster, 400x cheaper" than a normal LLM.
But that's TypeSafe's number, based on their own benchmarks. So, I built a small classifier application and integrated Tavily search results to test it out myself.
Admittedly, I didn't see the same astounding numbers that TypeSafe advertises. But still, the numbers that I did see were extremely impressive and make for an interesting conversation.
My test: Signal Sort

I built a tool called Signal Sort. This small app uses a single Tavily search and runs two identical classification pipelines. Jev is used on one side, while OpenAI’s GPT-6 Luna runs on the other.
I used a simple query for my test: Jev AI model benchmarks. With the amount of online attention that Jev has been getting, a query like this can result in a wide range of information. Because of this, it would be nice to be able to classify the information gathered.
Both pipelines answered the same two questions for every result:
- Is this result relevant to the query?
- If relevant, is its stance toward TypeSafe's claims skeptical, neutral, or promotional?
For the first pipeline, Jev was our decision maker.
For the second, I used GPT-6 Luna. This is OpenAI's smallest and fastest GPT-6 tier, built specifically for high-volume, latency-sensitive work like this.
In an effort to get a fair average for how these two approaches perform, I ran every result from the query through both pipelines three times and averaged the results.
Where the two pipelines agreed, and where they split
For the core task, classifying whether the result is relevant or not, Jev and GPT-6 Luna agreed that all 10 results returned were indeed relevant for each run. I wouldn't be a good Tavily employee if I didn't point out the fact that Tavily made this step largely unnecessary since it tries to only return relevant results by default.
Yet still, it's an important data point that neither Jev nor GPT-6 Luna misclassified relevant information as irrelevant.
However, when it comes to stance, the two pipelines disagreed 30% of the time, and on all three disagreeing results, GPT-6 Luna labeled the stance as "skeptical".
Both models were consistent across their three repeat rounds on every disagreement, and because of this, I needed to step in to determine who is actually right. I read the same short snippet each pipeline saw for all 10 results, making my own independent call with the same rubric: does it question or fact-check TypeSafe's claims, report them neutrally, or amplify and promote them. Scored against my own judgment, Jev and GPT-6 Luna each got eight out of ten right. Below is a table detailing classifications that they got wrong.
Result | Jev's stance | GPT-6 Luna's stance | Independent read |
|---|---|---|---|
GitHub — jevbench | neutral (53% conf.) | skeptical | neutral — Jev matched |
regolo.ai — benchmarks & alternatives | promotional (70% conf.) | skeptical | skeptical — LLM matched |
Reddit — "I benchmarked TypeSafe's Jev..." | promotional (60% conf.) | skeptical | neutral — neither matched |
My speed and cost numbers
Metric | Jev | GPT-6 Luna | LLM ÷ Jev |
|---|---|---|---|
Average latency per round | 831.9 ms | 7,542.1 ms | 9.1x |
Average cost per round | $0.0003478 | $0.0010201 | 2.9x |
With the results of the test, it's easy to see why Jev has garnered so much attention.
For my test, one query and ten results, each run three times to smooth out noise, Jev operated 9.1x faster and 2.9x cheaper. Although these didn't live up to the numbers TypeSafe advertised, imagine how much time and cost that would save at a cumulative scale.
Even at "only" 9.1x and 2.9x, that compounds aggressively once you're running the filtering step millions of times a month.
Final thoughts
The attention Jev has been getting is legitimate, and the speed and cost savings could be real for your workload, depending on what you're building.
However, the LLM that you pick for any comparison you run will have a huge impact on the improvement numbers that you see and the performance. GPT-6 Luna agreed with Jev on relevance 100% of the time, but they disagreed with each other on stance 30% of the time. When I checked both against an independent read, they came out tied, equally accurate, just wrong about different results. But the big takeaway is that Jev was able to achieve virtually the same result 9.1x faster and 2.9x cheaper.
Because of this, Jev is absolutely a great tool to reach for if inexpensive and fast classification is needed for your workflow, and can be a great alternative to traditional LLMs when it comes to making decisions based on your Tavily search results.
If you're building a search-driven agent and want to try this decision-making step yourself, sign up for a free Tavily account. You'll get 1,000 credits each month, on the house, to test with.
