Jevjitsu: Or, How We Tried Generalized Classifiers on Everything

Source: Qdrant•

Jevjitsu: Or, How We Tried Generalized Classifiers on Everything

Classification is one of the oldest problems in machine learning. Unlike modern decoding transformers, most classification models were non-autoregressive by nature. Before Jev, the common approaches were training your own classifier, which needs labels, using a zero-shot classifier (compared in…

Classification is one of the oldest problems in machine learning. Unlike modern decoding transformers, most classification models were non-autoregressive by nature. Before Jev, the common approaches were training your own classifier, which needs labels, using a zero-shot classifier (compared in this Hugging Face benchmark), or prompting an LLM with some hacky type safety on top. There were also meta-ML methods, but none of them reached the mainstream.

With the introduction of Jev, a decision model from TypeSafe served through OpenRouter, the taxonomy of classifiers has changed yet again:

There are two fundamental ideas that the Jev advancement brings to light:

  • Jev is the next step in making models usable without task-specific training. Pretrained transformers meant you no longer trained a model from scratch for each task, and LLMs let you describe what you want. Jev keeps “describe it, don’t train it” while addressing what makes LLMs awkward for classification: slow responses, high costs, and answers outside your options.
  • A conceptual shift in consumers of the “LLM Is All You Need” approach. Recently, the default was to stretch an LLM over everything. Jev is part of a move toward concrete methods. Even if the underlying concept still uses the transformer architecture, the constraints and promise are way different for the end customer.

Since Jev can take on any classification task, we tried it on several. Everyone has tested it as a reranker by now, and so did we, but the more interesting part came after: how to rerank better, how reading the query before searching helps, and what happens if you make Jev chunk a document. Here is what we expected, what surprised us, and what we’d try next. All of the experiments are in the many-jev-recipies repository if you want to run them yourself.

Reranking

Hybrid search usually finds the useful results. The harder part is the order: the difference between the best answer at position seven and a weaker one at position one can be too subtle for the first step, so it needs a “smarter cousin”, a reranker, to put the candidates in order. LLM judges are smart enough for that job but too slow and costly to run on every query. Jev fits naturally here: it returns numeric scores without generating text, so they sort directly, with nothing to parse.

We tried three ways to ask:

At this point, testing Jev as a reranker on a few hundred queries is practically a rite of passage. Every search vendor runs it on 80, 100, 200, or 300 queries, posts a table, and moves on. So naturally, we did too: 100 NFCorpus queries, the top 30 and top 50 candidates from Qdrant with BGE-small, and each method, plus two local cross-encoders for comparison, reordering the same candidates. The table shows each method’s nDCG@10 lift over BGE-small alone:

We expected the most expensive method, Iterative, to win by a wide margin. It had the highest score, but only by a small margin: all three Jev methods beat BGE-small at both depths, and the gaps between them are under 0.02. More effort doesn’t always buy a better result, so the cheap options win: Score is our default, and Iterative isn’t worth ten requests per query. The full comparison, including the cross-encoders, is in the reranking notebook.

Both cross-encoders gained less than any Jev method, and bge-reranker-base did not improve on BGE-small at all. Jev lets us adapt the relevance question without training a new model. Cross-encoders offer local execution and can be faster or cheaper. Which fits depends on how you weigh flexibility, quality, latency, and cost. But reranking is only the start: Jev’s judgments can guide other search decisions too.

Reducing Repetitive Answers

Imagine an e-commerce search for “iphone” that returns ten variants of the same iPhone 18. Sometimes that’s what the shopper wants, but often they’d rather see a range of products on the first page.

We measured this on 240 held-out queries from the WANDS benchmark, scoring relevance (nDCG@10) and repetition: near-duplicate pairs, results whose embeddings are 0.95 similar or more, on the first page of 10. Product classes count the different product categories in the top 10 results. This helps show whether shoppers see a useful range of products or mostly variations of the same thing.

Jev reranking improved relevance but left near-duplicates in place and reduced product variety. Even ordering by the human relevance labels narrowed the page to 2.66 product classes. Relevant results can still be repetitive.

Qdrant’s built-in Maximal Marginal Relevance (MMR) goes the other way: at diversity 0.5, it removes the near-duplicates and gives the widest page, but costs 0.12 nDCG@10, because it measures relevance as vector similarity to the query.

The fix was to use the reranker’s score as the relevance signal for diversity: either swap MMR’s relevance term for Jev’s answer, or run Qdrant’s MMR first and let Jev rerank its top 20. Both beat hybrid search on relevance and repetition, giving up part of the plain rerank’s gain for a more varied page. On this dataset, plain reranking gave the strongest relevance; combining Jev with MMR reduced repetition while keeping relevance above the hybrid baseline. The full comparison is in the pruning notebook.

Query Understanding

A shopper typing “kitchen faucet replacement” wants plumbing parts, and recognizing that category can improve the results. We followed the main rule from Doug Turnbull’s experiments: filter only when very sure, because a wrong filter hides useful products; otherwise, boost matching products.

We expected this part to work out of the box, but choosing category names from the products gave almost no gain. What helped was clustering the products, letting a cheap LLM name each group, and using Jev to check the names. The exact checks are in the notebook.

From here on, nothing is generated, only categorized:

  • At indexing time, Jev tags every product with these categories, stored as payload: a Nintendo Switch case gets device cases, under Electronics & Accessories.
  • At search time, Jev scores each product category against the query. At 0.9 or more, search filters to that category. From 0.5 to 0.9, it boosts matching products. Below that, it does nothing.

The filtering and boosting run in Qdrant: the category is a payload filter, and the boost is a score formula on top of hybrid search.

We compared setups on 300 Amazon-C4 searches, and that mix of filtering and boosting worked best:

Jev taxonomy routing on Amazon-C4300 searches, higher is betternDCG@10 (x100)TieredDefaultBoost onlyFilter only010203040Hybrid searchWith Jev taxonomy

Jev adds half a second to two seconds per query, depending on how many categories it reads. Small classifiers trained on Jev’s product tags handled 37% of queries with boosts only, skipping Jev with slightly higher average nDCG@10. The details are in section 6 of the .

The Amazon-C4 queries are long and LLM-generated, which gives Jev plenty of context. Real search boxes see much shorter queries, so we tried a few of our own. They have no relevance labels, so read them as examples, not measurements:

Filtering works when the category is the whole intent. Short queries are more ambiguous, and a filter hurts them most: once it removes products, MMR cannot bring them back. There are two ideas we have not tested yet. One is boosting instead of filtering on one- or two-word queries. The other is giving Jev more context, for example “apple” sent together with the shop it was typed in, say "shop": "electronics". Either way, a confident category prediction can still miss what the shopper wants. The lesson is to check your own queries and edge cases before you adopt even the most confident one.

Semantic Chunking

Retrieval often brings back the right chunk with only half the answer in it, because the chunker cut the paragraph in two. Fixed-size chunkers cut wherever the token count runs out. Most semantic chunkers cut where the embeddings of neighboring sentences drift apart, so a paragraph that changes vocabulary mid-argument can get split, and two articles that share vocabulary can get merged.

What if Jev judged each gap between sentences instead? We send it a window of numbered sentences from the document, trimmed here to four from a QASPER paper:

For each gap, we ask whether the next sentence starts a new topic, where a new topic means “a new section, story, question or theme begins”. For sentence 2, the question is “Does sentence 2 start a new topic?” Jev returned a yes probability of 0.97 for that section heading, followed by 0.13 and 0.03 for the next two sentences.

We cut where the answer passes a threshold, keeping each chunk between a minimum and maximum size. The cuts are a plain function of Jev’s answers, so changing the chunk size needs no new requests.

We tested this on QASPER’s 407 research papers and 1,297 questions with evidence paragraphs marked by researchers. For each question, Qdrant retrieved five chunks from its paper. We measured how much evidence they contained and how much unrelated text they carried, using character-based recall, precision, and intersection over union (IoU).

At matched mean chunk sizes, Jev improved all three metrics against fixed-size, recursive, and embedding-based chunkers. Average relative gains were 13.0% in IoU, 11.2% in precision, and 9.1% in recall. Every 95% confidence interval excluded zero. We expected a modest improvement and got it on the first try, without tuning the question. Here is one example:

The trade-off is cost and simplicity: the other chunkers run locally for free, while Jev sends your text to an API, at about $0.27 per million document tokens. The smallest gain was against a recursive splitter that breaks on blank lines first, since QASPER’s answers are whole paragraphs. Rewording the question may improve the chunks further. The implementation and the full comparison are in the notebook.

What We Would Try First

For RAG, start with the chunker: it was easy to set up and adds no overhead at query time. For product search, try combining Jev’s relevance scores with MMR to reduce repetition while keeping relevance above the hybrid baseline.

Query understanding has the most unexplored potential, but also needs more care. Test ambiguous queries like “apple” before trusting a category prediction to filter results. The fixes we proposed are still untested.

We would skip Iterative reranking for now: ten requests per query bought only a small gain in our experiment.

Summary

Jev’s most interesting uses came after reranking: balancing relevance with variety, choosing how to search, and deciding where to split a document. Those tasks benefit from different trade-offs, so there is no single recipe.

The industry is exploring the same idea, with OpenAI’s Decisions API and open-source approaches such as OpenJev, JevLite, and CLM.

Find a decision in your pipeline worth improving, then test whether Jev makes it better at a cost you can accept. Our give you a place to start.

What this article says