As query traffic grows, so does the compute bill for embedding it. On a low-power device, a large model may not fit in memory. A smaller query model could reduce that cost, but switching usually means re-embedding the collection.
Constella lets you make that switch. It’s a family of models on Hugging Face built around Stella, a 400M-parameter English embedding model. Stella encodes your documents. Zero, Nano, or Stella itself can then encode your queries, all searching the same Qdrant collection.
We’re sharing Constella as a research preview, with results across 15 BEIR datasets.
One Index, Three Query Models
A document vector can serve thousands of searches. A query vector usually serves one. Constella spends more compute on the document side, where you can reuse the result, and gives you a choice on the query side.
We trained Zero and Nano to reproduce Stella’s query embeddings. That’s what makes the models hot-swappable: you can change the query encoder per request while the stored document vectors stay fixed.
Inside Zero and Nano
Zero is a learned bag of tokens: it looks up each token’s vector, pools the vectors, and normalizes the result. That’s the entire query encoder. It makes search cheap, but loses word order: the same tokens with the same counts produce the same vector, even when rearranging them changes the meaning.
Nano adds context: its transformer models how tokens relate to each other and their positions. It combines features from layers 4, 8, and 12, then projects them into Stella’s 1024-dimensional space. We trained it on 199,999,721 examples to match frozen Stella embeddings, with roughly one-twelfth of Stella’s parameters. Our distillation and evaluation harness is on GitHub, including the training code and experiment records.
The Numbers: 15 BEIR Datasets
We evaluated Zero, Nano, and full Stella across BEIR-15, a collection of search tasks covering scientific papers, questions, claims, and other text. Each query model searches the same Stella document vectors.
The table reports exact-search nDCG@10, which measures how well relevant documents rank in the first 10 results. Higher is better. Each dataset has equal weight in the averages; CQADupStack combines its 12 forums into one dataset score. Shading shows each Zero and Nano score as a share of Full Stella’s.
† Training-contamination caveat: Stella reports training or evaluation exposure to these four datasets. Zero and Nano learn from Stella, so treat these scores as a comparison within the family, not a test on entirely unseen data. Full evaluation details.
Nano retains about 91% of full Stella’s average score across all 15 datasets, with a much smaller query transformer. Zero scores lower overall, but outscores Nano on FEVER, HotpotQA, and Climate-FEVER. The tradeoff varies by workload, which is why the choice of query model is worth testing on your own data.
We recommend pairing Zero with BM25 for hybrid search. Zero gained more from the combination than Nano in our benchmarks, improving retrieval quality without adding a transformer to the query path. Our Zero model card includes recommended fusion settings and practical guidance for getting started.
How Fast Is the Query Side?
On an Apple M5 Pro CPU, full Stella encoded a warm 20-word query in 38.95 ms. Nano took 3.13 ms, and Zero took 0.081 ms. That’s about 12 times faster for Nano and 480 times faster for Zero than full Stella.
The table separates model loading, the first query after loading, and warm p50, the median query time.
We measured all three with FastEmbed and ONNX Runtime on CPU, using four threads and batch size one. Values are medians across three fresh processes, each with five warmups and 20 synthetic 20-word queries. Stella receives its required query instruction in addition to those 20 words. These are encoding times; Qdrant search and network time are additional.
Loading includes imports and local model initialization, with assets already downloaded and the operating system’s disk cache left intact. Keep models loaded to avoid paying startup costs when switching. The raw timings and full protocol are available to inspect.
What You Can Build With It
- Offline search on low-power devices. Encode manuals or a knowledge base with Stella on a server, then ship Zero or Nano with the vectors. Search locally, even without a connection.
- High-volume retrieval APIs. Use Zero to reduce query-encoding compute, with Nano or Stella available for workloads where their relevance gain justifies the cost. All three search the same collection.
- Search as you type. Use Zero for queries on each keystroke, then Nano when the user pauses or submits. Our local Pokémon demo uses this pattern.
- Agents that search repeatedly. An agent may retrieve repeatedly before answering. Try Zero or Nano for those intermediate searches to reduce the time spent encoding queries.
- Hybrid search without a transformer. Pair Zero with BM25 to find related concepts alongside exact terms across your content. Combine both result sets in Qdrant without running a transformer for each query.
It’s a standard Qdrant + FastEmbed setup: embed your documents, store the vectors, and query the collection. The models are on Hugging Face, and native FastEmbed support is available on the research-preview branch.
Install the preview:
Create a collection and encode your documents once with Stella:
Now choose your query model. To switch from Zero to Nano, change one model name:
Same collection. Same query code. The stored document vectors stay exactly where they are. See the model card for the supported query paths.
What’s Next
Constella is in internal review ahead of a full release. This research preview is a chance to try the models and help shape what comes next. We’d love to hear where they work, where they fall short, and what you build with them. Share your feedback and what you build in our Discord community.










