Vecski: change embedding models without an all-at-once migration

Your search index is tied to the model that created its vectors. Switching embedding models can mean reading the original documents again, paying for another embedding pass, and rebuilding the index before you can use the new model.

Vecski adds another option: learn a mapping between the two vector spaces from paired examples, then translate vectors through that mapping. It is a small, self-hosted Rust API server for teams maintaining semantic search and retrieval-augmented generation systems.

The useful promise is a more gradual migration. Translation is an approximation, and the question is whether that approximation is good enough for your search workload.

What Vecski actually does

An embedding is a list of numbers that represents a piece of text for similarity search. Two models can encode the same text into different coordinate systems, even when their outputs have the same number of dimensions. You cannot assume that a query from one model can search an index created by another.

Vecski learns the relationship from matched pairs: each source vector and target vector must represent the same text, embedded by the two models. You produce those embeddings separately. Vecski receives numeric vectors, fits a translator, reports quality measurements, and applies the saved translator to further vectors.

The released v0.1.0 server supports centered orthogonal Procrustes and ridge regression. Both produce a compact affine map. Its automatic mode fits both candidates and favors Procrustes unless ridge improves held-out neighbor overlap by more than 0.005. It also supports different source and target dimensions.

You can fit, translate, evaluate, list, export, import, and delete translators through HTTP. JSON makes small experiments straightforward; raw little-endian float32 payloads reduce serialization overhead for larger batches. The server publishes an agent-readable guide at /llms.txt and an API specification at /openapi.json.

Vecski does not create embeddings, run a vector database, or move records between indexes. It supplies the translation step inside a migration you control.

Two migration paths, two different compromises

The direction of the map matters. Choose it before you prepare your pairs.

ApproachMap to fitWhat stays in placeWhat to test
Adapt new queries to an existing indexNew model → old modelExisting document vectors and indexWhether translated queries retrieve the documents your users need
Move existing document vectors toward a new spaceOld model → new modelThe original stored vectors can supply the migration inputWhether translated documents work well with native queries from the new model

For example, a team has a large support-document index and wants to switch the model used for incoming questions. A query-side translator can map those new query vectors into the old index's space while the team evaluates or rebuilds its document pipeline.

That keeps the old index useful, but it does not turn its document representations into those of the new model. Conversely, translating old document vectors into a new space cannot add distinctions the original encoder never captured. Wider output vectors do not create missing information.

Neither direction has a universal recall guarantee. Evaluate on your own corpus, queries, languages, and document types before changing live traffic.

A practical evaluation workflow

  1. Choose the migration direction. Label source and target models accordingly; a translator is directional.
  2. Create representative pairs. Embed the same sample texts with both models. Preserve row alignment, preprocessing, model versions, and dimensions. The right sample size depends on the task and dimensions; a few examples are an API demonstration, not evidence of production quality.
  3. Fit a candidate. Submit the paired vectors to /v1/translators. The default configuration holds out part of the sample and compares translation results with a mean-vector baseline.
  4. Inspect the report. Review neighbor overlap, matching accuracy, the baseline comparison, and warnings. A high cosine score alone is insufficient.
  5. Test on fresh data. Automatic estimator selection and ridge tuning use the internal holdout. Keep a separate, untouched evaluation set for the migration decision, and test retrieval against the actual index.
  6. Roll out gradually. Compare results against your existing system and, where feasible, a natively re-embedded sample. Keep a path back to the original index and model configuration.

Vecski exposes /v1/translators/{id}/evaluate for additional paired-vector checks. That is useful for inspecting a translator, but paired-vector metrics do not replace end-to-end search evaluation with relevant documents and real questions.

Read the report before trusting the map

Embedding vectors can cluster closely enough that predicting an average vector earns a superficially impressive cosine score. Vecski reports that mean-vector baseline alongside the translator's mean cosine so you can see whether the map learned something useful.

Its neighbor-overlap metric asks how much of a native query's top-k neighborhood is recovered by its translated counterpart within the evaluation set. Top-1 accuracy and mean reciprocal rank measure whether a translated vector finds its matching native vector. Pairwise cosine error measures differences in the target space's internal geometry.

These measurements answer different questions. Use them to investigate failure, then measure the search result quality your application actually needs. A translator that works for support tickets may degrade on legal documents, another language, or a new model version.

The repository includes performance measurements and a reproducible benchmark command. Those are project-reported measurements, not a throughput guarantee for your hardware, dimensions, batch size, or network path.

Install and start on localhost

Vecski v0.1.0 is Apache-2.0 licensed. The release provides binaries for macOS and Linux, each for ARM64 and x86-64. Its Homebrew formula builds the server from the tagged source and requires Rust at build time.

brew tap looskis/vecski https://github.com/looskis/vecski
brew install vecski
vecski --bind 127.0.0.1:8080 --data-dir ./vecski-data

In another terminal, check the service and discover its interface:

curl --fail http://127.0.0.1:8080/healthz
curl --fail http://127.0.0.1:8080/llms.txt
curl --fail http://127.0.0.1:8080/openapi.json

For a persistent Homebrew service, use brew services start vecski instead of the foreground command. The formula configures that service to bind to 127.0.0.1:8080 and store translators under Homebrew's var/vecski directory.

The repository's Python quick start walks through fitting, binary and JSON translation, downloading weights, and cleanup using synthetic pairs. It is an interface demonstration; use real paired embeddings for a migration assessment.

For your own data, put source and target arrays in a JSON file, with the same text represented at each matching row. Optional metadata includes name, source_model, and target_model; options.method can be auto, procrustes, or ridge. Submit the file with:

curl --fail http://127.0.0.1:8080/v1/translators \
  -H 'Content-Type: application/json' \
  --data-binary @paired-vectors.json

Use the returned translator ID or name for subsequent translation and evaluation requests. The README documents the full payloads and binary formats.

Where the data goes

The server fits and applies maps where you run it. Its fitting API accepts vectors rather than source documents, and translators persist as safetensors files in the configured data directory. You can export weights for use outside the server.

Producing paired embeddings is a separate step: if either embedding model is hosted, that step sends the sample text to that provider. Self-hosting Vecski does not change the data flow of your embedding pipeline. Treat vectors and fitted artifacts as data that still needs appropriate access controls.

A direct vecski invocation defaults to 0.0.0.0:8080, and API keys are optional. The explicit localhost binding above matters. For network access, configure API keys and a secured deployment. With keys configured, /v1/* requires a bearer token or X-API-Key; discovery endpoints remain public.

When it is worth trying

Vecski is useful when you need to keep an existing retrieval system working while changing models, or when you want to measure how far a mapping can carry you before committing to a full re-embedding job.

It is a poor substitute for re-embedding when the goal is to recover semantic information the old model missed. Mixing translated and natively embedded documents also needs testing: compatible dimensions do not guarantee comparable similarity scores.

Start with a representative sample and a clear quality threshold. If the translator passes, use it as part of a controlled migration. If it fails, you have an early signal that rebuilding the relevant embeddings is the better path.

Explore Vecski on GitHub, or browse the rest of the Looskis open-source collection.

Frequently asked questions

Does Vecski replace re-embedding?

It can provide an intermediate migration path when measured retrieval quality is acceptable. It does not reproduce every benefit of the new embedding model, and cannot recover information absent from the original vectors.

Does Vecski call an embedding provider?

The translation API works on vectors you supply. You generate paired embeddings separately, using local or hosted models. Any provider calls in that preparation step remain part of your own data pipeline.

Can the source and target dimensions differ?

Yes. Vecski supports different input and output dimensions. Matching the target dimension does not establish semantic equivalence; evaluate the translated vectors against native target embeddings and real retrieval tasks.

Is the internal holdout enough to approve a migration?

No. Vecski uses its internal holdout for automatic estimator selection and ridge tuning. Use an untouched evaluation set and end-to-end retrieval checks before deciding whether a translator meets your quality requirements.

Sources & further reading

Talk with us about your workflow →