Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Sentence Transformers v6.0 adds a fourth model type, MultiVectorEncoder, for ColBERT-style late interaction retrieval, alongside a complete training approach for it, according to a Hugging Face blog post. The author demonstrates finetuning a multi-vector model on domain data and states the method can also train new multi-vector models from scratch. Installation is via pip install -U "sentence-transformers[train]".
Multi-vector models keep one small vector per token and score a query against a document with the MaxSim operator, in which every query token finds its best-matching document token and the scores are summed. The post attributes stronger retrieval but a larger index to this token-level matching relative to single-vector dense embeddings.
Training components listed are the model, datasets, loss functions, training arguments, evaluators, and the trainer class. The post describes finetuning an existing multi-vector checkpoint such as lightonai/mLateOn-unsupervised, or building from a base transformer, where a randomly initialized token-level projection is appended and training is required before the model is useful.
On data, the post uses tomaarsen/miriad-4.4M-split with 4,467,542 medical question-passage rows averaging 941 tokens. Training used one epoch, a per-device train batch size of 128, learning rate 1e-4, bf16, and CachedMultiVectorMultipleNegativesRankingLoss with mini_batch_size 16.
The post reports that starting-point choice mattered in experiments: six starting points trained with an identical recipe on 25k medical question-passage pairs from MIRIAD and evaluated on 1,000 held-out questions against a 50,000-passage corpus, with unsupervised checkpoints adapting better than finished siblings. It also reports that document truncation at 180 to 512 tokens cost up to 0.24 NDCG@10 on the medical evaluation, and that a punctuation skiplist shrank the document index by 9.6% on that data.
The author states the finetuned multi-vector-encoder/mLateOn-medical model was trained in 14.5 hours on a single RTX 3090 and outperformed every general-purpose retrieval model found for the medical retrieval evaluation: dense, sparse, lexical, and multi-vector alike. The post links companion write-ups on dense embedding models, sparse embedding models, rerankers, and on using multi-vector embedding models.
Based on reporting from the original publisher. Visit the source for full context and later updates.