Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
4.6
TFLOPS
Antoine Louis
antoinelouis
74
8
84
Follow
Abdelkareem's profile picture
PeepDaSlan9's profile picture
SergeTouvoly's profile picture
63 followers
·
16 following
https://antoinelouis.co
antoinelouis_
ant-louis
antoine-louis
AI & ML interests
ML • NLP • IR • QA
Recent Activity
upvoted
an
article
about 4 hours ago
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
reacted
to
tomaarsen
's
post
with 🚀
about 4 hours ago
🚨 I've just published Sentence Transformers v6.0, introducing MultiVectorEncoder: ColBERT-style late interaction models are now a fourth model type, for training, inference, and interpretation, alongside the dense, sparse, and reranker models! Details: Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away. It is also the state of the art for visual document retrieval, where a text query is matched against page images directly, charts and tables included, with no OCR step in between. Any PyLate, Stanford ColBERT, or ColPali checkpoint loads straight into the same familiar API: model.encode_query(), model.encode_document(), and model.similarity() just work, whether the documents are texts or page images. Does it help? LightOn trained LateOn (multi-vector) and DenseOn (dense) on the same data with the same 149M ModernBERT backbone, and the multi-vector model wins on 9 of the 13 NanoBEIR datasets: 0.6868 vs 0.6764 mean NDCG@10. The price is a bigger index, and the new HierarchicalTokenPooling module halves it at roughly no retrieval cost. Antoine Chaffin, Raphaël Sourty, and I wrote a blog post walking through multi-vector models in practice: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable. Check it out if you want to get started, or just point your Agent to the URL: https://huggingface.co/blog/multi-vector-encoder pip install sentence-transformers==6.0.0 Release notes: https://github.com/huggingface/sentence-transformers/releases/tag/v6.0.0
new
activity
about 4 hours ago
antoinelouis/colbert-xm:
Add Sentence Transformers usage
View all activity
Organizations
antoinelouis
's Spaces
2
Sort: Recently updated
pinned
Running
14
MTEM Pruner
✂
Multilingual Text Embedding Model Pruner
pinned
Sleeping
Agents
13
DécouvrIR
🥇
Leaderboard of information retrieval models in French