illuin-tech / modernvbertView on GitHub
ModernVBERT is a 250M-parameter vision–language encoder that aligns a text-encoder (Ettin-150M) with a vision-encoder (SigLIP2-B) through a MLM objective. When fine-tuned for document retrieval, ModernVBERT sets a new state of the art for sub-1B models on ViDoRe tasks.
16Oct 16, 2025Updated 10 months ago

Alternatives and similar repositories for modernvbert

Users that are interested in modernvbert are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?