☆54Feb 10, 2025Updated last year
Alternatives and similar repositories for ModernBERT-Instruct-mini-cookbook
Users that are interested in ModernBERT-Instruct-mini-cookbook are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ModernVBERT is a 250M-parameter vision–language encoder that aligns a text-encoder (Ettin-150M) with a vision-encoder (SigLIP2-B) through…☆16Oct 16, 2025Updated 10 months ago
- ☆110Jun 2, 2025Updated last year
- ☆14Oct 21, 2024Updated last year
- Luth is a state-of-the-art series of fine-tuned LLMs for French☆47Oct 12, 2025Updated 10 months ago
- Generalist and Lightweight Model for Text Classification☆243Jul 21, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 中文预训练ModernBert☆101Apr 11, 2025Updated last year
- recipe for training fully-featured self supervised image jepa models☆14Jun 4, 2025Updated last year
- ☆37May 5, 2025Updated last year
- code for training & evaluating Contextual Document Embedding models☆207May 14, 2025Updated last year
- The official code of "PixelWorld: Towards Perceiving Everything as Pixels" [TMLR25]☆15Sep 12, 2025Updated 11 months ago
- ChunkNorris is a black belt in document chunking to feed your LLMs and RAG apps 🥋🔪☆29Aug 19, 2026Updated 2 weeks ago
- Code for SaGe subword tokenizer (EACL 2023)☆28Nov 30, 2024Updated last year
- This repository contains the training and evaluation code for llm-jp-modernbert-base.☆17Jun 17, 2025Updated last year
- Official implementation of "GPT or BERT: why not both?"☆65Jul 28, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Materials for the "My Workflow for Understanding LLM Architectures" tutorial☆30Apr 10, 2026Updated 4 months ago
- dspy-cli is a tool for creating, developing, testing, and deploying DSPy programs as HTTP APIs.☆135Mar 3, 2026Updated 5 months ago
- Label shift estimation for transfer difficulty with Familiarity.☆10Feb 4, 2025Updated last year
- Late Interaction Models Training & Retrieval☆888Jul 23, 2026Updated last month
- Python library to use Pleias-RAG models☆72Jul 1, 2026Updated 2 months ago
- ☆16Updated this week
- ☆18Jan 13, 2025Updated last year
- ☆104Jul 4, 2025Updated last year
- ☆59Aug 19, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆10Oct 22, 2024Updated last year
- Bringing BERT into modernity via both architecture changes and scaling☆1,716Mar 1, 2026Updated 6 months ago
- This repository contains papers for a comprehensive survey on accelerated generation techniques in Large Language Models (LLMs).☆11May 24, 2024Updated 2 years ago
- Code for the Avey-B paper (https://arxiv.org/abs/2602.15814)☆32Feb 21, 2026Updated 6 months ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 6 months ago
- Efficient and scalable zero-shot entity linking☆150Jul 20, 2026Updated last month
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- This is a question-output workflow template for shiny app!☆12May 17, 2019Updated 7 years ago
- ☆67Mar 4, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- jShiny server☆16Jan 24, 2017Updated 9 years ago
- [ICML2026] Official JAX code for Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying☆17Jul 3, 2026Updated last month
- [ICML 2024] Learning with Complementary Labels Revisited: The Selected-Completely-at-Random Setting Is More Practical☆13May 12, 2024Updated 2 years ago
- Causal Impact but with MFLES and conformal prediction intervals☆33Dec 31, 2024Updated last year
- Efficient encoder-decoder architecture for small language models (≤1B parameters) with cross-architecture knowledge distillation and visi…☆32Feb 7, 2025Updated last year
- LexiSignVQA: A Unified Training-free Multi-stage Approach to Multimodal Legal Question Answering on Traffic Sign Rules☆23Nov 18, 2025Updated 9 months ago
- ☆41Aug 21, 2026Updated last week