The training codes of Jasper-Token-Compression-600M
☆21Nov 19, 2025Updated 9 months ago
Alternatives and similar repositories for Jasper-Token-Compression-Training
Users that are interested in Jasper-Token-Compression-Training are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Jul 7, 2024Updated 2 years ago
- The Python Implementation of CRISP: Clustering Multi-Vector Representations for Denoising and Pruning☆27Jul 27, 2025Updated last year
- ☆16Aug 22, 2026Updated last week
- Official Code Repositiry for "RaDeR: Reasoning-aware Dense Retrieval Models" accepted at Main Conference EMNLP 2025☆18Jun 23, 2025Updated last year
- Jina VDR is a multilingual, multi-domain benchmark for visual document retrieval☆38Aug 4, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- AutoRAG example about benchmarking Korean embeddings.☆46Oct 2, 2024Updated last year
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆21Updated this week
- ☆31Jun 11, 2026Updated 2 months ago
- A massively multilingual modern encoder language model☆153Jan 20, 2026Updated 7 months ago
- Multi-Word Probabilistic based supertokenizer☆15May 15, 2025Updated last year
- 🎹 Instruct.KR 2025 Summer Meetup: 오픈소스 LLM, vLLM으로 Production까지 🎹☆23Aug 2, 2025Updated last year
- Training code for Sparse Autoencoders on Embedding models☆40Jul 11, 2026Updated last month
- MEXMA: Token-level objectives improve sentence representations☆44Jan 6, 2025Updated last year
- [ICLR 2026 Oral] Official Implementation of the paper "MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interactio…☆22Jul 2, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Draft grounded rebuttals to your paper's reviews, with the experiments actually run in your workspace☆17Jul 30, 2026Updated last month
- Korean-MTEB☆105May 12, 2026Updated 3 months ago
- ☆56Jun 21, 2025Updated last year
- [NeurIPS 2025] MergeBench: A Benchmark for Merging Domain-Specialized LLMs☆50Updated this week
- Bridge incompatible embedding spaces with a single SVD. When your embedding provider deprecates a model, adapt instead of re-embedding.☆36Apr 28, 2026Updated 4 months ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- Can VLMs understand students' hand-drawn math work?☆19Jan 20, 2026Updated 7 months ago
- Performs benchmarking on two Korean datasets with minimal time and effort.☆48Aug 6, 2026Updated 3 weeks ago
- State-of-the-art paired encoder and decoder models (17M-1B params)☆78Aug 6, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- It shows how to deploy and use an agent with LLM.☆19Mar 1, 2025Updated last year
- Label shift estimation for transfer difficulty with Familiarity.☆10Feb 4, 2025Updated last year
- LLM-as-a-judge using G-eval Scratch☆15Oct 12, 2025Updated 10 months ago
- Code for the paper "Multi-Field Adaptive Retrieval," a research project on a semi-structured document retrieval☆18Feb 13, 2026Updated 6 months ago
- Nearly Inference Free Embeddings: make your RAG queries 500x faster☆85Apr 27, 2026Updated 4 months ago
- Embedding Inversion via Conditional Masked Diffusion: recover original text from embedding vectors using parallel denoising. Live demo + …☆60Mar 7, 2026Updated 5 months ago
- Compression for unit-norm embedding vectors using spherical coordinates☆82Jan 23, 2026Updated 7 months ago
- A Multilingual Keyboard Layout-Based Typo Generator☆17Nov 23, 2025Updated 9 months ago
- Identify which embedding model produced a vector using digit-level tokenization and a tiny transformer☆23Mar 7, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL'26 Workshop] KoViDoRe: Korean Visual Document Retrieval Benchmark☆26Jul 2, 2026Updated last month
- ☆23May 30, 2025Updated last year
- A modular framework for training and inference of (compressed) multi-vector retrieval across any modality.☆22Apr 4, 2026Updated 4 months ago
- Korean Sentence Embedding Model Performance Benchmark for RAG☆49Jan 27, 2025Updated last year
- KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models☆25Aug 24, 2024Updated 2 years ago
- ☆18Jan 6, 2025Updated last year
- [ICML 2026] Scaling Beyond Masked Diffusion Language Models☆36Jul 3, 2026Updated last month