The training codes of Jasper-Token-Compression-600M
☆22Nov 19, 2025Updated 10 months ago
Alternatives and similar repositories for Jasper-Token-Compression-Training
Users that are interested in Jasper-Token-Compression-Training are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Jul 7, 2024Updated 2 years ago
- Make running benchmark simple yet maintainable, again. Now only supports Korean-based cross-encoder.☆36Aug 30, 2026Updated 3 weeks ago
- The Python Implementation of CRISP: Clustering Multi-Vector Representations for Denoising and Pruning☆27Jul 27, 2025Updated last year
- ☆15Aug 22, 2026Updated 3 weeks ago
- Official Code Repositiry for "RaDeR: Reasoning-aware Dense Retrieval Models" accepted at Main Conference EMNLP 2025☆18Jun 23, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Jina VDR is a multilingual, multi-domain benchmark for visual document retrieval☆38Aug 4, 2025Updated last year
- AutoRAG example about benchmarking Korean embeddings.☆46Oct 2, 2024Updated last year
- ☆31Aug 30, 2026Updated 3 weeks ago
- A massively multilingual modern encoder language model☆155Jan 20, 2026Updated 8 months ago
- Multi-Word Probabilistic based supertokenizer☆15May 15, 2025Updated last year
- 🎹 Instruct.KR 2025 Summer Meetup: 오픈소스 LLM, vLLM으로 Production까지 🎹☆23Aug 2, 2025Updated last year
- Training code for Sparse Autoencoders on Embedding models☆40Jul 11, 2026Updated 2 months ago
- MEXMA: Token-level objectives improve sentence representations☆44Jan 6, 2025Updated last year
- [ICLR 2026 Oral] Official Implementation of the paper "MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interactio…☆23Jul 2, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Korean-MTEB☆106May 12, 2026Updated 4 months ago
- ☆56Jun 21, 2025Updated last year
- [NeurIPS 2025] MergeBench: A Benchmark for Merging Domain-Specialized LLMs☆50Aug 28, 2026Updated 3 weeks ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- Can VLMs understand students' hand-drawn math work?☆19Jan 20, 2026Updated 8 months ago
- Evaluate state-of-the-art sparse embedding models on the LIMIT dataset (`limit-small` and `limit`) from google's paper `On the Theoretica…☆16Sep 4, 2025Updated last year
- Performs benchmarking on two Korean datasets with minimal time and effort.☆48Aug 6, 2026Updated last month
- State-of-the-art paired encoder and decoder models (17M-1B params)☆79Aug 6, 2025Updated last year
- It shows how to deploy and use an agent with LLM.☆19Mar 1, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Label shift estimation for transfer difficulty with Familiarity.☆10Feb 4, 2025Updated last year
- LLM-as-a-judge using G-eval Scratch☆15Oct 12, 2025Updated 11 months ago
- Code for the paper "Multi-Field Adaptive Retrieval," a research project on a semi-structured document retrieval☆19Feb 13, 2026Updated 7 months ago
- Nearly Inference Free Embeddings: make your RAG queries 500x faster☆86Apr 27, 2026Updated 4 months ago
- Embedding Inversion via Conditional Masked Diffusion: recover original text from embedding vectors using parallel denoising. Live demo + …☆61Mar 7, 2026Updated 6 months ago
- Compression for unit-norm embedding vectors using spherical coordinates☆84Jan 23, 2026Updated 7 months ago
- A Multilingual Keyboard Layout-Based Typo Generator☆17Nov 23, 2025Updated 9 months ago
- Identify which embedding model produced a vector using digit-level tokenization and a tiny transformer☆23Mar 7, 2026Updated 6 months ago
- [ACL'26 Workshop] KoViDoRe: Korean Visual Document Retrieval Benchmark☆26Jul 2, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"☆19Mar 15, 2024Updated 2 years ago
- ☆23May 30, 2025Updated last year
- A modular framework for training and inference of (compressed) multi-vector retrieval across any modality.☆22Apr 4, 2026Updated 5 months ago
- Korean Sentence Embedding Model Performance Benchmark for RAG☆49Jan 27, 2025Updated last year
- KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models☆25Aug 24, 2024Updated 2 years ago
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated last year
- ☆18Jan 6, 2025Updated last year