Datamodels for hugging face tokenizers
☆111Aug 11, 2026Updated last week
Alternatives and similar repositories for skeletoken
Users that are interested in skeletoken are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Nearly Inference Free Embeddings: make your RAG queries 500x faster☆84Apr 27, 2026Updated 3 months ago
- Code for SaGe subword tokenizer (EACL 2023)☆28Nov 30, 2024Updated last year
- Label shift estimation for transfer difficulty with Familiarity.☆10Feb 4, 2025Updated last year
- Trully flash implementation of DeBERTa disentangled attention mechanism.☆91Feb 10, 2026Updated 6 months ago
- A Multilingual Keyboard Layout-Based Typo Generator☆17Nov 23, 2025Updated 8 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Pre-train Static Word Embeddings☆110Jun 9, 2026Updated 2 months ago
- Check what an AI agent can access before you run it☆27Mar 8, 2026Updated 5 months ago
- [ICML'26] LEMUR reduces multi-vector retrieval for late interaction models such as ColBERT into regular single-vector retrieval.☆32Jun 21, 2026Updated last month
- ☆110Jun 2, 2025Updated last year
- Widgets that can automatically refresh in marimo☆32Nov 10, 2025Updated 9 months ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆19Aug 11, 2026Updated last week
- 🔢 Work with static vector models☆39Apr 21, 2025Updated last year
- ☆103Jul 4, 2025Updated last year
- Generalist and Lightweight Model for Text Classification☆240Jul 21, 2026Updated 3 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ANE accelerated embedding models!☆20Dec 11, 2024Updated last year
- State-of-the-art paired encoder and decoder models (17M-1B params)☆77Aug 6, 2025Updated last year
- C inference engine for running GLiClass (Generalist and Lightweight Classification) models☆17May 21, 2025Updated last year
- Late Interaction Models Training & Retrieval☆878Jul 23, 2026Updated 3 weeks ago
- Official code release for "SuperBPE: Space Travel for Language Models"☆98May 28, 2026Updated 2 months ago
- Official Rust Implementation of Model2Vec☆207May 24, 2026Updated 2 months ago
- Using open source LLMs to build synthetic datasets for direct preference optimization☆72Feb 29, 2024Updated 2 years ago
- NLP with Rust for Python 🦀🐍☆72Aug 8, 2026Updated last week
- Simply, faster, sentence-transformers☆144Aug 27, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Fast Diversification for Search & Retrieval☆496May 24, 2026Updated 2 months ago
- PyLate efficient inference engine☆88Jan 7, 2026Updated 7 months ago
- Fast Multimodal Semantic Deduplication & Filtering☆960May 24, 2026Updated 2 months ago
- My NER Experiments with ModernBERT and Ettin☆29Jul 17, 2025Updated last year
- Efficient and scalable zero-shot entity linking☆147Jul 20, 2026Updated 3 weeks ago
- bb25 is a fast, self-contained BM25 + Bayesian calibration implementation with a minimal Python API.☆149Mar 17, 2026Updated 5 months ago
- k8s operator and plugin for marimo deployment☆26Aug 10, 2026Updated last week
- Crispy reranking models by Mixedbread☆54Sep 17, 2025Updated 11 months ago
- Compression for unit-norm embedding vectors using spherical coordinates☆83Jan 23, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A massively multilingual modern encoder language model☆150Jan 20, 2026Updated 6 months ago
- 🤝 Trade any tensors over the network☆31Sep 27, 2023Updated 2 years ago
- This repository contains the code for the Form-Context Model and its Attentive Mimicking variant.☆30May 11, 2020Updated 6 years ago
- Optimus is a flexible and scalable framework built to train language models efficiently across diverse hardware configurations, including…☆70Dec 4, 2025Updated 8 months ago
- Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning☆30Jan 25, 2023Updated 3 years ago
- 👩🤝🤖 A curated list of datasets for large language models (LLMs), RLHF and related resources (continually updated)☆25May 2, 2023Updated 3 years ago
- A Rust rewrite of FastKMeans for CPU-based clustering☆16Jun 29, 2026Updated last month