[EMNLP 2022/2023] Fast Vocabulary Transfer & Multi-word Tokenization
☆27Jan 19, 2025Updated last year
Alternatives and similar repositories for fast-vocabulary-transfer
Users that are interested in fast-vocabulary-transfer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Show the time in Roman Numerals☆12Jan 23, 2020Updated 6 years ago
- Pytorch ImageNet1k Loader with Bounding Boxes.☆13Jan 23, 2022Updated 4 years ago
- Cosine Similary Search in ElasticSearch + FAISS GPU☆12Mar 24, 2022Updated 4 years ago
- Codebase for VAEL: Bridging Variational Autoencoders and Probabilistic Logic Programming☆26Jun 30, 2023Updated 3 years ago
- Code for the paper "Getting the most out of your tokenizer for pre-training and domain adaptation"☆22Feb 14, 2024Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- CSE201 Objected-Oriented Programming in C++: Teach an AI to produce pieces of music☆13Jan 23, 2019Updated 7 years ago
- This repository contains code used for our Multi Sentence Inference NAACL'22 paper.☆12Mar 6, 2023Updated 3 years ago
- Evaluate Transformers from the Hub 🔥☆14May 26, 2026Updated 3 months ago
- ☆18Mar 26, 2022Updated 4 years ago
- Perf monitoring CLI tool for Apple Silicon☆10Jan 25, 2023Updated 3 years ago
- ☆16Mar 4, 2024Updated 2 years ago
- Get up in the morning by striking a pose to stop your alarm from ringing.☆12Jun 9, 2021Updated 5 years ago
- A set of utilities to turn Dataclasses into useful configuration managers.☆11Mar 27, 2024Updated 2 years ago
- ☆15Oct 4, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Defeasible Natural Language Inference☆14Dec 4, 2020Updated 5 years ago
- Ongoing research training transformer models at scale☆18Jul 27, 2023Updated 3 years ago
- Font style transfer for Devanāgarī script using GANs☆13Jun 25, 2022Updated 4 years ago
- Random Number Generator NIST Test Suite framework for python 3.6 - SAILab - University of Siena☆56Mar 16, 2021Updated 5 years ago
- Projeto de aplicativo Web a fim de exibição e controle de animações 3D☆15Nov 15, 2023Updated 2 years ago
- Coala is a python package for Contextual Answer Sentence Selection.☆15Jun 12, 2023Updated 3 years ago
- Implementation of SENets by chainer (Squeeze-and-Excitation Networks: https://arxiv.org/abs/1709.01507)☆15Sep 15, 2017Updated 8 years ago
- CUDA keyring packaging for Debian☆14Apr 14, 2023Updated 3 years ago
- [NeurIPS 2024 D&B Track] DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation☆14Mar 5, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Code for Zero-Shot Tokenizer Transfer☆147Jan 14, 2025Updated last year
- C# bindings for llama.cpp for Unity☆18Dec 13, 2024Updated last year
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- This is the repo for constructing a comprehensive and rigorous evaluation framework for LLM calibration.☆14Apr 9, 2024Updated 2 years ago
- CaMML:Context-Aware MultiModal Learner for Large Models (ACL 2024 SAC Award)☆15May 21, 2025Updated last year
- Code and data from the paper 'Human Feedback is not Gold Standard'☆20Aug 9, 2026Updated 3 weeks ago
- ☆12Jun 30, 2024Updated 2 years ago
- ☆19Mar 12, 2026Updated 5 months ago
- CLIR version of ColBERT☆72May 28, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Repository for Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard Contexts, EMNLP22☆19Jun 23, 2023Updated 3 years ago
- [ICLR 2024] "Data Distillation Can Be Like Vodka: Distilling More Times For Better Quality" by Xuxi Chen*, Yu Yang*, Zhangyang Wang, Baha…☆15May 18, 2024Updated 2 years ago
- PyTorch implementation of the original evidental-deep-learning@https://github.com/aamini/evidential-deep-learning/☆13Sep 20, 2021Updated 4 years ago
- ☆14Mar 25, 2019Updated 7 years ago
- Diverse Demonstrations Improve In-context Compositional Generalization☆13Jul 7, 2023Updated 3 years ago
- Code for "Unlearning Traces the Influential Training Data of Language Models"☆13Jun 13, 2024Updated 2 years ago
- Official Code Repository for [AutoScale📈: Scale-Aware Data Mixing for Pre-Training LLMs] Published as a conference paper at **COLM 2025*…☆14Aug 8, 2025Updated last year