terashuf shuffles multi-terabyte text files using limited memory
☆232Feb 5, 2023Updated 3 years ago
Alternatives and similar repositories for terashuf
Users that are interested in terashuf are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆13Aug 23, 2024Updated last year
- Unsupervised factor-based text tokenizer for natural-language processing applications☆17Jul 24, 2020Updated 5 years ago
- This application shuffles the input file lines skipping (optionaly) the header. It's optimized for files bigger than available RAM.☆25Jan 9, 2017Updated 9 years ago
- Efficient teacher-student models and scripts to make them☆57Dec 16, 2023Updated 2 years ago
- Library for fast text representation and classification.☆31Jan 9, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Bicleaner fork that uses neural networks☆40Feb 23, 2026Updated 4 months ago
- Line shuffler for huge text file which does not fit in memory☆13Dec 1, 2022Updated 3 years ago
- Efficient Low-Memory Aligner☆148Jan 15, 2025Updated last year
- In-BoXBART: Get Instructions into Biomedical Multi-task Learning☆15Aug 23, 2022Updated 3 years ago
- Optimized inference with Ascend and Hugging Face☆12Apr 23, 2024Updated 2 years ago
- ☆14May 15, 2020Updated 6 years ago
- Twitter bot for generating photo descriptions (alt text)☆23Jul 1, 2021Updated 5 years ago
- Unsupervised text tokenizer focused on computational efficiency☆979Mar 29, 2024Updated 2 years ago
- A library for data streaming and augmentation☆22May 5, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pipeline for pulling and processing online language model pretraining data from the web☆179Jul 31, 2023Updated 2 years ago
- Randomly sample lines from massive text files efficiently☆16Apr 1, 2015Updated 11 years ago
- ☆93Feb 13, 2024Updated 2 years ago
- A utility for storing and reading files for Korean LM training 💾☆35Updated this week
- ☆21Feb 13, 2023Updated 3 years ago
- This repo contains details about USSR Eastern Front WWII veterans dataset extracted from Pamyat Naroda website☆13May 8, 2020Updated 6 years ago
- Tool to fix bitexts and tag near-duplicates for removal☆35Sep 4, 2025Updated 10 months ago
- ☆33Nov 22, 2021Updated 4 years ago
- FastFormers - highly efficient transformer models for NLU☆706Mar 21, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Efficient Sentence Embedding via Semantic Subspace Analysis☆14Feb 25, 2020Updated 6 years ago
- Tools to download and cleanup Common Crawl data☆1,046Apr 25, 2023Updated 3 years ago
- Korean Named Entity Corpus☆25May 12, 2023Updated 3 years ago
- Microsoft Speech Language Translation (MSLT) Corpus☆19Sep 18, 2017Updated 8 years ago
- ☆56May 12, 2018Updated 8 years ago
- Python scripts and other resources for tesing DetectNet on Nvidia DIGITS☆14Oct 10, 2017Updated 8 years ago
- All-in-one text de-duplication☆764Mar 9, 2026Updated 4 months ago
- ↔️ T5 Machine Translation from English to Korean☆18Aug 11, 2022Updated 3 years ago
- Bicleaner is a parallel corpus classifier/cleaner that aims at detecting noisy sentence pairs in a parallel corpus.☆160Jun 18, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆43Sep 16, 2020Updated 5 years ago
- Repository collecting resources and best practices to improve experimental rigour in deep learning research.☆27Mar 30, 2023Updated 3 years ago
- Code and data accompanying our ACL 2020 paper, "Unsupervised Domain Clusters in Pretrained Language Models".☆58Aug 22, 2020Updated 5 years ago
- ☆11Oct 3, 2021Updated 4 years ago
- Easy to use util for profiling in production☆11Aug 3, 2023Updated 2 years ago
- A library for preparing data for machine translation research (monolingual preprocessing, bitext mining, etc.) built by the FAIR NLLB te…☆309Updated this week
- Codes and files for the paper Are Emergent Abilities in Large Language Models just In-Context Learning☆33Jan 9, 2025Updated last year