Data augmentation for NLP
β4,663Jul 22, 2026Updated this week
Alternatives and similar repositories for nlpaug
Users that are interested in nlpaug are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TextAttack π is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocsβ¦β3,453Apr 17, 2026Updated 3 months ago
- Data augmentation for NLP, presented at EMNLP 2019β1,651Mar 19, 2023Updated 3 years ago
- State-of-the-Art Embeddings, Retrieval, and Rerankingβ18,944Updated this week
- Collection of papers and resources for data augmentation for NLP.β835Aug 12, 2022Updated 3 years ago
- TextAugment: Text Augmentation Libraryβ443Mar 4, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Beyond Accuracy: Behavioral Testing of NLP models with CheckListβ2,052Jan 9, 2024Updated 2 years ago
- NL-Augmenter π¦ β π A Collaborative Repository of Natural Language Transformationsβ786May 19, 2024Updated 2 years ago
- A very simple framework for state-of-the-art Natural Language Processing (NLP)β14,382Oct 27, 2025Updated 8 months ago
- Facebook AI Research Sequence-to-Sequence Toolkit written in Python.β32,253Sep 30, 2025Updated 9 months ago
- Unsupervised Data Augmentation (UDA)β2,205Aug 28, 2021Updated 4 years ago
- Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the moβ¦β22,956Jul 28, 2024Updated last year
- [EMNLP 2021] SimCSE: Simple Contrastive Learning of Sentence Embeddings https://arxiv.org/abs/2104.08821β3,655Oct 16, 2024Updated last year
- Transformers for Information Retrieval, Text Classification, NER, QA, Language Modelling, Language Generation, T5, Multi-Modal, and Conveβ¦β4,253May 31, 2026Updated last month
- An open-source NLP research library, built on PyTorch.β11,888Nov 22, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- BertViz: Visualize Attention in Transformer Modelsβ8,134Jan 8, 2026Updated 6 months ago
- The Learning Interpretability Tool: Interactively analyze ML models to understand their behavior in an extensible and framework agnostic β¦β3,658Jul 7, 2026Updated 2 weeks ago
- skweak: A software toolkit for weak supervision applied to NLP tasksβ925Sep 2, 2024Updated last year
- This repository contains the code for "Exploiting Cloze Questions for Few-Shot Text Classification and Natural Language Inference"β1,625Jun 12, 2023Updated 3 years ago
- π Scalable embedding, reasoning, ranking for images and sentences with CLIPβ12,832Jan 23, 2024Updated 2 years ago
- Leveraging BERT and c-TF-IDF to create easily interpretable topics.β7,756May 13, 2026Updated 2 months ago
- Open source annotation tool for machine learning practitioners.β10,715Apr 14, 2026Updated 3 months ago
- π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal modelβ¦β162,967Updated this week
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generatorsβ2,367Mar 23, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Longformer: The Long-Document Transformerβ2,201Feb 8, 2023Updated 3 years ago
- A data augmentations library for audio, image, text, and video.β5,087Jul 16, 2026Updated last week
- Unsupervised text tokenizer for Neural Network-based text generation.β11,983Updated this week
- Code for the paper "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer"β6,537Jul 8, 2026Updated 2 weeks ago
- Multi-Task Deep Neural Networks for Natural Language Understandingβ2,259Mar 7, 2024Updated 2 years ago
- Language-Agnostic SEntence Representationsβ3,661May 2, 2024Updated 2 years ago
- A system for quickly generating training data with weak supervisionβ5,994Jun 8, 2026Updated last month
- Explain, analyze, and visualize NLP language models. Ecco creates interactive visualizations directly in Jupyter notebooks explaining theβ¦β2,101Aug 15, 2024Updated last year
- Minimal keyword extraction with BERTβ4,207May 13, 2026Updated 2 months ago
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languagesβ7,853Updated this week
- PyTorch original implementation of Cross-lingual Language Model Pretraining.β2,923Feb 14, 2023Updated 3 years ago
- Super easy library for BERT based NLP modelsβ1,918Aug 19, 2024Updated last year
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,171Jan 23, 2026Updated 6 months ago
- SentAugment is a data augmentation technique for NLP that retrieves similar sentences from a large bank of sentences. It can be used in cβ¦β359Feb 22, 2022Updated 4 years ago
- A Unified Library for Parameter-Efficient and Modular Transfer Learningβ2,822Apr 26, 2026Updated 2 months ago
- Must-read Papers on pre-trained language models.β3,360Nov 6, 2022Updated 3 years ago