Wikipedia text corpus for self-supervised NLP model training
☆48Jul 17, 2022Updated 4 years ago
Alternatives and similar repositories for wikipedia2corpus
Users that are interested in wikipedia2corpus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is a german text corpus from Wikipedia. It is cleaned, preprocessed and sentence splitted. It's purpose is to train NLP embeddings l…☆24Feb 22, 2022Updated 4 years ago
- German word embeddings computed from a corpus of parliamentary transcripts (2017-2019)☆14Mar 5, 2020Updated 6 years ago
- German Alpaca Dataset (Cleaned + Translated)☆26Apr 6, 2023Updated 3 years ago
- Klexikon: A German Dataset for Joint Summarization and Simplification☆17Oct 5, 2022Updated 3 years ago
- Brave is a simple visualisation library for NLP information extraction, built on top of embedded BRAT.☆15Dec 25, 2019Updated 6 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- BERT and ELECTRA models trained on Europeana Newspapers☆39Dec 14, 2021Updated 4 years ago
- X-SCITLDR: Cross-Lingual Extreme Summarization of Scholarly Documents (JCDL 2022)☆14Jul 22, 2022Updated 4 years ago
- Repo for the simplified text alignment tools.☆21Dec 4, 2020Updated 5 years ago
- Alignment and annotation for comparable documents.☆22Oct 16, 2018Updated 7 years ago
- Code for our TSD paper "TOKEN is a MASK: Few-shot Named Entity Recognition with Pre-trained Language Models"☆14Aug 19, 2022Updated 4 years ago
- DWIE (Deutsche Welle corpus for Information Extraction) dataset. Introduced in our "DWIE: an entity-centric dataset for multi-task docume…☆51Jul 23, 2023Updated 3 years ago
- Apache OpenNLP Models☆17Updated this week
- Plan and train German transformer models.☆23Feb 22, 2021Updated 5 years ago
- [ACL 20] Probing Linguistic Features of Sentence-level Representations in Neural Relation Extraction☆13Apr 21, 2020Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An SDK and Library that is used in several Deutsche Telekom mobile apps☆12Sep 23, 2024Updated 2 years ago
- Polish data.☆13May 6, 2026Updated 4 months ago
- A Multilingual Keyboard Layout-Based Typo Generator☆17Nov 23, 2025Updated 10 months ago
- ☆18Feb 1, 2023Updated 3 years ago
- ☆12Oct 17, 2022Updated 3 years ago
- MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer☆40Jun 7, 2022Updated 4 years ago
- Code accompanying the paper "Knowledge Base Completion Meets Transfer Learning"☆15Feb 21, 2024Updated 2 years ago
- A tokenizer and sentence splitter for German and English web and social media texts.☆155Aug 7, 2026Updated last month
- Ukrainian ELECTRA model☆12Mar 11, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- XWikisCorpus, cross-lingual summarisation, multi-lingual summarisation, pre-trained language models, zero-shot and few-shot summarisation…☆10Nov 4, 2022Updated 3 years ago
- ☆12Oct 2, 2022Updated 3 years ago
- Codebase, data and models for the Re-Thinking the Shuffle Test paper at ACL2021☆10Oct 14, 2022Updated 3 years ago
- Tools relating to the CC-News-En Collection☆20Dec 8, 2023Updated 2 years ago
- Pytorch implementation of Google TCAV☆10Jan 11, 2019Updated 7 years ago
- Compound splitter for German language ("Komposita-Zerlegung") based on large dictionary combined with highly efficient multi-pattern stri…☆36Jul 7, 2022Updated 4 years ago
- Pixel Propagation for unsupervised visual representation learning☆11Feb 16, 2021Updated 5 years ago
- Curriculum training☆24Jun 25, 2025Updated last year
- SMiLER - Samsung MultiLingual Entity and Relation Extraction dataset☆18Feb 11, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- process your massive word2vec binary model file as a readable stream of records☆11Jan 28, 2018Updated 8 years ago
- Gradient accumulation on tf.estimator☆12Dec 15, 2020Updated 5 years ago
- This repository is part of an NLP course for humanities and cultural studies. This course uses historical newspapers as a source and appl…☆20Jun 5, 2025Updated last year
- Breaks a word into syllables using an LSTM-based neural network.☆20Aug 14, 2023Updated 3 years ago
- Building an effective preprocessing tool for African languages☆13Jan 24, 2024Updated 2 years ago
- The NLPStatTest project☆12Mar 12, 2022Updated 4 years ago
- ☆11Dec 8, 2022Updated 3 years ago