Minimal code to train ELMo models in recent versions of TensorFlow
☆14Jun 16, 2026Updated 2 months ago
Alternatives and similar repositories for simple_elmo_training
Users that are interested in simple_elmo_training are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆13Mar 27, 2020Updated 6 years ago
- Highly specialized crate to parse and use `google/sentencepiece` 's precompiled_charsmap in `tokenizers`☆23Jun 9, 2026Updated 2 months ago
- Data and scripts for the proper evaluation of cross-lingual embeddings in multiple languages☆15Apr 11, 2020Updated 6 years ago
- Getting interpretable dimensions in word embedding spaces.☆15Jul 6, 2023Updated 3 years ago
- [ACL 2025] 🔍 Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment☆11Apr 6, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- bin files☆13Jan 30, 2025Updated last year
- Low-code pre-built pipelines for experiments with huggingface/transformers for Data Scientists in a rush.☆16Oct 14, 2020Updated 5 years ago
- [ICML 2026] Improving GPT via a simple normalization strategy☆15May 22, 2026Updated 3 months ago
- ☆32Apr 4, 2020Updated 6 years ago
- ☆30Sep 27, 2021Updated 4 years ago
- ☆16May 6, 2021Updated 5 years ago
- pair2vec: Compositional Word-Pair Embeddings for Cross-Sentence Inference☆61Dec 8, 2022Updated 3 years ago
- A software for transferring pre-trained English models to foreign languages☆20Mar 20, 2023Updated 3 years ago
- A neural network that jointly part-of-speech tags and lemmatizes sentences, boosting accuracy for morphologically-rich languages (Czech, …☆34Apr 5, 2019Updated 7 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆11Jan 20, 2020Updated 6 years ago
- ☆45Nov 3, 2019Updated 6 years ago
- Distribution of word meanings in Wikipedia for English, Italian, French, German and Spanish.☆10Jan 4, 2021Updated 5 years ago
- Python tools for performing various operations on ALTO XML files☆50Jun 12, 2026Updated 2 months ago
- Code for the paper "Getting the most out of your tokenizer for pre-training and domain adaptation"☆22Feb 14, 2024Updated 2 years ago
- Part-of-speech tagging using BERT☆10Nov 14, 2019Updated 6 years ago
- Code for bidirectional sequence generation (BiSon) for generating from BERT pre-trained models.☆51Mar 17, 2020Updated 6 years ago
- ☆17Aug 19, 2024Updated 2 years ago
- Temporary remove unused tokens during training to save ram and speed.☆23Jun 15, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Price options by fitting a Lévy distribution☆10Jan 20, 2021Updated 5 years ago
- Source code repo for paper "TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation"☆10Aug 11, 2023Updated 3 years ago
- Example showing how to access R's serialization functions from C☆13May 17, 2025Updated last year
- Umbrella repository that describes the collections contained in any given release of ELTeC☆13Jan 26, 2022Updated 4 years ago
- ☆105Jan 14, 2021Updated 5 years ago
- Use spaCy for NLP and output to the FoLiA XML format.☆12Feb 27, 2024Updated 2 years ago
- A simple fighting game for teaching Python☆14Dec 12, 2016Updated 9 years ago
- Материалы курса "Компьютерная лингвистика и информационные технологии" для 4-го курса бакалавриата направления "Фундаментальная и приклад…☆10Mar 25, 2021Updated 5 years ago
- Legacy version of CNN neural net toolkit (now called dynet)☆19Oct 8, 2016Updated 9 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆11Jul 15, 2020Updated 6 years ago
- SentAugment is a data augmentation technique for NLP that retrieves similar sentences from a large bank of sentences. It can be used in c…☆358Feb 22, 2022Updated 4 years ago
- Presentations, tutorials and data for the OCR workshop at LMU☆16Jun 2, 2017Updated 9 years ago
- An example of how to use spaCy for extremely large files without running into memory issues☆35Sep 17, 2022Updated 3 years ago
- Text-Induced Corpus Clean-up☆21Jun 20, 2023Updated 3 years ago
- [IJCAI'23] Semantic-aware Generation of Multi-view Portrait Drawings (SAGE)☆10Feb 25, 2024Updated 2 years ago
- Research code for the paper "How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models"☆28Oct 3, 2021Updated 4 years ago