Tools for managing datasets for governance and training.
☆91Aug 24, 2026Updated last month
Alternatives and similar repositories for data_tooling
Users that are interested in data_tooling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code used for sourcing and cleaning the BigScience ROOTS corpus☆318Mar 20, 2023Updated 3 years ago
- ☆13Aug 23, 2024Updated 2 years ago
- Personal information identification standard☆21Jan 24, 2024Updated 2 years ago
- Generate BERT vocabularies and pretraining examples from Wikipedias☆17May 11, 2020Updated 6 years ago
- Submission for AIviVN sentiment analysis contest https://www.aivivn.com/contests/1☆15Oct 12, 2021Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- All-in-one text de-duplication☆762Mar 9, 2026Updated 6 months ago
- Thử nghiệm gần đây mô hình MLP-Mixer trên bài toán nhận diện cảm xúc (Sentiment sentiment analysis)☆13Jul 9, 2021Updated 5 years ago
- machine translation data process tools☆10Apr 29, 2024Updated 2 years ago
- Library for fast text representation and classification.☆31Jan 9, 2024Updated 2 years ago
- Code and Data for Evaluation WG☆42May 4, 2022Updated 4 years ago
- Central place for the engineering/scaling WG: documentation, SLURM scripts and logs, compute environment and data.☆1,020Jul 29, 2024Updated 2 years ago
- Multilingual bert retrained on news + squad2 for vietnamese☆23Feb 16, 2020Updated 6 years ago
- a ducttape workflow for neural machine translation☆14Mar 23, 2021Updated 5 years ago
- Pre-training script for BART in JAX/Flax☆37Aug 4, 2022Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- My PhD manuscript LaTeX code and the slides for the defense☆11Feb 2, 2022Updated 4 years ago
- I-SHEEP: Iterative Self-enHancEmEnt Paradigm of LLMs through Self-Instruct and Self-Assessment☆17Jan 16, 2025Updated last year
- ☆78Dec 7, 2023Updated 2 years ago
- Pipeline for pulling and processing online language model pretraining data from the web☆179Jul 31, 2023Updated 3 years ago
- A utility for storing and reading files for Korean LM training 💾☆35Jul 18, 2026Updated 2 months ago
- ☆12Dec 9, 2015Updated 10 years ago
- a compact audio-to-phoneme aligner for singing voice☆12Jan 17, 2024Updated 2 years ago
- This directory gathers the tools developed by the Data Sourcing Working Group☆31Oct 25, 2021Updated 4 years ago
- UFSAC is a resource containing all WordNet Sense Annotated Corpora, and a Java library for manipulating them☆39May 17, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of ECCV24 paper: POA☆24Aug 8, 2024Updated 2 years ago
- The pipeline for the OSCAR corpus☆178Nov 9, 2025Updated 10 months ago
- An ensemble system with a search engine for relevant document retrieval and a deep learning model (BERT) for machine comprehension in Vie…☆14Oct 17, 2019Updated 6 years ago
- ☆12Apr 3, 2026Updated 5 months ago
- Generative models of chemical data for PaccMann^RL☆13Jun 2, 2023Updated 3 years ago
- ☆13Jan 18, 2020Updated 6 years ago
- ☆1,270Jul 30, 2024Updated 2 years ago
- A fully differentiable set autoencoder☆17Apr 3, 2024Updated 2 years ago
- Machine Reading Comprehension special for the Vietnamese language☆41Mar 13, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 本项目主要对开源的MOSS SFT数据进行整理 ,转换成mnbvc多轮对话格式。MOSS-003涵盖用性、忠实性、无害性三个层面,共353w样本,MOSS-003 包含更细粒度的有用性类别标记、更广泛的无害性数据和更长对话轮数,共630w样本,☆13Dec 3, 2023Updated 2 years ago
- Preprocessing of datasets of chemical reactions: standardization, filtering, augmentation, tokenization, etc.☆17Sep 10, 2025Updated last year
- Anh - LAION's multilingual assistant datasets and models☆28Apr 5, 2023Updated 3 years ago
- kogpt를 oslo로 파인튜닝하는 예제.☆23Aug 26, 2022Updated 4 years ago
- A library for squeakily cleaning and filtering language datasets.☆50Jul 10, 2023Updated 3 years ago
- ☆94Jul 16, 2022Updated 4 years ago
- Các thí nghiệm liên quan tới LLMs cho tiếng Việt (insprised by Physics of LLMs Series)☆11Oct 21, 2024Updated last year