Data preparation code for CrystalCoder 7B LLM
☆45May 10, 2024Updated 2 years ago
Alternatives and similar repositories for crystalcoder-data-prep
Users that are interested in crystalcoder-data-prep are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pre-training code for CrystalCoder 7B LLM☆59May 10, 2024Updated 2 years ago
- Data preparation code for Amber 7B LLM☆96May 10, 2024Updated 2 years ago
- Pre-training code for Amber 7B LLM☆176May 10, 2024Updated 2 years ago
- Open Implementations of LLM Analyses☆111Oct 8, 2024Updated last year
- A Jupyter notebook building and training a Neural Network from scratch with NumPy☆12Aug 7, 2022Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Curso de Deep Learning desde las bases de Python, Fundamentos del Machine Learning, Fundamentos Matematicos del ML y DL, Redes Neuronale…☆16Jun 24, 2021Updated 5 years ago
- A list where most values will be None (or default)☆11Updated this week
- This is the implementation of CounterCurate, the data curation pipeline of both physical and semantic counterfactual image-caption pairs.☆19Jun 27, 2024Updated 2 years ago
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 4 months ago
- ☆12Jul 25, 2023Updated 3 years ago
- INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness☆15Jun 2, 2026Updated 2 months ago
- This repository contains the replication package of our paper "Assessing the Security of GitHub Copilot’s Generated Code - A Targeted Rep…☆10Nov 16, 2023Updated 2 years ago
- [NAACL 2025] Representing Rule-based Chatbots with Transformers☆23Feb 9, 2025Updated last year
- ☆10Apr 15, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Oct 11, 2024Updated last year
- Minimum Description Length probing for neural network representations☆20Jan 28, 2025Updated last year
- ☆16Oct 2, 2024Updated last year
- Implementation of "LM-Infinite: Simple On-the-Fly Length Generalization for Large Language Models"☆40Nov 11, 2024Updated last year
- [Findings of EMNLP22] From Mimicking to Integrating: Knowledge Integration for Pre-Trained Language Models☆19Mar 16, 2023Updated 3 years ago
- Text-2-SQL☆19Feb 21, 2025Updated last year
- Official repository for "Reweighting Strategy based on Synthetic Data Identification for Sentence Similarity (COLING2022)"☆18Sep 4, 2022Updated 3 years ago
- Source code for paper: Knowledge Inheritance for Pre-trained Language Models☆37Apr 24, 2022Updated 4 years ago
- ☆20Dec 13, 2020Updated 5 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- a Fine-tuned LLaMA that is Good at Arithmetic Tasks☆178Sep 15, 2023Updated 2 years ago
- BH hackathon☆14Apr 4, 2024Updated 2 years ago
- OOPSLA 2019 Artifact for AutoPandas. Website at https://rbavishi.github.io/autopandas☆31Nov 21, 2022Updated 3 years ago
- awesome-LLM-controlled-constrained-generation☆57Aug 16, 2024Updated 2 years ago
- notes, config, tools, etc. for kicking the tires on cockroachdb☆11Apr 8, 2025Updated last year
- ☆11Oct 11, 2023Updated 2 years ago
- AIxCC: automated vulnerability repair via LLMs, search, and static analysis☆13Jul 16, 2024Updated 2 years ago
- Paper notes for my PhD on Machine Learning (mostly focused on Reinforcement Learning)☆17Jul 22, 2019Updated 7 years ago
- A Dataset of 600k Java Source Code Changes Categorized by Diff Size http://arxiv.org/pdf/2108.04631☆23Mar 22, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- A Universal Discriminator for Zero-Shot Generalization☆18Jun 21, 2023Updated 3 years ago
- ☆11Dec 8, 2016Updated 9 years ago
- QuoteSum is a textual QA dataset containing Semi-Extractive Multi-source Question Answering (SEMQA) examples written by humans, based on …☆13Mar 25, 2024Updated 2 years ago
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks☆18Nov 2, 2021Updated 4 years ago
- Source code for the paper "Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data"☆20Feb 24, 2024Updated 2 years ago
- Build a level 1 coding agent.☆17Jan 28, 2025Updated last year