The code for paper "LLM-Neo: Parameter Efficient Knowledge Distillation for Large Language Models"
☆17Mar 2, 2025Updated last year
Alternatives and similar repositories for LLM-Neo
Users that are interested in LLM-Neo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆36Jan 20, 2026Updated 8 months ago
- ☆24Mar 7, 2025Updated last year
- Repo for the EMNLP'24 Paper "Dual-Space Knowledge Distillation for Large Language Models". A general white-box KD framework for both same…☆65Mar 21, 2026Updated 5 months ago
- PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation☆16Mar 28, 2023Updated 3 years ago
- Unofficial implementation of https://arxiv.org/pdf/2407.14679☆54Sep 7, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12May 20, 2022Updated 4 years ago
- Codebase of 'From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model'☆45Jun 27, 2026Updated 2 months ago
- Source code repo for paper "TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation"☆10Aug 11, 2023Updated 3 years ago
- LongAttn :Selecting Long-context Training Data via Token-level Attention☆15Jul 16, 2025Updated last year
- Learning MLPs to replace GNN☆10Jun 3, 2023Updated 3 years ago
- [ACL 2025] Cautious Next Token Prediction☆16Jul 24, 2025Updated last year
- ☆17Feb 21, 2025Updated last year
- Top 9 private leaderboard & Top 17 public leaderboard☆10Dec 1, 2022Updated 3 years ago
- ☆21Jul 3, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [NLPCC 2021] Shared Task on AutoIE2: Sub-Event Identification☆14Jul 19, 2021Updated 5 years ago
- ☆34Sep 14, 2024Updated 2 years ago
- [EMNLP 2023] Question Answering as Programming for Solving Time-Sensitive Questions☆12Dec 18, 2023Updated 2 years ago
- A Lightweight Multi-modality Image Segmentation Network via Domain Adaptation using Gradient Magnitude and Shape Constraint☆10Apr 3, 2023Updated 3 years ago
- Transform eyes into special Naruto forms using Dlib☆11Apr 21, 2020Updated 6 years ago
- [Preprint] Why is the State of Neural Network Pruning so Confusing? On the Fairness, Comparison Setup, and Trainability in Network Prunin…☆41Sep 9, 2025Updated last year
- EraX-VL-7B-V1 is the multimodal large language model developed by EraX team, base on Qwen2-VL.☆13Dec 31, 2024Updated last year
- LLM, Fine Tuning, Llama 2, Gemma, Mixtral, vLLM, LangChain, RAG, ChromaDB, FAISS☆13Mar 5, 2024Updated 2 years ago
- TabMini: A Benchmark Suite for Evaluating and Analyzing the Data Efficiency of Tabular Classifiers☆11Mar 31, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A multiphase field model based on machine learning method☆49Feb 10, 2022Updated 4 years ago
- Codes of the paper Deformable Butterfly: A Highly Structured and Sparse Linear Transform.☆16Nov 1, 2021Updated 4 years ago
- Toolkit for Bayesian scaling analysis☆14Sep 8, 2022Updated 4 years ago
- [ICML 2022] Learning Efficient and Robust Ordinary Differential \\ Equations via Invertible Neural Networks☆11Apr 14, 2023Updated 3 years ago
- EMMA [TMLR 2025]☆14Sep 25, 2025Updated 11 months ago
- This repository contains the scripts for reproducing the results presented in Costa AC, Ahamed T, Jordan D, Stephens GJ (2023) "A Markov…☆11Sep 25, 2025Updated 11 months ago
- Official Release of Multistep Quasimetric Estimation (MQE)☆21Mar 13, 2026Updated 6 months ago
- ☆18Jul 15, 2025Updated last year
- [CPAL 2026 oral] Offical implementation of "ROSE: Reordered SparseGPT for More Accurate One-Shot Large Language Models Pruning”☆17Jul 31, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Large Language Models Can Self-Improve in Long-context Reasoning☆72Nov 24, 2024Updated last year
- Links to recourses for the Lean Theorem Prover☆13Dec 3, 2019Updated 6 years ago
- ScisciJP2024で行ったチュートリアル。過去の論文で取り扱われることの多いトピックを取り上げ、定義、適用、注意点についてまとめ、再現実験のpythonコードを合わせて載せた☆17Jan 27, 2026Updated 7 months ago
- Tensorflow Custom Callbacks in Custom Training Loop☆10Apr 8, 2021Updated 5 years ago
- ☆126Jul 6, 2024Updated 2 years ago
- Revolutionize geospatial analysis with Swin-UNet – a cutting-edge solution for satellite imagery segmentation using Swin Transformers and…☆17Jun 6, 2026Updated 3 months ago
- ☆11Jan 23, 2017Updated 9 years ago