☆51May 31, 2024Updated 2 years ago
Alternatives and similar repositories for PruneGPT
Users that are interested in PruneGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆42Apr 23, 2024Updated 2 years ago
- The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.☆14Mar 30, 2024Updated 2 years ago
- ☆75Sep 5, 2023Updated 3 years ago
- [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models☆11Dec 13, 2023Updated 2 years ago
- various experiments for scaling inference time compute with small reasoning models☆17Jan 16, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Utilities for efficient fine-tuning, inference and evaluation of code generation models☆21Oct 3, 2023Updated 2 years ago
- Official code for ACL 2023 (short, findings) paper "Recursion of Thought: A Divide and Conquer Approach to Multi-Context Reasoning with L…☆45Jun 13, 2023Updated 3 years ago
- A sleek, customizable interface for managing LLMs with responsive design and easy agent personalization.☆19Aug 30, 2024Updated 2 years ago
- Using modal.com to process FineWeb-edu data☆21Sep 12, 2026Updated last week
- LLM backed Fantasy Tribe Game☆19Nov 21, 2024Updated last year
- a collection of resources around LLMs, aggregated for the workshop "Mastering LLMs: End-to-End Fine-Tuning and Deployment" by Dan Becker …☆108May 31, 2024Updated 2 years ago
- A fork of the PEFT library, supporting Robust Adaptation (RoSA)☆15Aug 16, 2024Updated 2 years ago
- ☆21Jan 25, 2025Updated last year
- "FiD-ICL: A Fusion-in-Decoder Approach for Efficient In-Context Learning" (ACL 2023)☆15Jul 24, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Comparison of Language Model Inference Engines☆244Dec 16, 2024Updated last year
- Evolutionary Search for expert-level performance on any task with environmental feedback☆14Oct 12, 2025Updated 11 months ago
- Like system requirements lab but for LLMs☆31Jun 10, 2023Updated 3 years ago
- 如需体验textin文档解析,请点击https://cc.co/16YSIy☆21Jul 9, 2024Updated 2 years ago
- SPLAA is an AI assistant framework that utilizes voice recognition, text-to-speech, and tool-calling capabilities to provide a conversati…☆29May 6, 2025Updated last year
- A data visualisation of a 100 responses when asking local LLMs to imagine a random person.☆24Nov 4, 2024Updated last year
- Proteus is an experimental platform that combines the power of Large Language Models with the Genesis physics engine☆25Dec 20, 2024Updated last year
- A Pointer Generator with a BERT encoder☆10Aug 12, 2019Updated 7 years ago
- Aspose.Email for Python via .NET Examples: https://products.aspose.com/email/python-net☆10Oct 9, 2025Updated 11 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆36Aug 14, 2024Updated 2 years ago
- This is the React sample used in the ZITADEL quick start guide.☆11Apr 13, 2026Updated 5 months ago
- A simple library for working with Hugging Face models.☆14Dec 30, 2024Updated last year
- Copy a bunch of files into your clipboard to provide context for LLMs☆115Feb 8, 2026Updated 7 months ago
- Distributed Inference for mlx LLm☆103Aug 1, 2024Updated 2 years ago
- A Realtime App to visualize votes on who folks think will die in Episode 3 of Game of Thrones Season 8. Built using Vue.js, Hasura and C…☆14Dec 9, 2022Updated 3 years ago
- a fast and customizable CUDA int4 tensor core gemm☆15Aug 2, 2024Updated 2 years ago
- Demo of an "always-on" AI assistant.☆23Feb 14, 2024Updated 2 years ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆267Dec 4, 2025Updated 9 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆52Oct 29, 2024Updated last year
- When real time Yoga Position classification meets GNN☆11Sep 17, 2023Updated 3 years ago
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆37Sep 13, 2026Updated last week
- ☆18Jul 13, 2024Updated 2 years ago
- A simple no-install web UI for Ollama and OAI-Compatible APIs!☆31Jan 30, 2025Updated last year
- Attempt at cog wrapper for segmind/SSD-1B☆10Dec 11, 2023Updated 2 years ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 6 months ago