☆21Jul 22, 2026Updated 2 weeks ago
Alternatives and similar repositories for DyLLM
Users that are interested in DyLLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆22Feb 26, 2023Updated 3 years ago
- ☆59Oct 14, 2025Updated 9 months ago
- ☆117Jul 4, 2024Updated 2 years ago
- Layered prefill changes the scheduling axis from tokens to layers and removes redundant MoE weight reloads while keeping decode stall fre…☆20Mar 9, 2026Updated 5 months ago
- Cheddar: A Swift Fully Homomorphic Encryption (FHE) GPU Library☆95Apr 9, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Mar 18, 2024Updated 2 years ago
- ☆164Jun 24, 2024Updated 2 years ago
- ☆11Aug 23, 2023Updated 2 years ago
- This is where gem5 based DRAM cache models live.☆20Mar 23, 2023Updated 3 years ago
- Artifact for "Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs" [arXiv '25]☆21Jul 26, 2026Updated 2 weeks ago
- ☆15Jan 7, 2023Updated 3 years ago
- Benchmarking OpenBLAS on the Apple M1☆17Dec 31, 2020Updated 5 years ago
- CNN simd based accelerator using Vitis HLS☆11Jul 15, 2022Updated 4 years ago
- DRAM Bender is the first open source DRAM testing infrastructure that can be used to easily and comprehensively test state-of-the-art HBM…☆132Jun 1, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ⚡ A seamless integration of HuggingFace Transformers & Diffusers with RBLN SDK for efficient inference on RBLN NPUs.☆19Updated this week
- Official PyTorch implementation of "LGViT: Dynamic Early Exiting for Accelerating Vision Transformer" (ACM MM 2023)☆16Nov 18, 2024Updated last year
- This is the respository that holds the artifacts of MICRO'23 -- Demystifying CXL Memory with True CXL-Ready Systems and CXL Memory Device…☆53Mar 17, 2024Updated 2 years ago
- The set of AI agent model implementations, benchmarks, and others used in our paper "The Cost of Dynamic Reasoning: Demystifying AI Agent…☆43Mar 26, 2026Updated 4 months ago
- ☆16Mar 10, 2026Updated 5 months ago
- Open Source Projects from Pallas Lab☆21Oct 10, 2021Updated 4 years ago
- A user level library for applications to transparently use Intel DSA.☆43Updated this week
- mNPUsim: A Cycle-accurate Multi-core NPU Simulator (IISWC 2023)☆77Dec 29, 2025Updated 7 months ago
- Codebase for layer wise N:M pruning pattern assignment for LLMs☆15Aug 5, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ECCV 2024] CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs☆20Jul 2, 2024Updated 2 years ago
- Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective☆19Jul 15, 2025Updated last year
- A Caffe2 implementation of the YOLO v3 object detection algorithm☆30Dec 7, 2018Updated 7 years ago
- Selected problems and their solutions from the book on "Machine Intelligence in Design Automation"☆27Dec 9, 2018Updated 7 years ago
- Dynamically Reconfigurable Architecture Template and Cycle-level Microarchitecture Simulator for Dataflow AcCelerators☆29Jul 17, 2023Updated 3 years ago
- The official implementation of the DAC 2024 paper GQA-LUT☆25Dec 20, 2024Updated last year
- ☆23Apr 3, 2026Updated 4 months ago
- Code for ICLR2022 Decoupled Adaptation for Cross-Domain Object Detection (D-adapt) https://arxiv.org/abs/2110.02578☆30Mar 14, 2022Updated 4 years ago
- ESL-CGRA-simulator☆22Apr 10, 2026Updated 3 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆83Jun 23, 2025Updated last year
- ☆41Jul 28, 2026Updated last week
- [ACL 2026 Main] Code for the paper "ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs"☆29Jun 1, 2026Updated 2 months ago
- Artifacts for "ZenHammer: Rowhammer Attacks on AMD Zen-based Platforms" (USENIX Security '24).☆64Jun 19, 2025Updated last year
- (Not actively updating)Vision Transformer Accelerator implemented in Vivado HLS for Xilinx FPGAs.☆26Dec 29, 2024Updated last year
- vLLM plugin for RBLN NPU☆57Updated this week
- ☆26Jun 12, 2026Updated last month