☆76Nov 22, 2024Updated last year
Alternatives and similar repositories for DIOPI
Users that are interested in DIOPI are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆76Oct 31, 2024Updated last year
- ☆13May 23, 2025Updated last year
- A benchmark suited especially for deep learning operators☆42Feb 13, 2023Updated 3 years ago
- ☆31Jan 7, 2025Updated last year
- Composable and Embeddable Communication Runtime for Distributed AI Services☆102Jun 5, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- CVFusion is an open-source deep learning compiler to fuse the OpenCV operators.☆33Aug 31, 2022Updated 3 years ago
- [DAC2024] A Holistic Functionalization Approach to Optimizing Imperative Tensor Programs in Deep Learning☆15Jan 13, 2024Updated 2 years ago
- triton for dsa☆68Jul 10, 2026Updated last week
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆22May 5, 2026Updated 2 months ago
- LightRFT (Light Reinforcement Fine-Tuning) is an advanced reinforcement learning fine-tuning framework designed for Large Language Models…☆19Jan 12, 2026Updated 6 months ago
- ☆35Mar 27, 2026Updated 3 months ago
- Sequence-level 1F1B schedule for LLMs.☆37Aug 26, 2025Updated 10 months ago
- A demonstrative example of running SGLang Diffusion with DP router☆17Mar 15, 2026Updated 4 months ago
- ☆17Jan 1, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆74Updated this week
- ☆26Mar 31, 2026Updated 3 months ago
- The code about TC-Bench and CircuitMind☆17Jun 7, 2025Updated last year
- torch_remat fine-grained activation checkpointing API☆15Updated this week
- InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencie…☆421Aug 21, 2025Updated 11 months ago
- [Archived] For the latest updates and community contribution, please visit: https://github.com/Ascend/TransferQueue or https://gitcode.co…☆16Jan 16, 2026Updated 6 months ago
- This is a comprehensive guide on how you can automate your feature engineering process.☆11Jun 25, 2018Updated 8 years ago
- Early-stage Rust drop-in alternative frontend for vLLM☆72May 22, 2026Updated last month
- mllm-npu: training multimodal large language models on Ascend NPUs☆95Aug 29, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official repository for the paper DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines☆19Dec 8, 2023Updated 2 years ago
- ☆16Mar 4, 2026Updated 4 months ago
- A benchmarking tool for comparing different LLM API providers' DeepSeek model deployments.☆31Mar 28, 2025Updated last year
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆23Apr 10, 2026Updated 3 months ago
- A benchmark and playground for Completely Fair Scheduling in Go☆11Feb 12, 2022Updated 4 years ago
- a data collection of related work: Toward Understanding Deep Learning Framework Bugs☆18Oct 23, 2023Updated 2 years ago
- A prefill & decode disaggregated LLM serving framework with shared GPU memory and fine-grained compute isolation.☆127Dec 25, 2025Updated 6 months ago
- Convert a dynamically linked binary to a statically linked binary going thorugh LLVM IR, using mcsema☆12May 27, 2019Updated 7 years ago
- an implementation of parallel skills like amp, ddp, pp, tp for learning purposes☆14Nov 18, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A model compilation solution for various hardware☆473Aug 20, 2025Updated 11 months ago
- Allow torch tensor memory to be released and resumed later☆259Updated this week
- A throughput-oriented high-performance serving framework for LLMs☆968Mar 29, 2026Updated 3 months ago
- lightsocks client implements by golang☆13Sep 11, 2015Updated 10 years ago
- 中文版 Parallel Programming for FPGAs☆15Nov 20, 2019Updated 6 years ago
- Simple intermediate representation language for learning and research.☆22Mar 27, 2020Updated 6 years ago
- Multimodal RAG using LlamaIndex, Qdrant, llama.cpp for document QA with local VisonLLM and embedding models☆20Nov 8, 2024Updated last year