Course materials for MIT6.5940: TinyML and Efficient Deep Learning Computing
☆85Jan 8, 2025Updated last year
Alternatives and similar repositories for MIT6.5940_TinyML
Users that are interested in MIT6.5940_TinyML are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Skeleton code for new 6.858 final project --- an encrypted and authenticated file system☆24Apr 20, 2022Updated 4 years ago
- Expert Specialization MoE Solution based on CUTLASS☆27Apr 14, 2026Updated 5 months ago
- The official implementation for the intra-stage fusion technique introduced in https://arxiv.org/abs/2409.13221☆32Apr 22, 2025Updated last year
- Some funny cute/cuteDSL code snippets☆35Mar 2, 2026Updated 7 months ago
- ☆54May 19, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The repository has collected a batch of noteworthy MLSys bloggers (Algorithms/Systems)☆345Jan 5, 2025Updated last year
- Flash Attention from Scratch on CUDA Ampere☆201Sep 1, 2025Updated last year
- NCCL communication API layer, and transport layer created from first principles.☆16Aug 20, 2025Updated last year
- MPI Code Generation through Domain-Specific Language Models☆16Nov 19, 2024Updated last year
- 使用 cutlass 实现 flash-attention 精简版,具有教学意义☆59Aug 12, 2024Updated 2 years ago
- Cute layout visualization☆45Jan 18, 2026Updated 8 months ago
- CUDA 算子手撕与面试指南☆1,132Aug 23, 2025Updated last year
- patches for huggingface transformers to save memory☆37May 9, 2026Updated 5 months ago
- ☆16Nov 28, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [DAC'25] Official implement of "HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference"☆121Dec 15, 2025Updated 9 months ago
- Implement Flash Attention using Cute.☆112Dec 17, 2024Updated last year
- ☆26Apr 2, 2026Updated 6 months ago
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated last year
- MIRAGE (USENIX Security 2021)☆15Nov 8, 2023Updated 2 years ago
- Learning material for CMU10-714: Deep Learning System☆326May 12, 2024Updated 2 years ago
- 基于 CUDA Driver API 的 cuda 运行时环境☆16Jul 30, 2025Updated last year
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 4 months ago
- MoFlo — an opinionated, local-first AI agent orchestration toolkit for Claude Code: semantic memory, learned routing, gates, and spells. …☆18Oct 1, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆39Oct 16, 2025Updated 11 months ago
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)☆31Jan 22, 2026Updated 8 months ago
- MICRO 2024 Evaluation Artifact for FuseMax☆19Aug 26, 2024Updated 2 years ago
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆25Apr 10, 2026Updated 6 months ago
- 算子库(Rust)☆15Jul 24, 2025Updated last year
- use TJCTM24024-SPI module with xpt2046 and TFT☆10Jan 15, 2018Updated 8 years ago
- fake CUTLASS to get peformance☆25Apr 28, 2026Updated 5 months ago
- Since the emergence of chatGPT in 2022, the acceleration of Large Language Model has become increasingly important. Here is a list of pap…☆284Mar 6, 2025Updated last year
- ☆12Apr 25, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12Feb 23, 2022Updated 4 years ago
- 第十五届全国大学生智能汽车竞赛-双车组三轮图像处理☆11Dec 10, 2022Updated 3 years ago
- Repo for PyChart 1.39, refs http://download.gna.org/pychart/☆10Sep 29, 2014Updated 12 years ago
- 本项目为2023年全国大学生嵌入式芯片与系统设计竞赛——FPGA创新设计竞赛(高云赛道)项目,题目基于高云FPGA的多路网络视频监控编码系统。☆71Dec 11, 2023Updated 2 years ago
- 2022年哈尔滨工业大学(深圳)《操作系统》课程实验的xv6实验部分 | xv6 labs of the course "Operating System", HITSZ, 2022.☆12Jan 2, 2023Updated 3 years ago
- ☆10Mar 14, 2021Updated 5 years ago
- ☆16Mar 26, 2025Updated last year