Mojo Opset is a collection of different high-performance kernel implementations for LLM and multimodal.
☆55Sep 16, 2026Updated this week
Alternatives and similar repositories for mojo_opset
Users that are interested in mojo_opset are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A torch compile backend for multi-targets☆51May 27, 2026Updated 3 months ago
- ☆23Jun 29, 2026Updated 2 months ago
- 面向多平台编译优化的深度学习中间表示☆10Oct 28, 2024Updated last year
- A comprehensive knowledge base for Huawei Ascend NPU development, structured as distributed Agent Skills. https://ascend-ai-coding.github…☆170Updated this week
- [HPCA 2026] AI Accelerator Benchmark focuses on evaluating AI Accelerators from a practical production perspective, including the ease of…☆385Apr 22, 2026Updated 4 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Learning and Debugging for FSDP/FSDP2 Training☆18Feb 7, 2026Updated 7 months ago
- Ascend operator generation☆38Jun 17, 2026Updated 3 months ago
- Provide performance insight capabilities for RL frameworks.☆83Updated this week
- Guide to build and use Tensorflow XLA/AOT on Windows☆13Dec 26, 2018Updated 7 years ago
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- My tests and experiments with some popular dl frameworks.☆17Sep 11, 2025Updated last year
- A model compilation solution for various hardware☆475Aug 20, 2025Updated last year
- 模型量化工程 Base pretrained models and datasets in pytorch (MNIST, SVHN, CIFAR10, CIFAR100, STL10, AlexNet, VGG16, VGG19, ResNet, Inception,…☆12Aug 3, 2018Updated 8 years ago
- ☆59Mar 15, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 上海交通大学Xflops超算队2024招新第一轮考核试题☆16Oct 15, 2024Updated last year
- A LogGOPS (LogP, LogGP, LogGPS) Simulator and Simulation Framework☆16Aug 20, 2024Updated 2 years ago
- ☆14Feb 7, 2020Updated 6 years ago
- The public blockchain vulnerability dataset released in our FSE'22 paper☆10Aug 22, 2022Updated 4 years ago
- SGLang is a high-performance serving framework for large language models and multimodal models.☆22Updated this week
- (Pytorch ver) Code for "Fully Neural Network based Model for General Temporal Point Process"☆21Sep 15, 2020Updated 6 years ago
- Triton adapter for Ascend. Mirror of https://gitcode.com/ascend/triton-ascend☆128May 18, 2026Updated 4 months ago
- ☆26Jul 24, 2020Updated 6 years ago
- ☆23Mar 28, 2022Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- compare the theory attention gradient with PyTorch attention gradient☆16Apr 1, 2024Updated 2 years ago
- Evaluating Program Reasoning of LLMs via Formal Specification Inference (ACL 2025)☆18Sep 21, 2025Updated 11 months ago
- gups mirror☆12Oct 25, 2015Updated 10 years ago
- [NeurIPS 2023] Token-Scaled Logit Distillation for Ternary Weight Generative Language Models☆18Dec 6, 2023Updated 2 years ago
- Bayesian optimization based on Gaussian processes☆13Dec 2, 2022Updated 3 years ago
- A record of coursework in AI Computing Systems, mainly focusing on high performance computing development for MLU.☆14Jul 14, 2022Updated 4 years ago
- Code for paper "FuSeConv Fully Separable Convolutions for Fast Inference on Systolic Arrays" published at DATE 2021☆18Aug 23, 2021Updated 5 years ago
- [TMLR] Official PyTorch implementation of paper "Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precisio…☆49Sep 27, 2024Updated last year
- A "standard library" of Triton kernels.☆26Oct 2, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A Triton JIT runtime and ffi provider in C++☆41Sep 9, 2026Updated last week
- A simple and experimental c/c++ package manager☆12Jul 10, 2026Updated 2 months ago
- DLBlas: clean and efficient kernels☆47Updated this week
- (AAAI 2026) OSVBench, a new benchmark for evaluating Large Language Models (LLMs) in generating complete specification code pertaining to…☆16May 13, 2025Updated last year
- ☆28Oct 21, 2020Updated 5 years ago
- ☆14May 28, 2023Updated 3 years ago
- Public skills collected from well-known open-source projects focused on LLM infrastructure, GPU kernels, compiler/operator development☆31May 7, 2026Updated 4 months ago