Triton implementation of GPT/LLAMA
☆22Aug 28, 2024Updated last year
Alternatives and similar repositories for gpt-triton
Users that are interested in gpt-triton are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Iterate fast on your RAG pipelines☆24Jun 21, 2025Updated last year
- Fast SGEMM emulation on Tensor Cores☆17Feb 16, 2025Updated last year
- ☆14Jun 24, 2024Updated 2 years ago
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- Because it's there.☆16Sep 22, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Mamba support for transformer lens☆20Sep 17, 2024Updated last year
- ☆11Feb 22, 2025Updated last year
- Implementation of Direct Preference Optimization☆17Jul 17, 2023Updated 3 years ago
- Built quantitative models to measure value at risk (VaR) and Expected Shortfall (ES).☆13Aug 30, 2018Updated 7 years ago
- ☆17Jan 1, 2025Updated last year
- Full End-to-End examples showing how to use First-gen Gaudi and Gaudi2 in common use cases☆13Dec 2, 2024Updated last year
- CLIP is an open source, multimodal computer vision model and it's awesome!☆17Dec 16, 2024Updated last year
- Repository for NLP project. Name to be changed when we decide on a project☆16Apr 19, 2022Updated 4 years ago
- Minimal implementation of TokenFormer for inference and learning☆13Nov 6, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Jan 30, 2023Updated 3 years ago
- ☆15Jul 13, 2026Updated last month
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆17Aug 6, 2026Updated last week
- ☆15Apr 24, 2026Updated 3 months ago
- ☆51Jun 9, 2026Updated 2 months ago
- Self Reproduction Code of Paper "Reducing Transformer Key-Value Cache Size with Cross-Layer Attention (MIT CSAIL)☆17May 24, 2024Updated 2 years ago
- Simple and scalable tools for data-driven pretraining data selection.☆30Jun 9, 2025Updated last year
- Repository for ACM India Summer School on Generative AI for Text☆13Jul 11, 2024Updated 2 years ago
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆18Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Replication package for evaluation of code generation metrics☆17Nov 24, 2025Updated 8 months ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"☆19May 30, 2025Updated last year
- Workshop materials for AI Engineer World's Fair☆19Jun 3, 2025Updated last year
- Exercises & Notes from the LinkedIn Learning courses about Behavioral, Creational & Structural Design Patterns by Bethan Palmer☆18Jun 7, 2025Updated last year
- ☆19Jan 4, 2024Updated 2 years ago
- Writing FLUX in Triton☆42Sep 22, 2024Updated last year
- ☆23Jan 10, 2025Updated last year
- ☆15Aug 26, 2023Updated 2 years ago
- Writing and Citation Assistant Tool☆39Dec 21, 2025Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Converting Mixtral-8x7B to Mixtral-[1~7]x7B☆22Mar 4, 2024Updated 2 years ago
- Winning solution of the Kaggle "Google Brain - Ventilator Pressure Prediction" competition☆10Nov 12, 2021Updated 4 years ago
- Implementation of BERT-based Language Models☆29Aug 3, 2026Updated last week
- ☆63Jun 2, 2021Updated 5 years ago
- Build a Resume Parser in Python using Spacy☆21Jul 22, 2022Updated 4 years ago
- Centerface ONNX accelerated with Deepstream 5.1☆10Jun 13, 2021Updated 5 years ago
- Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parall…☆20Updated this week