☆30Oct 26, 2024Updated last year
Alternatives and similar repositories for flaxattention
Users that are interested in flaxattention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆34Jun 3, 2024Updated 2 years ago
- Translation of the databricks-dolly-15k dataset to Chinese for commercial use.☆19Apr 17, 2023Updated 3 years ago
- Two implementations of ZeRO-1 optimizer sharding in JAX☆14Jun 11, 2023Updated 3 years ago
- EncryptedClipboard☆13Sep 24, 2020Updated 6 years ago
- ☆13Apr 16, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆35Sep 6, 2025Updated last year
- A FlashAttention implementation for JAX with support for efficient document mask computation and context parallelism.☆175Nov 11, 2025Updated 10 months ago
- Code for the reproduction of counting manifolds☆16Feb 26, 2026Updated 7 months ago
- Particle filtering in JAX☆18Aug 27, 2026Updated last month
- GGUF parser in Python☆28May 1, 2026Updated 5 months ago
- Microbenchmarking hyperparameter tuning for JAX functions.☆24Sep 24, 2026Updated last week
- Python library for evaluating and comparing generative models. Unified interface for computing quality metrics across images, videos, te…☆19Dec 18, 2025Updated 9 months ago
- 2021 Spring☆18Oct 12, 2024Updated last year
- Official implementation for "Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size", NeurIPS 2022, Offline RL Worksho…☆21Feb 27, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- triton ver of gqa flash attn, based on the tutorial☆12Aug 4, 2024Updated 2 years ago
- Sparse Autoencoder Training Library☆58May 1, 2025Updated last year
- Dockerfile and instructions for human pose estimation implementation using Caffe, OpenCV 3.1.0 and Python 2.7.☆12Mar 3, 2019Updated 7 years ago
- Tokamax: A GPU and TPU kernel library.☆306Updated this week
- ☆19Dec 4, 2025Updated 9 months ago
- Improving Neural Text Generation with Reinforcement Learning☆23Jan 13, 2021Updated 5 years ago
- QWOP AI using Q-learning☆12Jul 13, 2016Updated 10 years ago
- Minimal yet performant LLM examples in pure JAX☆285Aug 23, 2026Updated last month
- Accelerate, Optimize performance with streamlined training and serving options with JAX.☆374Sep 26, 2026Updated last week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official Implementation for the ICLR2023 paper "Fuzzy Alignments in Directed Acyclic Graph for Non-autoregressive Machine Translation"☆14Mar 1, 2023Updated 3 years ago
- All the tools that allow me to never ever open up Final Cut☆11Feb 16, 2025Updated last year
- a common lisp library for data analysis and manipulation☆12May 28, 2023Updated 3 years ago
- Convert StableHLO models into Apple Core ML format☆22Sep 12, 2026Updated 3 weeks ago
- A library of speech gadgets.☆16Oct 15, 2022Updated 3 years ago
- a simple tool to translate caffe model to keras model☆10Oct 26, 2015Updated 10 years ago
- Mxnet implementation of an ICLR 2018 paper: A new method of region embedding for text classification.☆10Oct 14, 2018Updated 7 years ago
- Reversal Curse Experiment☆15Sep 24, 2023Updated 3 years ago
- Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph Learning.☆16Nov 7, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- SIMD implementation of argmin and argmax☆12Jul 2, 2024Updated 2 years ago
- Using YouTube to prepare a speech recognition dataset for any language☆10Mar 30, 2021Updated 5 years ago
- Trainable H-Net Package☆34Sep 3, 2025Updated last year
- Yet Another Linter Aggregator☆15Oct 17, 2023Updated 2 years ago
- ☆15Oct 13, 2025Updated 11 months ago
- Pytorch/XLA SPMD Test code in Google TPU☆23Apr 3, 2024Updated 2 years ago
- 本项目主要对开源的MOSS SFT数据进行整理 ,转换成mnbvc多轮对话格式。MOSS-003涵盖用性、忠实性、无害性三个层面,共353w样本,MOSS-003 包含更细粒度的有用性类别标记、更广泛的无害性数据和更长对话轮数,共630w样本,☆13Dec 3, 2023Updated 2 years ago