huggingface / candle-paged-attentionLinks
☆12Updated last year
Alternatives and similar repositories for candle-paged-attention
Users that are interested in candle-paged-attention are comparing it to the libraries listed below
Sorting:
- ☆19Updated last year
- CLI utility to inspect and explore .safetensors and .gguf files☆32Updated 3 months ago
- A collection of optimisers for use with candle☆43Updated 2 months ago
- implement llava using candle☆15Updated last year
- Experimental GPU language with meta-programming☆23Updated last year
- Rust crate for some audio utilities☆25Updated 7 months ago
- Experimental compiler for deep learning models☆67Updated last month
- Read and write tensorboard data using Rust☆23Updated last year
- ☆24Updated 6 months ago
- ☆134Updated last year
- Sample Python extension using Rust/PyO3/tch to interact with PyTorch☆37Updated last year
- ☆27Updated 2 years ago
- Automatically derive Python dunder methods for your Rust code☆20Updated 6 months ago
- Simple (fast) transformer inference in PyTorch with torch.compile + lit-llama code☆10Updated 2 years ago
- ☆21Updated 7 months ago
- GPU based FFT written in Rust and CubeCL☆24Updated 4 months ago
- An implementation of the Llama architecture, to instruct and delight☆21Updated 4 months ago
- Your one stop CLI for ONNX model analysis.☆47Updated 2 years ago
- Low rank adaptation (LoRA) for Candle.☆163Updated 6 months ago
- ☆58Updated 2 years ago
- 8-bit floating point types for Rust☆60Updated 2 months ago
- 👷 Build compute kernels☆163Updated this week
- ☆12Updated 9 months ago
- A high-performance constrained decoding engine based on context free grammar in Rust☆55Updated 5 months ago
- Experiment of using Tangent to autodiff triton☆80Updated last year
- ☆93Updated 9 months ago
- research impl of Native Sparse Attention (2502.11089)☆62Updated 8 months ago
- JAX bindings for Flash Attention v2☆97Updated last week
- ☆70Updated last year
- Graph model execution API for Candle☆16Updated 3 months ago