Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
☆13Aug 2, 2026Updated 2 weeks ago
Alternatives and similar repositories for Griffin
Users that are interested in Griffin are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of Griffin from the paper: "Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models"☆58Oct 27, 2025Updated 9 months ago
- Code for MICCAI2023 paper: TransLiver: A Hybrid Transformer Model for Multi-phase Liver Lesion Classification☆18Jan 10, 2024Updated 2 years ago
- 4-bit Shampoo for Memory-Efficient Network Training (NeurIPS 2024)☆13Feb 13, 2025Updated last year
- A Transformer-based Prediction Method for Depth of Anesthesia During Target-controlled Infusion of Propofol and Remifentanil.☆16Feb 17, 2025Updated last year
- Reproduce the artice Hoy et al. 2014 Optimization of a free water elimination two-compartment model for diffusion tensor imaging. Neuroim…☆14Apr 12, 2019Updated 7 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- This is the official implementation for IVA '19 paper "Analyzing Input and Output Representations for Speech-Driven Gesture Generation".☆10Jul 12, 2022Updated 4 years ago
- ☆15Jan 19, 2024Updated 2 years ago
- Code for ICRA 2022 paper - KinoJGM: A framework for efficient and accurate quadrotor trajectory generation and tracking in dynamic enviro…☆14Jul 18, 2024Updated 2 years ago
- ☆17May 31, 2024Updated 2 years ago
- Official Implementation of AQ-GT: a Temporally Aligned and Quantized GRU-Transformer for Co-Speech Gesture Synthesis with the extension (…☆21Apr 19, 2024Updated 2 years ago
- ☆13Mar 30, 2022Updated 4 years ago
- This is the official repository for our publication "The IVI Lab entry to the GENEA Challenge 2022 – A Tacotron2 Based Method for Co-Spee…☆13May 2, 2023Updated 3 years ago
- 《Python机器学习实践指南》代码和笔记☆12Aug 26, 2020Updated 5 years ago
- ☆14May 24, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Feb 5, 2021Updated 5 years ago
- Jax implementation of "Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models"☆16May 10, 2024Updated 2 years ago
- BrainHGT: A Hierarchical Graph Transformer for Interpretable Brain Network Analysis☆23Nov 25, 2025Updated 8 months ago
- Training GPTs to solve interaction nets☆18Aug 14, 2024Updated 2 years ago
- A publishing website of a table collecting meta-learning-related papers in the area of human language processing.☆17Aug 2, 2021Updated 5 years ago
- ☆21Jun 22, 2022Updated 4 years ago
- a tool to segment the hypothalamus and associated subunits on T1-weighted MRI scans☆27Jan 6, 2023Updated 3 years ago
- code of paper "Robust and kernelized data-enabled predictive control for nonlinear systems"☆20Sep 4, 2023Updated 2 years ago
- Disentangled Implicit Content and Rhythm Learning for Diverse Co-Speech Gestures Synthesis [ACMMM 2022]☆26Jun 26, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆49Mar 31, 2024Updated 2 years ago
- ☆12Sep 21, 2023Updated 2 years ago
- My implementation of diffusion (like) models☆11Apr 14, 2023Updated 3 years ago
- Time Series Representation Models☆13Jul 17, 2025Updated last year
- Official implementation of MMPD: Diverse Time Series Forecasting via Multi-Mode Patch Diffusion Loss (ICLR 2026).☆17Apr 3, 2026Updated 4 months ago
- BOLD-CSF coupling☆24Feb 23, 2024Updated 2 years ago
- Teaching Addition to Small Transformers☆19Nov 28, 2023Updated 2 years ago
- ☆16Jan 3, 2025Updated last year
- This is an official pytorch implementation for paper "Hi-Patch: Hierarchical Patch GNN for Irregular Multivariate Time Series" (ICML-25).☆19May 2, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Dec 17, 2021Updated 4 years ago
- Fast Kolmogorov-Arnold Network in JAX, initial experiments☆17May 20, 2024Updated 2 years ago
- ☆24Apr 27, 2023Updated 3 years ago
- Implementation and pretrained model for the SingLEM paper.☆15Aug 1, 2026Updated 2 weeks ago
- TabDDPM is the state of the art synthetic data generation tool using diffusion models. Here I wrap the diffusion model in an easier plug …☆20Jun 26, 2025Updated last year
- An Open-access Dataset for Liver Lesion Diagnosis on Multi-phase MRI☆40Apr 7, 2025Updated last year
- ☆33Aug 14, 2023Updated 3 years ago