A HuggingFace compatible Small Language Model trainer.
☆77Feb 2, 2025Updated last year
Alternatives and similar repositories for helibrunna
Users that are interested in helibrunna are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MLX implementation of xLSTM model by Beck et al. (2024)☆31Jun 5, 2024Updated 2 years ago
- Official repository of the xLSTM.☆2,205Sep 7, 2026Updated 2 weeks ago
- [ICLR'25] "Understanding Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing" by Peihao Wang, Ruisi Cai, Yue…☆18Mar 21, 2025Updated last year
- docker for HF wav2vec2-sprint☆13Mar 26, 2021Updated 5 years ago
- (updated 2026, We will soon release new version of dataset and a playground) An electric guitar transcription model for the real world so…☆15Jan 11, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Label shift estimation for transfer difficulty with Familiarity.☆10Feb 4, 2025Updated last year
- ☆13Jun 16, 2021Updated 5 years ago
- Official code for the NeurIPS25 paper "RAT: Bridging RNN Efficiencyand Attention Accuracy in Language Modeling" (https://arxiv.org/abs/25…☆26Aug 31, 2026Updated 3 weeks ago
- Official code repo for paper "Great Memory, Shallow Reasoning: Limits of kNN-LMs"☆24Apr 30, 2025Updated last year
- Jax implementation of x-LSTM: Extended Long Short-Term Memory by Beck et al. (2024)☆16Aug 6, 2024Updated 2 years ago
- ☆16Oct 21, 2025Updated 11 months ago
- ☆21Apr 26, 2026Updated 5 months ago
- Trully flash implementation of DeBERTa disentangled attention mechanism.☆93Sep 20, 2026Updated last week
- ☆10Oct 2, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Feb 7, 2023Updated 3 years ago
- Code and data to explore neural scaling laws of xLSTM and Transformer models.☆24Apr 8, 2026Updated 5 months ago
- ☆14May 29, 2026Updated 3 months ago
- DeltaProduct is a new linear recurrent neural network architecture that uses products of generalized Householder matrices as state-transi…☆18Oct 13, 2025Updated 11 months ago
- A benchmarking harness for coding agents.☆17Updated this week
- SDLG is an efficient method to accurately estimate aleatoric semantic uncertainty in LLMs☆28Sep 2, 2026Updated 3 weeks ago
- Implementation of Cascaded Head-colliding Attention (ACL'2021)☆11Sep 16, 2021Updated 5 years ago
- Code for the paper "A Data-Driven Methodology for Considering Feasibility and Pairwise Likelihood in Deep Learning Based Guitar Tablature…☆20Dec 14, 2022Updated 3 years ago
- Official implementation of "BERTs are Generative In-Context Learners"☆32Mar 14, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for the paper Don't Pay Attention☆60Sep 25, 2025Updated last year
- A Python Toolbox for Sonifying Music Annotations and Feature Representations☆26Mar 24, 2025Updated last year
- Official implementation of the paper - GD-Retriever: Controllable generative text-music retrieval with diffusion models (Accepted at ISMI…☆19Sep 25, 2025Updated last year
- ☆12Feb 9, 2021Updated 5 years ago
- Tries to UI development. Clone of https://www.perplexity.ai/☆11Sep 30, 2023Updated 2 years ago
- Implementation and experiments for "Partially Supervised NER via Expected Entity Ratio" in TACL 2022☆14Nov 7, 2022Updated 3 years ago
- Exploring an idea where one forgets about efficiency and carries out attention across each edge of the nodes (tokens)☆56Mar 25, 2025Updated last year
- A CLI for generating synthetic data☆43May 14, 2025Updated last year
- Researchers who published code, models (in some cases), and demo apps (in few cases) along with their SOTA paper☆12Oct 19, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tiled Flash Linear Attention library for fast and efficient mLSTM Kernels.☆94Sep 14, 2026Updated last week
- Linear Attention for Efficient Bidirectional Sequence Modeling☆18May 13, 2025Updated last year
- [NeurIPS 2022]MorphTE: Injecting Morphology in Tensorized Embeddings☆17Oct 29, 2022Updated 3 years ago
- Code for "Theoretical Foundations of Deep Selective State-Space Models" (NeurIPS 2024)☆17Jan 7, 2025Updated last year
- ☆16Mar 22, 2023Updated 3 years ago
- Neve Programming Language: Fast and Simple☆16Updated this week
- Cuda implemenation of flash-kmeans, 2x faster☆23May 8, 2026Updated 4 months ago