Code to train and evaluate Neural Attention Memory Models to obtain universally-applicable memory systems for transformers.
β360Oct 22, 2024Updated last year
Alternatives and similar repositories for evo-memory
Users that are interested in evo-memory are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for Discovering Preference Optimization Algorithms with and for Large Language Modelsβ196Jun 13, 2024Updated 2 years ago
- A Self-adaptation Frameworkπ that adapts LLMs for unseen tasks in real-time!β1,221Jan 30, 2025Updated last year
- Automating the Search for Artificial Life with Foundation Models!β476Oct 23, 2025Updated 9 months ago
- Official implementation of "TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models"β123Oct 6, 2025Updated 9 months ago
- Official repository of Evolutionary Optimization of Model Merging Recipesβ1,437Nov 29, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Experiments Notebook of "Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism"β16Apr 30, 2025Updated last year
- β16Jul 16, 2024Updated 2 years ago
- Memory layers use a trainable key-value lookup mechanism to add extra parameters to a model without increasing FLOPs. Conceptually, sparsβ¦β379Dec 12, 2024Updated last year
- Learning to route instances for Human vs AI Feedback (ACL Main '25)β29Jul 23, 2025Updated last year
- An open source replication of the stawberry method that leverages Monte Carlo Search with PPO and or DPOβ30Updated this week
- Example of using Epochraft to train HuggingFace transformers models with PyTorch FSDPβ11Jan 29, 2024Updated 2 years ago
- [ICLR 2025 & COLM 2025] Official PyTorch implementation of the Forgetting Transformer and Adaptive Computation Pruningβ150Feb 25, 2026Updated 4 months ago
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddingβ219Jan 12, 2026Updated 6 months ago
- Code for Fast-weight Product Key Memory (FwPKM)β19Mar 18, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β48Jul 3, 2026Updated 3 weeks ago
- DeciMamba: Exploring the Length Extrapolation Potential of Mamba (ICLR 2025)β32Apr 9, 2025Updated last year
- An AI character interaction system with emotional modeling and advanced memory managementβ17Oct 26, 2024Updated last year
- Tools for merging pretrained large language models.β7,260Jun 17, 2026Updated last month
- Code for EMNLP'24 paper - On Diversified Preferences of Large Language Model Alignmentβ16Aug 6, 2024Updated last year
- The code repository of the paper: Competition and Attraction Improve Model Fusionβ170Aug 25, 2025Updated 11 months ago
- Official repository of paper "RNNs Are Not Transformers (Yet): The Key Bottleneck on In-context Retrieval"β27Apr 17, 2024Updated 2 years ago
- β284Jun 6, 2025Updated last year
- OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code (ICLR 2025).β81Dec 26, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for the EMNLP24 paper "A simple and effective L2 norm based method for KV Cache compression."β19Dec 13, 2024Updated last year
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention (NeurIPS'25 Spotlight)β26Feb 22, 2026Updated 5 months ago
- OMNI: Open-endedness via Models of human Notions of Interestingnessβ66Jan 28, 2025Updated last year
- [Technical Report] Official PyTorch implementation code for realizing the technical part of Phantom of Latent representing equipped with β¦β63Oct 9, 2024Updated last year
- Evaluating majors LLMs on the Abstraction and Reasoning Corpusβ17Nov 9, 2023Updated 2 years ago
- DeMo: Decoupled Momentum Optimizationβ202Dec 2, 2024Updated last year
- β93Aug 18, 2024Updated last year
- LongRoPE is a novel method that can extends the context window of pre-trained LLMs to an impressive 2048k tokens.β290Oct 28, 2025Updated 8 months ago
- Repo to reproduce the First-Explore paper resultsβ39May 6, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- train with kittens!β67Oct 25, 2024Updated last year
- Repo for "LoLCATs: On Low-Rank Linearizing of Large Language Models"β260Jan 31, 2025Updated last year
- β14Mar 2, 2025Updated last year
- The code for "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by VIdeo SpatioTemporal Augmentation" [CVPR2025]β20Feb 27, 2025Updated last year
- Entropy Based Sampling and Parallel CoT Decodingβ3,433Nov 13, 2024Updated last year
- OpenCoconut implements a latent reasoning paradigm where we generate thoughts before decoding.β173Jan 16, 2025Updated last year
- β33Jan 7, 2025Updated last year