[ICLR 2026] Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
☆143Aug 31, 2026Updated 3 weeks ago
Alternatives and similar repositories for BEAM
Users that are interested in BEAM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆38Feb 13, 2026Updated 7 months ago
- ☆1,178Aug 13, 2024Updated 2 years ago
- Membenchmark repository☆60Nov 27, 2025Updated 9 months ago
- Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)☆1,096May 11, 2026Updated 4 months ago
- ☆30Feb 27, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Benchmarking LLMs in Real-World Memory-Driven Interaction☆51Apr 7, 2026Updated 5 months ago
- Open-source evaluation suite for memory-augmented LLM systems☆112Aug 6, 2026Updated last month
- Synthetic data generation and benchmark implementation for "Episodic Memories Generation and Evaluation Benchmark for Large Language Mode…☆69Oct 3, 2025Updated 11 months ago
- MemSyco-Bench: Benchmarking Sycophancy in Agent Memory☆20Sep 11, 2026Updated last week
- Official implementation for the paper "Video-Based Reward Modeling for Computer-Use Agents"☆17Mar 14, 2026Updated 6 months ago
- Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language☆39May 27, 2026Updated 3 months ago
- [ESWC '24] This repo is official implementation for the paper "Towards Harnessing Large Language Models as Autonomous Agents for Semantic…☆10May 25, 2024Updated 2 years ago
- A Unified Memory Benchmark Suite for Memory-Augmented Agents☆148Jul 5, 2026Updated 2 months ago
- MULTITQ is a large-scale dataset featuring ample relevant facts and multiple temporal granularities.☆27Apr 28, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [IPDPS 2024] Adaptive neighbor sampling for temporal GNN☆16Feb 17, 2025Updated last year
- Official repository for LongMemEval-V2☆162Aug 9, 2026Updated last month
- An RL Recipe for Building Agentic LLMs via Self-Imitation on Long-Horizon Agentic Tasks☆39Jan 30, 2026Updated 7 months ago
- Hierarchical RAG architecture scaling to 693K chunks on consumer hardware (4GB VRAM). Features 3-address routing, hybrid vector+graph fus…☆39Feb 11, 2026Updated 7 months ago
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- Implementation of Influence Function approximations for differently sized ML models, using PyTorch☆18Sep 15, 2023Updated 3 years ago
- An AI agent memory framework that converts an agent’s own interaction traces—both successes and failures—into reusable, high-level reason…☆64Feb 9, 2026Updated 7 months ago
- Collection of Elixir language skills for Cursor, Claude and Codex☆34Jun 24, 2026Updated 2 months ago
- Biomedical concept relatedness benchmark sampled from electronic health records☆11Jul 14, 2022Updated 4 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Official reference implementation of our paper "Long Range Propagation on Continuous-Time Dynamic Graphs" accepted at ICML24 and "Effecti…☆15Sep 11, 2025Updated last year
- Learned Metric Index (LMI) is a machine learning based data structure for fast look-up of approximate nearest neighbors in complex data.☆11Dec 14, 2025Updated 9 months ago
- The paper list of "Memory in the Age of AI Agents: A Survey"☆2,392Mar 4, 2026Updated 6 months ago
- mcp wrapper for openai built-in tools☆12Mar 13, 2025Updated last year
- ☆13Sep 9, 2026Updated last week
- The code of “Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning”☆17Feb 26, 2024Updated 2 years ago
- Deterministic graph runtime for agent workflows in Swift. Same input, same output, every time — golden-testable, checkpoint-resumable, an…☆25May 17, 2026Updated 4 months ago
- Benchmarking Long-Term Memory for AI Clones☆30Apr 7, 2026Updated 5 months ago
- Paper list for "From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms".☆54Apr 13, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official repository for K-EXAONE built by LG AI Research☆87May 15, 2026Updated 4 months ago
- Prototyp MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism☆34Apr 4, 2025Updated last year
- Official repository for paper Auto-scaling Continuous Memory for GUI Agent☆31Feb 2, 2026Updated 7 months ago
- HaluMem is the first operation level hallucination evaluation benchmark tailored to agent memory systems.☆162Sep 3, 2026Updated 2 weeks ago
- MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning☆69Jun 14, 2026Updated 3 months ago
- ☆52May 4, 2026Updated 4 months ago
- ☆13Oct 9, 2025Updated 11 months ago