Membenchmark repository
☆60Nov 27, 2025Updated 9 months ago
Alternatives and similar repositories for Membench
Users that are interested in Membench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Comprehensive Library for Memory of LLM-based Agents.☆114May 13, 2025Updated last year
- The official repository for "MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants".☆18Oct 10, 2024Updated last year
- Code for MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems☆92Sep 15, 2026Updated last week
- [NeurIPS 2025] CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension☆22Oct 8, 2025Updated 11 months ago
- Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions☆458Aug 20, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A comprehensive evaluation framework for Backboard's memory system using the LoCoMo (Long Conversation Memory) benchmark.☆18Nov 3, 2025Updated 10 months ago
- Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)☆1,102May 11, 2026Updated 4 months ago
- SECOM: On Memory Construction and Retrieval for Personalized Conversational Agents, ICLR 2025☆63Mar 1, 2025Updated last year
- ☆1,182Aug 13, 2024Updated 2 years ago
- Official repository of DialSim☆35Oct 31, 2025Updated 10 months ago
- The official implementation of the paper "Self-Updatable Large Language Models by Integrating Context into Model Parameters"☆16May 18, 2025Updated last year
- A Comprehensive Benchmarking Framework for Long-Term Conversational Memory Layers☆49Sep 3, 2026Updated 2 weeks ago
- 基于 PyDracula 的 Qt 客户端框架☆11May 10, 2022Updated 4 years ago
- ☆63Jun 1, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [COLM 2025] CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions☆18Dec 16, 2025Updated 9 months ago
- Mem-T: Densifying Rewards for Long-Horizon Memory Agents☆43Mar 22, 2026Updated 6 months ago
- Source code and demo for memory bank and SiliconFriend☆451May 24, 2023Updated 3 years ago
- The official implementation of the paper "Mem-α: Learning Memory Construction via Reinforcement Learning"☆228Dec 25, 2025Updated 8 months ago
- [ICLR 2026] LightMem: Lightweight and Efficient Memory-Augmented Generation☆1,174Sep 5, 2026Updated 2 weeks ago
- ☆336Jan 3, 2026Updated 8 months ago
- Benchmarking LLMs in Real-World Memory-Driven Interaction☆51Apr 7, 2026Updated 5 months ago
- The code for NeurIPS 2025 paper "A-Mem: Agentic Memory for LLM Agents"☆975Mar 5, 2026Updated 6 months ago
- ☆34Oct 13, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆25Sep 15, 2026Updated last week
- The implement of FedCyBGD☆12Jul 19, 2024Updated 2 years ago
- A Multi-domain Benchmark for Personalized Search Evaluation☆12Sep 7, 2023Updated 3 years ago
- Benchmarking Long-Term Memory for AI Clones☆30Apr 7, 2026Updated 5 months ago
- ☆243Dec 20, 2024Updated last year
- ☆38Feb 13, 2026Updated 7 months ago
- WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World Scenario (COLING 2025)☆13Jan 5, 2025Updated last year
- [ICML 2025] ReflectionBench: Evaluating Epistemic Agency in Large Language Models☆23Jun 24, 2025Updated last year
- [ICML 2025] Official repository for paper "OR-Bench: An Over-Refusal Benchmark for Large Language Models"☆30Mar 4, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆14Jan 3, 2025Updated last year
- A minimalist MVP demonstrating a simple yet profound insight: aligning AI memory with human episodic memory granularity. Shows how this s…☆209Apr 16, 2026Updated 5 months ago
- Code repo for "LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners"☆99May 30, 2025Updated last year
- [NeurIPS 2025 D&B (Spotlight🌟)] TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenario☆33Oct 5, 2025Updated 11 months ago
- ☆507Jul 28, 2025Updated last year
- Evaluate your agent memory on real-world dialogues, not LLM-simulated dialogues.☆51Jul 3, 2025Updated last year
- [EMNLP 2022] RLET: A Reinforcement Learning Based Approach for Explainable QA with Entailment Trees☆11Jul 15, 2023Updated 3 years ago