Code for M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
☆23Jul 27, 2024Updated 2 years ago
Alternatives and similar repositories for M4LE
Users that are interested in M4LE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] Sparser Block-Sparse Attention via Token Permutation☆33May 22, 2026Updated 3 months ago
- Use the tokenizer in parallel to achieve superior acceleration☆20Mar 21, 2024Updated 2 years ago
- Implementation of "Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation"☆21Jul 31, 2023Updated 3 years ago
- [ACL 24 Findings] Implementation of Resonance RoPE and the PosGen synthetic dataset.☆24Mar 5, 2024Updated 2 years ago
- ☆14Nov 20, 2022Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆11May 10, 2018Updated 8 years ago
- ☆22May 13, 2019Updated 7 years ago
- [KDD24-ADS] R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models☆11Apr 9, 2024Updated 2 years ago
- Speaker Role Contextual Model for Dialogues☆15Sep 30, 2017Updated 8 years ago
- Notes on Deep Reinforcement Learning for Natural Language Processing papers☆30Jul 17, 2017Updated 9 years ago
- A PyTorch re-implementation of the persona-based neural conversation model proposed by Jiwei Li, Michel Galley, Chris Brockett, Georgios …☆26Apr 30, 2020Updated 6 years ago
- ☆14Nov 2, 2025Updated 10 months ago
- ☆36Mar 25, 2024Updated 2 years ago
- Use strategy in stock transaction for high revenue.☆10Dec 24, 2015Updated 10 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration☆15Jun 4, 2024Updated 2 years ago
- Code and data for "MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models"☆57Nov 18, 2025Updated 9 months ago
- ☆38Apr 2, 2026Updated 5 months ago
- [ICLR 2025 Spotlight] Weak-to-strong preference optimization: stealing reward from weak aligned model☆18Feb 24, 2025Updated last year
- This is for EMNLP 2024 Paper: AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction☆16Nov 4, 2024Updated last year
- ACL 2024 | LooGLE: Long Context Evaluation for Long-Context Language Models☆200Aug 6, 2026Updated last month
- ☆19Jul 5, 2024Updated 2 years ago
- [ACL 2024] "Understanding and Patching Compositional Reasoning in LLMs"☆14Aug 8, 2026Updated 3 weeks ago
- We aim to provide the best references to search, select, and synthesize high-quality and large-quantity data for post-training your LLMs.☆66Oct 3, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆10Apr 16, 2024Updated 2 years ago
- ☆15Nov 12, 2025Updated 9 months ago
- Implementation of ICLR 2018 paper "Loss-aware Weight Quantization of Deep Networks"☆27Oct 24, 2019Updated 6 years ago
- This repository contains the ToolSelect dataset which was used to fine-tune Llama-2 70B for tool selection.☆23Mar 11, 2024Updated 2 years ago
- Paradigm shift in natural language processing☆42May 29, 2022Updated 4 years ago
- A survey of long-context language models covering architecture, infrastructure, training, and evaluation☆64Mar 31, 2025Updated last year
- ☆14Sep 17, 2020Updated 5 years ago
- Collection of papers, benchmarks and newest trends in the domain of End-to-end ToDs☆14Nov 18, 2023Updated 2 years ago
- 随机扒取古诗文词语作为git的commit msg☆11Jan 16, 2017Updated 9 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [EMNLP'22] Title2Event: Benchmarking Open Event Extraction with a Large-scale Chinese Title Dataset☆20Apr 4, 2023Updated 3 years ago
- [NeurIPS 2025] Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models☆17Nov 2, 2025Updated 10 months ago
- Official code repository for the EMNLP 2021 paper☆26Jan 30, 2022Updated 4 years ago
- Task-oriented Dialog Policy Learning with Multi-Agent Reinforcement Learning☆53Jun 23, 2020Updated 6 years ago
- Favorite AI papers☆16Jul 28, 2017Updated 9 years ago
- Data and code for ACL 2026 Paper "Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems…☆20Apr 30, 2026Updated 4 months ago
- Fine-Tuning Pre-trained Transformers into Decaying Fast Weights☆20Oct 9, 2022Updated 3 years ago