A benchmark to evaluate search-augmented LLMs
☆17Aug 28, 2025Updated 11 months ago
Alternatives and similar repositories for research-eval
Users that are interested in research-eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Python SDK for the Reka AI API☆18Sep 30, 2024Updated last year
- Finding Generalizable Evidence by Learning to Convince Q&A Models☆25Jan 5, 2023Updated 3 years ago
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated 11 months ago
- This repository contains papers for a comprehensive survey on accelerated generation techniques in Large Language Models (LLMs).☆11May 24, 2024Updated 2 years ago
- Official repository of the R2-D2's pipeline☆21Nov 16, 2021Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This repository helps you evaluate your models on the FreshStack benchmark!☆34Dec 9, 2025Updated 8 months ago
- An official repository for MIA 2022 (NAACL 2022 Workshop) Shared Task on Cross-lingual Open-Retrieval Question Answering.☆31Jun 26, 2022Updated 4 years ago
- Code for "On the Expressiveness of Approximate Inference in Bayesian Neural Networks"☆13Aug 16, 2021Updated 4 years ago
- A framework for few-shot evaluation of autoregressive language models.☆13Jul 14, 2025Updated last year
- Code repo for EMNLP 2019 WIQA dataset paper☆13Jun 12, 2023Updated 3 years ago
- ☆15Jun 19, 2026Updated last month
- A tool for extracting plain text from Wikipedia dumps☆15Sep 13, 2018Updated 7 years ago
- AgentIR is a retriever specialized for Deep Research agents.☆62Apr 16, 2026Updated 3 months ago
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification☆21Oct 8, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ReConsider is a re-ranking model that re-ranks the top-K (passage, answer-span) predictions of an Open-Domain QA Model like DPR (Karpukhi…☆50Apr 26, 2021Updated 5 years ago
- Progressively Pretrained Dense Corpus Index for Open-Domain QA and Information Retrieval☆43Jun 12, 2023Updated 3 years ago
- The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"☆120Sep 29, 2025Updated 10 months ago
- This is the open source version of HPL-MXP. The code performance has been verified on Frontier☆18Jul 9, 2025Updated last year
- This repository contains the code for the publication "Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Lang…☆10Oct 26, 2023Updated 2 years ago
- ☆13Mar 27, 2020Updated 6 years ago
- ☆12Sep 2, 2021Updated 4 years ago
- Representation Learning of Entities and Documents from Knowledge Base Descriptions☆18Oct 6, 2018Updated 7 years ago
- 🎓 STiNE CLI / library in Go☆11Apr 3, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official eval scripts for JobBench☆35Updated this week
- Entity-Based Knowledge Conflicts in Question Answering. Code repo for EMNLP2021 paper: https://aclanthology.org/2021.emnlp-main.565/☆77Aug 29, 2022Updated 3 years ago
- you.com's framework for evaluating deep research systems.☆76May 15, 2025Updated last year
- My PyTorch playground for NLP☆13Sep 20, 2018Updated 7 years ago
- Replication code for "With Little Power Comes Great Responsibility"☆39Oct 15, 2020Updated 5 years ago
- ☆126Aug 1, 2026Updated last week
- Reparameterize your PyTorch modules☆70Dec 31, 2020Updated 5 years ago
- Meta-Learning for End-to-End ASR☆10Aug 8, 2020Updated 6 years ago
- ☆23Dec 26, 2023Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Reference implementation of algorithms for reinforcement learning and Markov decision processes.☆13Jan 28, 2021Updated 5 years ago
- [ACL2025 Findings] Benchmarking Multihop Multimodal Internet Agents☆54Feb 27, 2025Updated last year
- ☆48Jun 8, 2020Updated 6 years ago
- The dataset consists of public social media url pairs and the corresponding entailment label for an external conference (ACL 2021). Each …☆14Aug 16, 2021Updated 4 years ago
- The official implementation of the EMNLP 2023 paper "Paraphrase Types for Generation and Detection"☆12Oct 20, 2024Updated last year
- German Alpaca Dataset (Cleaned + Translated)☆26Apr 6, 2023Updated 3 years ago
- Large-batch Training, Neural Network Optimization☆10Nov 8, 2019Updated 6 years ago