A benchmark to evaluate search-augmented LLMs
☆17Aug 28, 2025Updated 10 months ago
Alternatives and similar repositories for research-eval
Users that are interested in research-eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Finding Generalizable Evidence by Learning to Convince Q&A Models☆25Jan 5, 2023Updated 3 years ago
- This repository contains papers for a comprehensive survey on accelerated generation techniques in Large Language Models (LLMs).☆11May 24, 2024Updated 2 years ago
- Code for "On the Expressiveness of Approximate Inference in Bayesian Neural Networks"☆13Aug 16, 2021Updated 4 years ago
- FaVIQ: Fact Verification from Information-seeking Questions☆43Nov 23, 2022Updated 3 years ago
- A framework for few-shot evaluation of autoregressive language models.☆13Jul 14, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code repo for EMNLP 2019 WIQA dataset paper☆13Jun 12, 2023Updated 3 years ago
- ☆15Jun 19, 2026Updated last month
- AgentIR is a retriever specialized for Deep Research agents.☆62Apr 16, 2026Updated 3 months ago
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification☆21Oct 8, 2025Updated 9 months ago
- ReConsider is a re-ranking model that re-ranks the top-K (passage, answer-span) predictions of an Open-Domain QA Model like DPR (Karpukhi…☆50Apr 26, 2021Updated 5 years ago
- Progressively Pretrained Dense Corpus Index for Open-Domain QA and Information Retrieval☆43Jun 12, 2023Updated 3 years ago
- The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"☆120Sep 29, 2025Updated 9 months ago
- ☆13Mar 27, 2020Updated 6 years ago
- ☆12Sep 2, 2021Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official eval scripts for JobBench☆28Updated this week
- Simple (fast) transformer inference in PyTorch with torch.compile + lit-llama code☆10Aug 29, 2023Updated 2 years ago
- Entity-Based Knowledge Conflicts in Question Answering. Code repo for EMNLP2021 paper: https://aclanthology.org/2021.emnlp-main.565/☆77Aug 29, 2022Updated 3 years ago
- you.com's framework for evaluating deep research systems.☆75May 15, 2025Updated last year
- A simple checkers game created using C++ and the SDL2's frameworks.☆13Nov 2, 2016Updated 9 years ago
- My PyTorch playground for NLP☆13Sep 20, 2018Updated 7 years ago
- ☆27Aug 13, 2025Updated 11 months ago
- Meta-Learning for End-to-End ASR☆10Aug 8, 2020Updated 5 years ago
- Reference implementation of algorithms for reinforcement learning and Markov decision processes.☆12Jan 28, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL2025 Findings] Benchmarking Multihop Multimodal Internet Agents☆54Feb 27, 2025Updated last year
- Description and applications of OpenAI's paper about DALL-E (2021) and implementation of other (CLIP-guided) zero-shot text-to-image gene…☆33Aug 11, 2022Updated 3 years ago
- ☆48Jun 8, 2020Updated 6 years ago
- The official implementation of the EMNLP 2023 paper "Paraphrase Types for Generation and Detection"☆12Oct 20, 2024Updated last year
- ☆38May 7, 2026Updated 2 months ago
- Large-batch Training, Neural Network Optimization☆10Nov 8, 2019Updated 6 years ago
- Companion code for FanOutQA: Multi-Hop, Multi-Document Question Answering for Large Language Models (ACL 2024)☆62Apr 3, 2026Updated 3 months ago
- Model implementation for the contextual embeddings project☆47Jun 2, 2025Updated last year
- Supporting example for "A Rust SentencePiece implementation"☆20Jun 7, 2020Updated 6 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Use contrastive learning to train a large language model (LLM) as a retriever☆12Jul 19, 2024Updated 2 years ago
- The dataset used in the CVPR 2022 paper (SimAN: Exploring Self-Supervised Representation Learning of Scene Text via Similarity-Aware Norm…☆34Jun 21, 2022Updated 4 years ago
- C^3-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking☆38Mar 1, 2026Updated 4 months ago
- ☆10Nov 15, 2020Updated 5 years ago
- An original implementation of EMNLP 2019, "A Discrete Hard EM Approach for Weakly Supervised Question Answering"☆134Jul 3, 2020Updated 6 years ago
- a deep recurrent model for exchangeable data☆34Jul 1, 2020Updated 6 years ago
- Official code for AAAI'20 paper "Merging Weak and Active Supervision for Semantic Parsing"☆11Dec 8, 2022Updated 3 years ago