⚔️ [ICLR 2026] Official code of "Search Arena: Analyzing Search-Augmented LLMs".
☆58Feb 23, 2026Updated 5 months ago
Alternatives and similar repositories for search-arena
Users that are interested in search-arena are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Every Call is Precious: Global Optimization of Black-Box Functions with Unknown Lipschitz Constants☆16Apr 23, 2026Updated 3 months ago
- ☆24Aug 20, 2025Updated 11 months ago
- Analysis code for Neurips 2025 paper "SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks"☆56Aug 6, 2025Updated 11 months ago
- Ordinal Common-sense Inference☆27May 15, 2018Updated 8 years ago
- Intro to using DSPy with Kuzu to enrich the data within the Nobel Laureate mentorship network☆16Sep 16, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆42Apr 9, 2025Updated last year
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval And Synthesis For SLMs☆62Oct 7, 2025Updated 9 months ago
- ☆54Jul 10, 2026Updated 2 weeks ago
- TSQA: Tabular Scenario Based Question Answering (AAAI 2021)☆18Dec 17, 2020Updated 5 years ago
- [ACL'26 Findings] Steering LLM Thinking with Budget Guidance☆33Feb 19, 2026Updated 5 months ago
- This repository contains the code for implementation of RAG approach with company policies data, evaluation of RAG solution and smart chu…☆16Sep 18, 2025Updated 10 months ago
- ☆10Sep 25, 2019Updated 6 years ago
- This repo contains the code for "MEGA-Bench Scaling Multimodal Evaluation to over 500 Real-World Tasks" [ICLR 2025]☆81Jul 1, 2025Updated last year
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)☆319May 28, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A benchmark to evaluate search-augmented LLMs☆17Aug 28, 2025Updated 11 months ago
- ☆13Aug 26, 2024Updated last year
- Source code for COLING 2022 paper "Automatic Label Sequence Generation for Prompting Sequence-to-sequence Models"☆24Sep 21, 2022Updated 3 years ago
- [NeurIPS'25] Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning☆16Dec 12, 2025Updated 7 months ago
- ☆27Jul 23, 2025Updated last year
- ☆20Dec 14, 2024Updated last year
- ☆13Dec 9, 2024Updated last year
- LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild☆16Oct 31, 2024Updated last year
- ☆18Mar 25, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Awesome-RL-Reasoning☆17Updated this week
- ☆16Aug 5, 2025Updated 11 months ago
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆16Feb 23, 2026Updated 5 months ago
- A spoken version of the textual story cloze benchmark☆22Aug 6, 2023Updated 2 years ago
- Official repository for paper "ReasonIR Training Retrievers for Reasoning Tasks".☆230Jul 2, 2026Updated 3 weeks ago
- Exploring aspects of similarity between spoken personal narratives by disentangling them into narrative clause types -- Supplementary inf…☆12Jul 14, 2020Updated 6 years ago
- benchmarks for evaluating MT models☆11Jun 26, 2024Updated 2 years ago
- Official implementation of "Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data" (ICLR 2024)☆36Oct 16, 2024Updated last year
- a set of tools for computer vision processing☆18Jul 9, 2016Updated 10 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Official repository for the EMNLP 2019 paper, "Learning Invariant Representations of Social Media Users."☆12Aug 27, 2021Updated 4 years ago
- This repository contains the code for the paper "Botometer 101: Social bot practicum for computational social scientists."☆11Oct 6, 2022Updated 3 years ago
- [ICLR 2025] Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization☆25Oct 5, 2025Updated 9 months ago
- ☆23Sep 7, 2025Updated 10 months ago
- ☆16Aug 18, 2025Updated 11 months ago
- Wikipedia based dataset to train relationship classifiers and fact extraction models☆25May 25, 2021Updated 5 years ago
- ☆15Apr 25, 2026Updated 3 months ago