☆369Dec 1, 2025Updated 10 months ago
Alternatives and similar repositories for copilot-arena
Users that are interested in copilot-arena are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆31Apr 7, 2026Updated 6 months ago
- This library supports evaluating disparities in generated image quality, diversity, and consistency between geographic regions.☆20Jun 3, 2024Updated 2 years ago
- Codebase exploration with AI research agents☆21Feb 25, 2025Updated last year
- AI agent webscrapers☆13Feb 4, 2024Updated 2 years ago
- The project to find correlation between tweets and future stock prices☆12Feb 28, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Trim and timestamp audio, in the terminal☆14Oct 14, 2024Updated last year
- ☆24Nov 5, 2024Updated last year
- ☆31Sep 23, 2024Updated 2 years ago
- [ICLR 2025] 🚀 CodeMMLU Evaluator: A framework for evaluating LM models on CodeMMLU MCQs benchmark.☆30Apr 21, 2025Updated last year
- An open-source chat text to control actions agentic workflow framework/showcase powered by Agently AI application development framework.☆30Mar 13, 2026Updated 6 months ago
- Code for verifying deep neural feature ansatz☆23May 3, 2023Updated 3 years ago
- ☆18Jun 23, 2026Updated 3 months ago
- This project is concerned with my participating in the RuNNE competition https://github.com/dialogue-evaluation/RuNNE☆13Jun 28, 2023Updated 3 years ago
- [COLM-LLA 2026] The official implementation for paper "AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficien…☆23Aug 23, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- https://scale.com/research/mrt☆20Mar 16, 2026Updated 6 months ago
- Nexusflow function call, tool use, and agent benchmarks.☆28Dec 13, 2024Updated last year
- II-Thought-RL is our initial attempt at developing a large-scale, multi-domain Reinforcement Learning (RL) dataset☆30Apr 8, 2025Updated last year
- Code for steering and monitoring with concepts vectors in LLMs (initial draft)☆33Aug 10, 2025Updated last year
- A benchmark of programming tasks for LLMs that supports almost any programming language.☆13Jun 30, 2025Updated last year
- A model implementation of sessions for koa using postgres as the backend☆10Oct 16, 2017Updated 8 years ago
- ⚔️ OpenHands PR Arena ⚔️ is a platform for evaluating and benchmarking agentic coding assistants through paired pull request (PR) generat…☆18Dec 15, 2025Updated 9 months ago
- A self-improving swarm of local-LLM agents that mine, smelt, build, farm, and fight their way through Minecraft as a coordinated team. Bu…☆31Sep 29, 2026Updated last week
- ☆15Jan 17, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆866Mar 18, 2025Updated last year
- helpful scripts for llamafiles☆14Jul 21, 2026Updated 2 months ago
- [Obsolete] OpenAI GPT hosted Agent Framework for Windows and MacOS from 2024☆36Jul 8, 2024Updated 2 years ago
- ☆32Oct 4, 2024Updated 2 years ago
- ☆43Jan 6, 2025Updated last year
- uvx is now uvenv☆16Dec 4, 2024Updated last year
- A distributed, extensible, secure solution for evaluating machine generated code with unit tests in multiple programming languages.☆63Oct 21, 2024Updated last year
- Training and data processing code for Saiga☆56Jan 2, 2026Updated 9 months ago
- GRadient-INformed MoE☆264Sep 25, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for Husky, an open-source language agent that solves complex, multi-step reasoning tasks. Husky v1 addresses numerical, tabular and …☆349Jun 16, 2024Updated 2 years ago
- ☆22Oct 24, 2025Updated 11 months ago
- PhD/MBA-level human-annotated rubrics dataset across Physics, Chemistry, Finance and Consulting☆34Oct 30, 2025Updated 11 months ago
- Telegram bot for different language models. Supports system prompts and images☆67Aug 19, 2026Updated last month
- A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models☆77Feb 25, 2025Updated last year
- A DSPy integration for the Agent Skills specification, enabling ReAct agents to dynamically discover and use modular, sandboxed skills☆28Jan 8, 2026Updated 9 months ago
- ☆17May 10, 2024Updated 2 years ago