General benchmarking apparatus for running multi-agent systems against benchmarks
☆46Apr 13, 2026Updated 4 months ago
Alternatives and similar repositories for neuro-san-benchmarking
Users that are interested in neuro-san-benchmarking are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆30Sep 16, 2023Updated 2 years ago
- Open-source repository for the OOPSLA'24 paper "CYCLE: Learning to Self-Refine Code Generation"☆10Mar 8, 2024Updated 2 years ago
- A Prompt Learning Framework for Source Code Summarization☆14Dec 26, 2023Updated 2 years ago
- This repo contains the source code for the paper "Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning"☆381Jun 26, 2026Updated last month
- TypeScript agents for real applications.☆26Aug 5, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Repo for "AlphaResearch: Accelerating New Algorithm Discovery with Language Models"☆59Nov 12, 2025Updated 9 months ago
- False LLM endpoints for testing☆14May 21, 2025Updated last year
- ☆10Oct 11, 2022Updated 3 years ago
- ☆11Nov 8, 2023Updated 2 years ago
- Learn DSPy's core abstractions while building a deep research agent.☆45Mar 8, 2026Updated 5 months ago
- ☆15Dec 12, 2024Updated last year
- An open-source framework for evaluating out-of-knowledge base robustness☆28Feb 11, 2026Updated 6 months ago
- Training and testing code from our CVPR 2023 paper "Are Deep Neural Networks SMARTer than Second Graders?"☆11Aug 10, 2023Updated 3 years ago
- Code for paper https://arxiv.org/abs/2501.00522☆15Apr 28, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆10Jun 11, 2023Updated 3 years ago
- Roo Code AIMS 开发模式配置,提供自定义 modes 的规则与说明。☆19Jun 15, 2026Updated 2 months ago
- Reference code for the AWS S3 section in the Dive into AWS Course.☆15Dec 8, 2022Updated 3 years ago
- Providing the answer to "How to do patching on all available SAEs on GPT-2?". It is an official repository of the implementation of the p…☆14Jan 26, 2025Updated last year
- A set of custom nodes that I've either written myself or adapted from other authors for my own convenience.☆11Sep 18, 2024Updated last year
- Code for "Demonstration-free Autonomous Reinforcement Learning via Implicit and Bidirectional Curriculum" (ICML 2023)☆10Jul 6, 2023Updated 3 years ago
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- RunwayML + Pure Data 🦜☆16Oct 24, 2019Updated 6 years ago
- ☆10Nov 16, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [EMNLP 2023] Official Code of "JointMatch: A Unified Approach for Diverse and Collaborative Pseudo-Labeling to Semi-Supervised Text Class…☆24May 13, 2024Updated 2 years ago
- Code for training on Imagenet to SOTA results using PyTorch☆13Updated this week
- ☆10Apr 26, 2023Updated 3 years ago
- A PyTorch implementation of the paper Multimodal Transformer with Multiview Visual Representation for Image Captioning☆25Sep 4, 2020Updated 5 years ago
- A compact high-signal benchmark for evaluating frontier agents☆25Aug 3, 2026Updated 2 weeks ago
- Using Bayesian inference to mine rule sets☆12Jan 9, 2020Updated 6 years ago
- The official data and code for EMNLP 2023 main conference paper: CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular D…☆13May 19, 2025Updated last year
- Groq-powered MAD: The first work to explore Multi-Agent Debate with Large Language Models :D☆12Jul 5, 2024Updated 2 years ago
- LLM Benchmark problems for SWE tasks in julia☆15Nov 24, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Jun 7, 2023Updated 3 years ago
- Workflow for AI Agent☆16Jun 12, 2025Updated last year
- This is the project for IRM methods☆12Sep 13, 2021Updated 4 years ago
- The code for our NeurIPS 2021 paper "Kernelized Heterogeneous Risk Minimization".☆13Oct 13, 2021Updated 4 years ago
- This repository contains the code and data download links to reproduce building the WDC Products Benchmark.☆15Jul 13, 2023Updated 3 years ago
- WIP - A starting location I am putting together to make GitHub's Spec Kit and Roo Code work together as seamlessly as possible. Suggestio…☆20Oct 22, 2025Updated 10 months ago
- ☆13May 23, 2024Updated 2 years ago