Research artifacts from Recursive's automated AI research system
☆236Jun 11, 2026Updated 3 months ago
Alternatives and similar repositories for first-steps-toward-automated-ai-research
Users that are interested in first-steps-toward-automated-ai-research are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An approach to utomatically generating browser environment with verifiable tasks☆69Mar 24, 2026Updated 5 months ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆52May 30, 2026Updated 3 months ago
- Official implementation of Paper "System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving"☆30Apr 17, 2026Updated 4 months ago
- ☆636May 24, 2026Updated 3 months ago
- AI-Driven Scientific, Algorithmic, and Systems Discovery☆664Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- autonomous nanogpt optimizer speedrun☆110May 14, 2026Updated 3 months ago
- Code repository for the ICML 2026 Oral paper "Characterizing, Evaluating, and Optimizing Complex Reasoning".☆20Jun 21, 2026Updated 2 months ago
- Continual Learning Bench☆221Jul 19, 2026Updated last month
- ☆78Apr 30, 2026Updated 4 months ago
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆554Updated this week
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆145May 6, 2026Updated 4 months ago
- https://scale.com/research/mrt☆20Mar 16, 2026Updated 5 months ago
- Recursive Coding Agents — AI Engineer World's Fair 2026 talk + website (SvelteKit, Cloudflare Workers)☆22Jun 30, 2026Updated 2 months ago
- ☆23Jun 12, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ALMA (Automated meta-Learning of Memory designs for Agentic systems) is a framework that meta-learns memory designs to replace human-engi…☆294Apr 8, 2026Updated 5 months ago
- [COLM 2026] Official implementation for "MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo…☆21Updated this week
- A benchmark of real-world DL kernel problems☆295Jul 15, 2026Updated last month
- [MLSys 2026] AccelOpt: Self-improving Agents for AI Accelerator Kernel Optimization☆73Jul 28, 2026Updated last month
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆2,810Updated this week
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆40May 31, 2026Updated 3 months ago
- Source code for paper "On the Pareto Front of Multilingual Neural Machine Translation" @ NeurIPS 2023☆17Sep 27, 2023Updated 2 years ago
- mKernel: fast multi-node, multi-GPU fused kernels☆270Updated this week
- mini-swe-agent-plus: a tiny (~100 LOC) GitHub issue fixer—now with a robust multi-line text edit tool.☆26Jan 20, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Post-training with Tinker☆4,117Updated this week
- ThetaEvolve: Test-time Learning on Open Problems, enabling RL training on AlphaEvolve/OpenEvolve and emphasizing scaling test-time comput…☆177Feb 27, 2026Updated 6 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 6 months ago
- Streamline on-policy/off-policy distillation workflows in a few lines of code☆109Aug 5, 2026Updated last month
- AI powered Virtual Desktop☆17Aug 25, 2026Updated 2 weeks ago
- Experimenting Heuristic Learning with ImageNet☆66May 25, 2026Updated 3 months ago
- Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.☆992Updated this week
- ☆23Jan 5, 2025Updated last year
- KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels☆51Jun 1, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards☆40Jun 1, 2026Updated 3 months ago
- A compact high-signal benchmark for evaluating frontier agents☆37Aug 3, 2026Updated last month
- Code of "Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment" (2025).☆14Apr 4, 2025Updated last year
- Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average…☆159Updated this week
- This repository contains the official code for the paper: "Prompt Injection: Parameterization of Fixed Inputs"☆32Sep 13, 2024Updated 2 years ago
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,295Updated this week
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 7 months ago