Your Company Bench: Long-horizon coherence benchmark in simulated time to test AI agent abilities to manage resources and maximize returns as a tech startup founder
☆131Aug 13, 2026Updated 2 weeks ago
Alternatives and similar repositories for yc-bench
Users that are interested in yc-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SimLab is the data layer for creating simulations to QA, evaluate, hillclimb, and refine agents.☆24Aug 5, 2026Updated 3 weeks ago
- Continual Learning Bench☆212Jul 19, 2026Updated last month
- ☆24Feb 5, 2025Updated last year
- [NeurIPS Datasets & Benchmarks 2025] SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning☆16Dec 2, 2025Updated 8 months ago
- Code repo for paper: Effective Strategies for Asynchronous Software Engineering Agents☆70Apr 2, 2026Updated 4 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Harness for running and evaluating AI agents against RL environments☆247Updated this week
- ☆15Jun 11, 2025Updated last year
- An enterprise deep research benchmark☆42Apr 22, 2026Updated 4 months ago
- RTT timing measurements across OpenClaw npm releases.☆32Updated this week
- ☆25Apr 8, 2026Updated 4 months ago
- The official repo for DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graph☆19Oct 13, 2024Updated last year
- convert entire directories into structured markdown files. handles large files gracefully, skips unwanted directories (node_modules, .git…☆43Mar 12, 2025Updated last year
- Using modal.com to process FineWeb-edu data☆20Updated this week
- Code for "Reasoning to Learn from Latent Thoughts"☆134Mar 28, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Efficient subtyping of ovarian cancer histopathology whole slide images using active sampling in multiple instance learning☆10Mar 19, 2024Updated 2 years ago
- Source code of SACD(Super-resolution with Auto-Correlation two-step Deconvolution)☆13Dec 12, 2022Updated 3 years ago
- Public repository for the Remote Labor Index (RLI)☆75Nov 3, 2025Updated 9 months ago
- ☆44Apr 26, 2026Updated 4 months ago
- ☆19Oct 25, 2025Updated 10 months ago
- Evaluation of LLMs on latest math competitions☆278Jun 23, 2026Updated 2 months ago
- autonomous nanogpt optimizer speedrun☆110May 14, 2026Updated 3 months ago
- [NeurIPS'25 Spotlight] ARM: Adaptive Reasoning Model☆68Apr 6, 2026Updated 4 months ago
- LMTuner: Make the LLM Better for Everyone☆38Sep 21, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Codebase for EnterpriseOps-Gym from ServiceNow☆121Aug 20, 2026Updated last week
- Training API and CLI☆720Updated this week
- ☆11Jul 21, 2024Updated 2 years ago
- Inspect AI interface to Harbor tasks☆21Updated this week
- The repository for papaer "Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs"☆14Dec 16, 2024Updated last year
- ☆10Jan 28, 2019Updated 7 years ago
- Gym-Anything: Turn any Software into an Agent Environment☆278Updated this week
- ☆13Jul 2, 2025Updated last year
- Sequential Monte Carlo Speculative Decoding☆52Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- Smart commit messages☆18Oct 25, 2024Updated last year
- 一个用于课程小论文排版的LaTeX模板。☆10Oct 21, 2019Updated 6 years ago
- [ICML 2024] Code for the paper "MoE-RBench: Towards Building Reliable Language Models with Sparse Mixture-of-Experts"☆11Jul 1, 2024Updated 2 years ago
- ☆12Aug 21, 2024Updated 2 years ago
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,732Jul 23, 2026Updated last month
- A conversational voice-to-voice open-weights LLM-powered assistant designed to run on high-end consumer or workstation class hardware☆75Jul 6, 2026Updated last month