☆41May 16, 2026Updated 4 months ago
Alternatives and similar repositories for harbor-datasets
Users that are interested in harbor-datasets are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An agent for auditing repositories of traces for violations of safety properties. Automatically finds cheating (task-level gaming and har…☆15Jun 6, 2026Updated 3 months ago
- Framework for evaluating and improving agents☆5,411Updated this week
- ☆411Apr 30, 2026Updated 4 months ago
- Multi-agent synthetic data generation pipeline capable of generating and validating long horizon terminal/coding tasks for RL training☆74Jul 28, 2025Updated last year
- 💻 SETA: Scaling Environments for Terminal Agents - Environments☆147Feb 16, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Measuring and evolving with the frontier of agent work☆735Updated this week
- ☆12Mar 18, 2024Updated 2 years ago
- MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following☆16Oct 31, 2024Updated last year
- Simple WebSockets API☆10Nov 18, 2021Updated 4 years ago
- Python package for extractive NLP using the OpenAI API☆17Aug 28, 2024Updated 2 years ago
- ☆14Mar 15, 2022Updated 4 years ago
- ☆15Jan 20, 2026Updated 8 months ago
- Released code for「Stance Detection on Social Media with Background Knowledge」in EMNLP2023.☆20Apr 23, 2024Updated 2 years ago
- [ICLR2026🔥Oral] SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving☆15Feb 26, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A benchmark for LLMs on complicated tasks in the terminal☆2,590Jul 11, 2026Updated 2 months ago
- My implementation of the model KosmosG from "KOSMOS-G: Generating Images in Context with Multimodal Large Language Models"☆13Nov 11, 2024Updated last year
- ☆21Jul 16, 2024Updated 2 years ago
- A Datasette instance for searching WebVid-10M☆15Sep 30, 2022Updated 3 years ago
- Run SWE-bench evaluations remotely☆81Aug 14, 2025Updated last year
- 登录脚本☆12Nov 4, 2022Updated 3 years ago
- ☆13Jan 21, 2022Updated 4 years ago
- ☆15Aug 4, 2024Updated 2 years ago
- Hasura GraphQL Engine on Render☆15Aug 28, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆83Jun 25, 2026Updated 2 months ago
- ☆23Nov 7, 2023Updated 2 years ago
- GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's T…☆411Aug 24, 2025Updated last year
- Fast Topological Clustering with Wasserstein Distance (ICLR 2022)☆12Jun 24, 2022Updated 4 years ago
- ☆23Mar 20, 2025Updated last year
- Python SDK for Weaver.☆17Updated this week
- Demo showing how to sync data with ElectricSQL from Postgres to Cloudflare's Workers KV☆17Aug 20, 2024Updated 2 years ago
- Benchmarking execution environments ability to prevent reward hacking in agent evals.☆18Updated this week
- Open ChatGLM Eyes to See the World☆13Mar 30, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is the source code of FUSION, a safety-aware causal representation for generalizable driving agents.☆29Oct 23, 2024Updated last year
- A compact high-signal benchmark for evaluating frontier agents☆39Aug 3, 2026Updated last month
- 💻 SETA: Scaling Environments for Terminal Agents☆155Jul 28, 2026Updated last month
- Code for our EMNLP 2020 paper "Uncertainty-Aware Label Refinement for Sequence Labeling"☆22Oct 4, 2020Updated 5 years ago
- [ICLR 2026]InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research☆17Feb 3, 2026Updated 7 months ago
- ☆52Jan 30, 2026Updated 7 months ago
- Decoding Echo Chambers: LLM-Powered Simulations Revealing Polarization in Social Networks☆25Dec 11, 2024Updated last year