Measuring General Intelligence With Generated Games (Preprint)
☆25Jul 30, 2025Updated last year
Alternatives and similar repositories for gg-bench
Users that are interested in gg-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆22Aug 18, 2024Updated 2 years ago
- ☆34Jun 21, 2024Updated 2 years ago
- Official Code Repository for EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents (COLM 2024)☆41Jul 13, 2024Updated 2 years ago
- [TMLR'25] "Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents"☆108Oct 5, 2025Updated 11 months ago
- Repository for the paper Stream of Search: Learning to Search in Language☆152Feb 3, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LLM play 20questions with itself☆13Mar 31, 2023Updated 3 years ago
- ☆15May 9, 2024Updated 2 years ago
- ☆10Nov 6, 2024Updated last year
- A lightweight computational physics framework, based on the organization of turboWAVE. Implements a "Simulation, PhysicsModule, ComputeTo…☆12Jul 23, 2026Updated 2 months ago
- ☆23Dec 11, 2024Updated last year
- 同济大学数据挖掘课程期末作业:股票走势预测☆11Jan 11, 2021Updated 5 years ago
- Simple, extensible implementations of some meta-learning algorithms in Jax☆11Oct 6, 2020Updated 5 years ago
- a benchmark to evaluate the situated inductive reasoning☆20Jan 7, 2025Updated last year
- Multi-robot Reinforcement Learning Scalable Training School (MRST) is a training and evaluation platform for reinforcement learning rease…☆11Sep 6, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Low-rank sparse attention decomposition for LLM interpretability; active development continues in Llamascopium☆30Nov 9, 2025Updated 10 months ago
- ☆39Sep 6, 2021Updated 5 years ago
- Official Code Release for "Training a Generally Curious Agent"☆50May 18, 2025Updated last year
- Official code for "A General Learning Framework for Open Ad Hoc Teamwork Using Graph-based Policy Learning"☆15Mar 1, 2023Updated 3 years ago
- An automatic prompt iteration and optimization generator suitable for any scenario☆16Jan 31, 2025Updated last year
- ☆18Feb 13, 2025Updated last year
- hsvgbkhgbv / Thermostat-assisted-continuously-tempered-Hamiltonian-Monte-Carlo-for-Bayesian-learningThermostat-assisted continuously-tempered Hamiltonian Monte Carlo for Bayesian learning☆10Dec 10, 2018Updated 7 years ago
- ☆15Mar 26, 2024Updated 2 years ago
- Cooperative Graph-based Networked Agent Challenges for Multi-Agent Reinforcement Learning☆16Jan 26, 2026Updated 8 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆17Dec 16, 2025Updated 9 months ago
- ☆22Feb 29, 2024Updated 2 years ago
- A Codex skill (via CLI) which runs codex in parallel to reflect on previous conversations to brainstorm some skills you could add to your…☆20Jan 23, 2026Updated 8 months ago
- ☆17Jun 10, 2025Updated last year
- [CVPRW'23] The official PyTorch implementation of NamedMask☆23Jun 12, 2023Updated 3 years ago
- About Official PyTorch implementation of "Query-Efficient Black-Box Red Teaming via Bayesian Optimization" (ACL'23)☆15Jul 9, 2023Updated 3 years ago
- Official PyTorch implementation of "Rethinking Value Function Learning for Generalization in Reinforcement Learning" (NeurIPS 2022)☆15Feb 20, 2023Updated 3 years ago
- Rewarded soups official implementation☆66Sep 27, 2023Updated 3 years ago
- Official PyTorch implementation of "Neural Relation Graph: A Unified Framework for Identifying Label Noise and Outlier Data" (NeurIPS'23)☆15Dec 4, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 5 months ago
- SmartPlay is a benchmark for Large Language Models (LLMs). Uses a variety of games to test various important LLM capabilities as agents. …☆145Apr 11, 2024Updated 2 years ago
- ☆29Aug 27, 2025Updated last year
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 8 months ago
- Fine-tuned MARL algorithms on SMAC (100% win rates on most scenarios)☆19Aug 20, 2023Updated 3 years ago
- ☆33Sep 22, 2025Updated last year
- ☆20Sep 16, 2025Updated last year