π΅ Code for our EMNLP 2025 Main paper: "FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games"
β28Apr 26, 2026Updated 4 months ago
Alternatives and similar repositories for FlashAdventure
Users that are interested in FlashAdventure are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π§π» Code and benchmark for our Findings of ACL 2024 paper - "TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playingβ¦β23Dec 20, 2024Updated last year
- [ICML 2025] Adaptive Self-improvement LLM Agentic System for ML Library Developmentβ17Jan 6, 2026Updated 8 months ago
- An Ultra-Long Output Reinforcement Learning Approachβ23Jul 31, 2025Updated last year
- [CVPR 2024] "Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition"β11Feb 27, 2024Updated 2 years ago
- Benchmark environment for evaluating vision-language models (VLMs) on popular video games!β373May 30, 2025Updated last year
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICLR 2024] DMBP: Diffusion Model-Based Predictor for Robust Offline Reinforcement Learning against State Observations Perturbations.β17May 24, 2024Updated 2 years ago
- [ACL '26 Findings] V-MAGE: A Game Evaluation Framework for Assessing Visual-Centric Capabilities in MLLMsβ27Apr 28, 2026Updated 4 months ago
- β14Dec 8, 2025Updated 9 months ago
- [ICLR 2026] LLM/VLM gaming agents and model evaluation through games.β982Nov 16, 2025Updated 10 months ago
- β61Oct 18, 2024Updated last year
- A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Modelsβ29Nov 25, 2024Updated last year
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"β19Mar 15, 2024Updated 2 years ago
- This is the repository for paper EscapeBench: Pushing Language Models to Think Outside the Boxβ18Dec 19, 2024Updated last year
- [COLM 2026] An efficient 3D sampling method for long-CoT LLM.β16May 25, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β25Nov 30, 2020Updated 5 years ago
- A secure, zero-touch API provisioning protocol driven by contracts and powered by AIβ13May 4, 2025Updated last year
- β24Dec 30, 2024Updated last year
- β17Apr 14, 2026Updated 5 months ago
- The code and data for "Summary-Oriented Vision Modeling for Multimodal Abstractive Summarization"β11May 16, 2023Updated 3 years ago
- Official code and dataset for our NAACL 2024 paper: DialogCC: An Automated Pipeline for Creating High-Quality Multi-modal Dialogue Dataseβ¦β13Jun 24, 2024Updated 2 years ago
- β25Nov 22, 2024Updated last year
- β12Jul 21, 2022Updated 4 years ago
- ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning. In ICCV, 2021.β64Nov 18, 2021Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACM MM '24 Poster] Official repository of paper titled "Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIPβs Zeroβ¦β10Aug 6, 2024Updated 2 years ago
- Data and scripts for the proper evaluation of cross-lingual embeddings in multiple languagesβ15Apr 11, 2020Updated 6 years ago
- [ICCV 25] Official repository of "Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialβ¦β32Apr 1, 2026Updated 5 months ago
- β129Oct 3, 2025Updated 11 months ago
- β76Jun 10, 2025Updated last year
- [ICLR 2024 Spotlight] Code for ICLR 2024 paper "Towards Robust Offline Reinforcement Learning under Diverse Data Corruption"β22Nov 25, 2024Updated last year
- β12May 17, 2022Updated 4 years ago
- Code for paper OpenWebRL: Online Multi-Turn Reinforcement Learning for Visual Web Agentsβ51Aug 17, 2026Updated last month
- Implementation of the dataset defined in Spiking Neural Networks for event-based action recognition: A new task to understand their advaβ¦β16Aug 9, 2023Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- β11Oct 16, 2023Updated 2 years ago
- Uncertainty-Guided Pseudo-Labelling with Model Averagingβ11Mar 17, 2026Updated 6 months ago
- The source code of ExFunTubeβ10Aug 8, 2025Updated last year
- Official Implementation for "In-Context Reinforcement Learning from Noise Distillation"β35Sep 18, 2024Updated 2 years ago
- The code for paper "EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning"β40Jul 13, 2026Updated 2 months ago
- A Sober Look at Language Model Reasoningβ92Nov 18, 2025Updated 10 months ago
- Official project page and code repository for WiT, a pixel space diffusionβ17Sep 14, 2026Updated last week