Code for Paper: Benchmarking Multi-step Scientific Tool-use in LLM Agents
☆38Jul 5, 2026Updated last month
Alternatives and similar repositories for SciAgentGYM
Users that are interested in SciAgentGYM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2026] Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models☆28Apr 19, 2026Updated 3 months ago
- Heterogeneous Scientific Foundation Model Collaboration☆23May 1, 2026Updated 3 months ago
- This is Official implementation for T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasonin…☆24Mar 5, 2026Updated 5 months ago
- [TrimKV] Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs - [DBTrimKV] Make Each Token Count: Towards Improving Lo…☆17Jul 26, 2026Updated last week
- ☆29Mar 10, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official implementation for paper "Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe"☆36May 12, 2026Updated 2 months ago
- Executive Memory for Coherent Long-Horizon Reasoning!☆86Jan 14, 2026Updated 6 months ago
- ☆137May 12, 2026Updated 2 months ago
- Jointly Optimizing Large Language Models for Reasoning and Self-Refinement☆15Apr 22, 2026Updated 3 months ago
- The official implementation of "LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation"☆22Apr 22, 2025Updated last year
- ☆15Apr 15, 2026Updated 3 months ago
- ☆15Nov 6, 2020Updated 5 years ago
- MUA-RL: MULTI-TURN USER-INTERACTING AGENT REINFORCEMENT LEARNING FOR AGENTIC TOOL USE☆67Nov 5, 2025Updated 9 months ago
- Lab Cookbook☆39Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆12May 23, 2018Updated 8 years ago
- Imbalanced Gradients: A New Cause of Overestimated Adversarial Robustness. (MD attacks)☆11Aug 29, 2020Updated 5 years ago
- Open database of system prompts extracted from frontier LLMs using JustAsk☆39Apr 4, 2026Updated 4 months ago
- this is for the ACM MM paper---Backdoor Attack on Crowd Counting☆17Jul 10, 2022Updated 4 years ago
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 8 months ago
- [CVPRW 2026 Oral] Less Detail, Better Answers: Degradation-Driven Prompting for VQA☆20Apr 25, 2026Updated 3 months ago
- ☆12Jun 16, 2023Updated 3 years ago
- [CVPR2020] Clean-Label Backdoor Attacks on Video Recognition Models☆41Jun 19, 2020Updated 6 years ago
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆12Nov 18, 2023Updated 2 years ago
- ☆30Jun 13, 2026Updated last month
- Data for the paper "A Dataset for Learning University STEM Courses at Scale" by Zhang et al., 2022.☆15Nov 22, 2022Updated 3 years ago
- ☆20Jan 18, 2026Updated 6 months ago
- This is AI implementation (not official) of the DreamGym framework from the paper "Scaling Agent Learning via Experience Synthesis" (arXi…☆45Nov 9, 2025Updated 8 months ago
- Official source code repository for paper BubbleRAG.☆17Jun 1, 2026Updated 2 months ago
- Official repo for "TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders"☆25Apr 9, 2026Updated 3 months ago
- [JMLR] Gradual Domain Adaptation: Theory and Algorithms☆11Jan 14, 2025Updated last year
- ☆21Apr 3, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An arbitrage bot is a smart contract connected to an external automation script that controls its operation.☆2,288Updated this week
- ☆47Apr 8, 2026Updated 3 months ago
- Catalogue of Life toolkit for Python☆11Aug 4, 2020Updated 6 years ago
- Agent-RRM: Exploring Reasoning Reward Model for Agents☆70Mar 17, 2026Updated 4 months ago
- Solv@TUM - The Solvation Free Energy Database☆15Mar 31, 2024Updated 2 years ago
- ☆25Feb 12, 2026Updated 5 months ago
- The code for ACM MM2024 (Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning)☆15Jul 18, 2024Updated 2 years ago