[AAAI 2026] Data and Code for Paper IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
☆47Nov 24, 2025Updated 7 months ago
Alternatives and similar repositories for IS-Bench
Users that are interested in IS-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repository of "Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Ste…☆28Mar 9, 2026Updated 4 months ago
- [EMNLP 2025] The code repo of paper "X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Com…☆41Nov 24, 2025Updated 7 months ago
- [NeurIPS 2025] Official repository of RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents☆123Dec 2, 2025Updated 7 months ago
- Official Implementation of "Geometrically-Constrained Agent for Spatial Reasoning"☆89Apr 7, 2026Updated 3 months ago
- Diagnostic Framework for LLMs and MLLMs☆39Mar 2, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official Repo of Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents☆90Jun 2, 2026Updated last month
- (ICLR 2026 🔥) Code for "The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs"☆79Feb 9, 2026Updated 5 months ago
- 😎 A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, Agent, and Beyond☆355Jan 22, 2026Updated 5 months ago
- Codes for paper "SafeAgentBench: A Benchmark for Safe Task Planning of \\ Embodied LLM Agents"☆74Feb 25, 2025Updated last year
- Responsible Robotic Manipulation☆16Aug 31, 2025Updated 10 months ago
- A Diagnostic Guardrail Framework for AI Agent Safety and Security☆669Jun 8, 2026Updated last month
- The officalimplement of dLLM-Factory☆25Jul 12, 2025Updated last year
- Official repository of DARE: Diffusion Large Language Models Alignment and Reinforcement Executor☆213Updated this week
- ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents☆58May 13, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆27Oct 9, 2025Updated 9 months ago
- ☆21Feb 3, 2025Updated last year
- Repo for paper "Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability"☆108Apr 23, 2026Updated 2 months ago
- [NeurIPS 2024] Data exporter for SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh Dataset☆16Nov 8, 2024Updated last year
- AI-powered scientific figure generator using LLM analysis and nano banana for publication-quality visualizations☆19Feb 7, 2026Updated 5 months ago
- [AAAI26] Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilitie…☆11Feb 7, 2026Updated 5 months ago
- 北京大学 2024 年秋 ICS 相关资料☆12May 14, 2025Updated last year
- AgenTracer: A Lightweight Failure Attributor for Agentic Systems☆97Nov 12, 2025Updated 8 months ago
- [ACL2026]Code Repo for paper "Scaling Behaviors of LLM Reinforcement Learning Post-Training"☆24Jul 1, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆19Dec 23, 2025Updated 6 months ago
- The repository of the paper "REEF: Representation Encoding Fingerprints for Large Language Models," aims to protect the IP of open-source…☆79Jan 16, 2025Updated last year
- ☆25Jan 29, 2026Updated 5 months ago
- Code for our paper titled "Lens: Rethinking Multilingual Enhancement for Large Language Models"☆12Oct 15, 2024Updated last year
- ☆10Mar 19, 2024Updated 2 years ago
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- ENACT is a benchmark that evaluates embodied cognition through world modeling from egocentric interaction. It is designed to be simple an…☆52Nov 27, 2025Updated 7 months ago
- 🧨 TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?☆80Nov 27, 2025Updated 7 months ago
- ☆30May 22, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official code for "From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation" (ICLR2026)☆37Mar 1, 2026Updated 4 months ago
- [NeurIPS 2023] "Diversified Outlier Exposure for Out-of-Distribution Detection via Informative Extrapolation"☆11Oct 6, 2023Updated 2 years ago
- This repository includes the code to download the curated HuggingFace papers into a single markdown formatted file☆16Jul 26, 2024Updated last year
- 实现了《编译原理实践与指导教程》一书中的C--编译器☆11Oct 10, 2024Updated last year
- Official GitHub repository for the paper "Adversarial Attacks on Robotic Vision Language Action Models"☆35May 28, 2025Updated last year
- Code for "Adversarial Illusions in Multi-Modal Embeddings"☆32Aug 4, 2024Updated last year
- Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction Metric☆16Nov 12, 2025Updated 8 months ago