MyPhoneBench: Do Phone-Use Agents Respect Your Privacy?
☆24Apr 3, 2026Updated 5 months ago
Alternatives and similar repositories for MyPhoneBench
Users that are interested in MyPhoneBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PhoneHarness runtime harness for mixed-action phone agents☆68Jun 17, 2026Updated 2 months ago
- ☆15Jan 27, 2025Updated last year
- Training open models for agentic phone use with real-app and mock-app environments.☆59Aug 24, 2026Updated 2 weeks ago
- LiveClin is a live benchmark designed for the faithful replication of clinical practice☆18Feb 27, 2026Updated 6 months ago
- [ICLR'25] ApolloMoE: Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts☆53Nov 20, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆35Feb 12, 2026Updated 7 months ago
- Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model☆34May 15, 2026Updated 3 months ago
- Official codebase for the paper "WorldMemArena: Evaluating Multimodal Agent Memory Through Action–World Interaction"☆26May 29, 2026Updated 3 months ago
- ☆74Oct 23, 2025Updated 10 months ago
- Official resources of "Hierarchical Verbalizer for Few-Shot Hierarchical Text Classification" (ACL 2023 long).☆27Jul 30, 2023Updated 3 years ago
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols☆20Nov 19, 2025Updated 9 months ago
- Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)☆59May 12, 2025Updated last year
- [ECCV 2024] Official implementation of "Uncertainty Calibration with Energy Based Instance-wise Scaling in the Wild Dataset"☆11Aug 13, 2024Updated 2 years ago
- Towards Fine-grained Audio Captioning with Multimodal Contextual Cues☆90Jan 4, 2026Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The code and dataset for "FastRE: Towards Fast Relation Extraction with Convolutional Encoder and Improved Cascade Binary Tagging Framewo…☆25Aug 13, 2022Updated 4 years ago
- MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria☆78Oct 16, 2024Updated last year
- A Paper collection for LLM based Patient Simulators☆132Jan 7, 2026Updated 8 months ago
- A collection of research on specialized medical LLMs for specific diseases and distinct medical specialties, organized by ICD-10 chapters…☆66Oct 10, 2025Updated 11 months ago
- Code, Data and Model for Paper "Learning from Peers in Reasoning Models"☆26May 13, 2025Updated last year
- ☆10Feb 6, 2025Updated last year
- ☆13Jul 14, 2024Updated 2 years ago
- Official codebase for the paper "Auditing Agent Harness Safety"☆53May 19, 2026Updated 3 months ago
- ☆10Jul 13, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"☆38Jul 11, 2024Updated 2 years ago
- FastLongSpeech is a novel framework designed to extend the capabilities of Large Speech-Language Models for efficient long-speech process…☆16Jul 22, 2025Updated last year
- ☆11May 24, 2024Updated 2 years ago
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 5 months ago
- ☆13Jun 16, 2021Updated 5 years ago
- A trainable user simulator☆34Jun 30, 2025Updated last year
- distill large scale web page text☆12Jul 29, 2023Updated 3 years ago
- Download, parse, and filter data from Court Listener, part of the FreeLaw projects. Data-ready for The-Pile.☆16Jun 3, 2023Updated 3 years ago
- Code for ICLR 2022 Paper (HyperDQN: A Randomized Exploration Method for Deep Reinforcement Learning)☆12Nov 28, 2023Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆14May 30, 2019Updated 7 years ago
- [ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning☆405Mar 30, 2026Updated 5 months ago
- ☆11May 9, 2022Updated 4 years ago
- Code for Paper (ReMax: A Simple, Efficient and Effective Reinforcement Learning Method for Aligning Large Language Models)☆201Dec 16, 2023Updated 2 years ago
- ☆19Nov 11, 2023Updated 2 years ago
- This is the code repo for the paper AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play (NeurIPS 2025 Spotl…☆25Sep 29, 2025Updated 11 months ago
- Official code for "ConTSG-Bench: A Unified Benchmark for Conditional Time Series Generation" (ICML 2026)☆17May 2, 2026Updated 4 months ago