NeurIPS 2025 Poster
☆26Feb 4, 2025Updated last year
Alternatives and similar repositories for Adaptive_Distractions
Users that are interested in Adaptive_Distractions are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ACL 2025: Search-based multilingual LLM evaluation that discovers English-correct, target-language-wrong failures; code and 6,713-pair da…☆42Sep 21, 2026Updated 2 weeks ago
- NeurIPS 2025 Poster☆24Oct 17, 2025Updated 11 months ago
- ProbeLLM: Automating Principled Diagnosis of LLM Failures☆18Feb 11, 2026Updated 7 months ago
- [NeurIPS 2024] HonestLLM: Toward an Honest and Helpful Large Language Model☆29Jun 10, 2025Updated last year
- [ICLR'25] DataGen: Unified Synthetic Dataset Generation via Large Language Models☆69Mar 8, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- SDE-Harness (Scientific Discovery Evaluation Framework)☆64Sep 19, 2026Updated 2 weeks ago
- [ICLR'26, NAACL'25 Demo] Toolkit & Benchmark for evaluating the trustworthiness of generative foundation models.☆137Aug 22, 2025Updated last year
- PostgreSQL extension which allows to translate a given source SQL statement into another pre-defined SQL statement.☆24Sep 25, 2025Updated last year
- Code for the 2025 ACL publication "Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs"☆33Jun 25, 2025Updated last year
- Official implementation for the paper "Video-Based Reward Modeling for Computer-Use Agents"☆17Mar 14, 2026Updated 6 months ago
- Implements the [TPCH benchmark](http://www.tpc.org/tpch/) for Postgres☆31Apr 11, 2022Updated 4 years ago
- [ACL 2026] A benchmark for evaluating the reliability of text-to-infographic generation with curated test cases and automated question-ba…☆17Jun 8, 2026Updated 4 months ago
- Benchmark for Agentic Powerpoint Editing Tasks☆29Sep 25, 2026Updated 2 weeks ago
- NeurIPS 2026: Official implementation of “VISD: Enhancing Video Reasoning via Structured Self-Distillation”.☆25Oct 1, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML 2024 Oral] Official code repository for MLLM-as-a-Judge.☆95Feb 17, 2025Updated last year
- An Efficient "Factory" to Build Multiple LoRA Adapters☆385Feb 13, 2025Updated last year
- [ICCV 2025] MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation☆23Sep 5, 2025Updated last year
- study notes for IT☆11Feb 22, 2020Updated 6 years ago
- Benchmarking data and script used for LLM multi-agent collaboration systems from AWS Bedrock Agents Science team.☆18Dec 10, 2024Updated last year
- Learned Query Optimizer☆13Mar 16, 2022Updated 4 years ago
- ☆15Sep 2, 2022Updated 4 years ago
- Fleming-VL: Towards Universal Medical Visual Understanding with Multimodal LLMs☆18Nov 6, 2025Updated 11 months ago
- 本项目分享了本人在四川大学计算机学院计算机科学与技术专业的各类课程的资料、学习建议以及作业。欢迎使用,也希望其他校友能为此库提供缺失资料,如果喜欢就Star吧。☆10May 18, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A instruction data generation system for multimodal language models.☆37Jan 31, 2025Updated last year
- From Agentic Intelligence to Interactive Intelligence. Give your AI agent a body and home.☆24Feb 22, 2026Updated 7 months ago
- C++ 实现 BP 神经网络识别手写数字数据集 MNIST☆15Jan 7, 2024Updated 2 years ago
- ☆12May 6, 2022Updated 4 years ago
- ☆23Nov 15, 2022Updated 3 years ago
- Can We Trust Large Language Models?: A Benchmark for Responsible Large Language Models via Toxicity, Bias, and Value-alignment Evaluation☆25Oct 12, 2023Updated 2 years ago
- The official repository for "MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants".☆19Oct 10, 2024Updated last year
- ☆11Dec 23, 2024Updated last year
- Koishi's Day 2025 Paper (NeurIPS 2025): "Codifying Character Logic in Role-Playing"☆26Jan 15, 2026Updated 8 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [ICML 2024] TrustLLM: LLM trustworthiness evaluation across truthfulness, safety, fairness, robustness, privacy and ethics. Python/CLI to…☆633Updated this week
- Fine-tuning, DPO, RLHF, RLAIF on LLMs - Qwen3, Zephyr 7B GPTQ with 4-Bit Quantization, Mistral-7B-GPTQ☆15Jul 5, 2025Updated last year
- [CVPR 2024] Narrative Action Evaluation with Prompt-Guided Multimodal Interaction☆43May 16, 2024Updated 2 years ago
- ☆22May 21, 2025Updated last year
- Official Repo for SvS: A Self-play with Variational Problem Synthesis strategy for RLVR training☆86Sep 25, 2026Updated 2 weeks ago
- [ICLR 2026] "VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?", Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu, L…☆41Jan 30, 2026Updated 8 months ago
- ☆10Jun 29, 2020Updated 6 years ago