A framework for evolving and testing question-answering datasets with various models.
☆26Feb 28, 2024Updated 2 years ago
Alternatives and similar repositories for Self-Evolving-Benchmark
Users that are interested in Self-Evolving-Benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repository contains two datasets with multi-turn adversarial conversations generated by human agents interacting with a dialog model…☆35Jul 16, 2024Updated 2 years ago
- ☆24Jun 16, 2025Updated last year
- FGLA: Fast Generation-Based Gradient Leakage Attacks against Highly Compressed Gradients☆15Mar 17, 2026Updated 6 months ago
- Companion code to https://arxiv.org/abs/2402.15491☆22Sep 18, 2025Updated last year
- ☆13Jul 16, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆12Sep 23, 2024Updated last year
- chinese ner based on rnn☆12Oct 14, 2016Updated 9 years ago
- pytorch implementation of mvp: a multi-stage vision-language pre-training framework☆12Apr 23, 2022Updated 4 years ago
- TabLeak: Tabular Data Leakage in Federated Learning☆16Jul 4, 2024Updated 2 years ago
- [ICASSP '26] This is the code repo for our paper: LegalΔ: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thou…☆32Jul 1, 2026Updated 2 months ago
- The rule-based evaluation subset and code implementation of Omni-MATH☆29Dec 23, 2024Updated last year
- ☆12Sep 8, 2020Updated 6 years ago
- LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning☆39Apr 4, 2024Updated 2 years ago
- Codes and data for CIKM 2022 paper "RuDi: Explaining Behavior Sequence Models by Automatic Statistics Generation and Rule Distillation"☆12Aug 16, 2022Updated 4 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Code for the paper "Jailbreak Large Vision-Language Models Through Multi-Modal Linkage"☆35Dec 6, 2024Updated last year
- ☆11Jun 12, 2023Updated 3 years ago
- Open-source strong baseline for domain generlization re-ID. We will udpate the strong baseline and CFD method~☆10Nov 30, 2021Updated 4 years ago
- ☆12Mar 22, 2025Updated last year
- ☆18May 17, 2025Updated last year
- ☆14May 20, 2025Updated last year
- The source code of Paper "PathQG: Neural Question Generation from Facts".☆23Jan 4, 2021Updated 5 years ago
- [ACL 2024] An Easy-to-use Hallucination Detection Framework for LLMs.☆42Feb 25, 2025Updated last year
- ☆10Dec 16, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- To calculate the BLUE score☆11Jun 7, 2016Updated 10 years ago
- A new dataset of difficult graduate-level applied mathematics problems; evaluations demonstrate that leading LLMs currently exhibit low a…☆30Feb 14, 2025Updated last year
- Trying out diffusion training in federated learning setting.☆18Jan 23, 2024Updated 2 years ago
- 利用大语言模型进行卧底游戏,包括谁是卧底及衍生的发现AI卧底游戏等。☆12Sep 6, 2024Updated 2 years ago
- ☆13Jun 17, 2024Updated 2 years ago
- ☆12Aug 15, 2022Updated 4 years ago
- [ACL 2024] LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement☆196Mar 25, 2024Updated 2 years ago
- Official Implementation of "Semantics-Consistent Feature Search for Self-Supervised Visual Representation Learning" in AAAI2024.☆13Feb 28, 2024Updated 2 years ago
- Research on "Many-Shot Jailbreaking" in Large Language Models (LLMs). It unveils a novel technique capable of bypassing the safety mechan…☆18Aug 6, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Implementation of our ACL2023 paper: Unifying Cross-Lingual and Cross-Modal Modeling Towards Weakly Supervised Multilingual Vision-Langua…☆20Jul 5, 2023Updated 3 years ago
- Robust Point Cloud Processing through Positional Embedding☆14Sep 7, 2023Updated 3 years ago
- Chinese Generation Evaluation☆13Aug 14, 2023Updated 3 years ago
- ☆64Jan 27, 2023Updated 3 years ago
- A simple implementation of DP-RAG☆21Mar 17, 2025Updated last year
- [EMNLP 2025 Findings] Retrieval-Augmented Machine Translation with Unstructured Knowledge☆15Sep 4, 2025Updated last year
- [ACL2024] Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios☆71Aug 5, 2025Updated last year