☆58Apr 4, 2025Updated last year
Alternatives and similar repositories for super-benchmark
Users that are interested in super-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code to reproduce LREC Paper Simplifying Semantic Annotations of SMCalFlow☆24Mar 28, 2024Updated 2 years ago
- ☆78Nov 23, 2025Updated 8 months ago
- Diverse Demonstrations Improve In-context Compositional Generalization☆13Jul 7, 2023Updated 3 years ago
- Reasoning by Communicating with Agents☆30Apr 29, 2025Updated last year
- https://footprints.baulab.info☆17Oct 4, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implementation code for ACL2024:Advancing Parameter Efficiency in Fine-tuning via Representation Editing☆15Apr 20, 2024Updated 2 years ago
- ☆28Nov 19, 2025Updated 8 months ago
- Discovering Data-driven Hypotheses in the Wild☆157Jun 9, 2025Updated last year
- Bridging the Generalization Gap in Text-to-SQL Parsing with Schema Expansion☆13Jul 26, 2023Updated 3 years ago
- XmodelLM☆38Nov 19, 2024Updated last year
- 👻 Code and benchmark for our EMNLP 2023 paper - "FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions"☆62May 31, 2024Updated 2 years ago
- Efficient Scaling laws and collaborative pretraining.☆23Jul 19, 2026Updated 3 weeks ago
- WorldSense benchmark for grounded reasoning in language models☆25Nov 28, 2023Updated 2 years ago
- Code for "[COLM'25] RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing"☆24Mar 18, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆22Mar 23, 2025Updated last year
- Official implementation of ECCV24 paper: POA☆24Aug 8, 2024Updated 2 years ago
- Project for SNARE benchmark☆11Jun 5, 2024Updated 2 years ago
- AI Assistance for Writing Scientific Alt Text☆14Feb 7, 2024Updated 2 years ago
- [ICLR 2026] Rectifying LLM Thought From Lens of Optimization☆15Dec 5, 2025Updated 8 months ago
- ☆13Dec 12, 2025Updated 8 months ago
- HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation Perturbation [ACL 2023]☆14Jul 11, 2023Updated 3 years ago
- Control LLM☆23Apr 6, 2025Updated last year
- Learning to route instances for Human vs AI Feedback (ACL Main '25)☆29Jul 23, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- NeuralWOZ: Learning to Collect Task-Oriented Dialogue via Model-based Simulation (ACL-IJCNLP 2021)☆36Jul 22, 2021Updated 5 years ago
- Source code for Grounded Adaptation for Zero-shot Executable Semantic Parsing☆21Feb 1, 2021Updated 5 years ago
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum Masking☆13Feb 5, 2023Updated 3 years ago
- ☆21Oct 22, 2021Updated 4 years ago
- The official github repo for MixEval-X, the first any-to-any, real-world benchmark.☆17Feb 15, 2025Updated last year
- CodeRepoQA dataset☆15Feb 19, 2025Updated last year
- This is an implementation of the paper "Are We Done with Object-Centric Learning?"☆14Jun 21, 2026Updated last month
- Evaluating Reward Models in Multilingual Settings (ACL Main '25)☆44May 16, 2025Updated last year
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models☆60Nov 26, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆14Jan 22, 2025Updated last year
- Code for paper: Aligning Large Language Models with Representation Editing: A Control Perspective☆35Jan 31, 2025Updated last year
- To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models☆34May 21, 2025Updated last year
- ☆15Apr 14, 2025Updated last year
- ☆19Dec 20, 2025Updated 7 months ago
- ☆10Feb 6, 2025Updated last year
- Documentation and common scenarios for newsroom usage of https://github.com/cantino/huginn/☆16Aug 19, 2016Updated 9 years ago