OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
☆19Jun 25, 2024Updated 2 years ago
Alternatives and similar repositories for opencompass
Users that are interested in opencompass are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2024] MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues☆152Jul 24, 2024Updated 2 years ago
- ☆20Jan 9, 2024Updated 2 years ago
- Gantt Chart using echarts☆13Mar 31, 2021Updated 5 years ago
- [ICLR 26] Context Tokens are Anchors: Understanding the Repeat Curse in dMLLMs from an Information Flow Perspective☆16Mar 6, 2026Updated 4 months ago
- Measuring if attention is explanation with ROAR☆22Mar 3, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The implementation of the IEEE S&P 2024 paper MM-BD: Post-Training Detection of Backdoor Attacks with Arbitrary Backdoor Pattern Types Us…☆16May 12, 2024Updated 2 years ago
- [NeurIPS 2021] "Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models" by Boxin Wang*, Chejian Xu*, Shuoh…☆13Apr 3, 2023Updated 3 years ago
- [ACL 2026 Main] MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools.☆25Apr 8, 2026Updated 3 months ago
- TKDE'23: A Survey and Experimental Study on Privacy-Preserving Trajectory Data Publishing☆12May 5, 2023Updated 3 years ago
- ☆13Mar 15, 2022Updated 4 years ago
- MemSearcher is a search agent that keeps a compact, iteratively-updated memory instead of the full interaction history, trained end-to-en…☆27Jun 29, 2026Updated 3 weeks ago
- Several variations of a dot product benchmark.☆11Dec 10, 2012Updated 13 years ago
- This is a niche collection of research papers which are proven to be gradients pushing the field of Natural Language Processing, Deep Lea…☆25Nov 19, 2024Updated last year
- (ACM MM24) This is the offical repository of GIST: Improving Parameter Efficient Fine Tuning via Knowledge Interaction.☆11Jan 28, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2023] The official implementation of our CVPR 2023 paper "Detecting Backdoors During the Inference Stage Based on Corruption Robust…☆25May 25, 2023Updated 3 years ago
- 本项目 实现自《差分隐私下满足一致性的轨迹流量发布方法》,作者蔡剑平☆12Sep 2, 2019Updated 6 years ago
- Ladder Side-Tuning在CLUE上的简单尝试☆23Jun 20, 2022Updated 4 years ago
- ☆29Jun 17, 2024Updated 2 years ago
- EMNLP 2022: Analyzing and Evaluating Faithfulness in Dialogue Summarization☆13Mar 20, 2025Updated last year
- MetricEval: A framework that conceptualizes and operationalizes four main components of metric evaluation, in terms of reliability and va…☆12Nov 6, 2023Updated 2 years ago
- The code for paper "ProQA: Structural Prompt-based Pre-training for Unified Question Answering"☆11Feb 7, 2023Updated 3 years ago
- A curated list of personalized Language model / Large language model (continually updated)☆10Nov 17, 2023Updated 2 years ago
- [ICLR 2026] Evaluating Text Creativity across Diverse Domains: A Dataset and a Large Language Model Evaluator☆18Feb 28, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- LBSN based on foursquare dataset☆14Apr 26, 2019Updated 7 years ago
- Codes for "Benchmarking the Generation of Fact Checking Explanations"☆10Aug 16, 2024Updated last year
- 2024编译系统实现赛RISC-V赛道一等奖作品(A compiler of SysY (subset of C) )☆24Sep 4, 2024Updated last year
- This repository contains our implementation of the ontology matching framework based on representation learning.☆15May 7, 2018Updated 8 years ago
- Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems☆38Jul 16, 2026Updated last week
- (ICME24) This is the offical repository of iDAT: inverse Distillation Adapter-Tuning.☆13Apr 3, 2024Updated 2 years ago
- Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs☆13Feb 13, 2024Updated 2 years ago
- ☆17Aug 6, 2023Updated 2 years ago
- Investigating Cultural Alignment of Large Language Models☆13Aug 14, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- https://openreview.net/forum?id=OC1o4_OI6Jw☆13May 27, 2022Updated 4 years ago
- Uncertainty-Aware Curriculum Learning for Neural Machine Translation (ACL 2020)☆11Jun 12, 2020Updated 6 years ago
- code for promptCSE, emnlp 2022☆11Apr 10, 2023Updated 3 years ago
- BackTime: Backdoor Attacks on Multivariate Time Series Forecasting☆32Apr 14, 2025Updated last year
- ☆16Mar 27, 2023Updated 3 years ago
- 福州大学博士研究生毕业论文Latex模板☆18May 27, 2024Updated 2 years ago
- Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments (Zhou et al., EMNLP 2024)☆14Oct 3, 2024Updated last year