LLM evaluation.
☆16Nov 7, 2023Updated 2 years ago
Alternatives and similar repositories for llm_eval
Users that are interested in llm_eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆27Jul 20, 2024Updated 2 years ago
- The official implementation of the paper "Large Scale Knowledge Washing"☆10Jun 12, 2024Updated 2 years ago
- ☆10Mar 13, 2023Updated 3 years ago
- [COLING 2025] Official repo of paper: "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jail…☆12Jul 26, 2024Updated 2 years ago
- [ACL 2024 Demo] Official GitHub repo for UltraEval: An open source framework for evaluating foundation models.☆257Oct 30, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Scalable Meta-Evaluation of LLMs as Evaluators☆43Feb 15, 2024Updated 2 years ago
- Create a LangChain ReAct agent with multiple tools (Python REPL and DuckDuckGo Search)☆14Updated this week
- Code for paper "Point and Ask: Incorporating Pointing into Visual Question Answering"☆19Oct 4, 2022Updated 3 years ago
- Concept-Pointer-Network-for-Abstractive-Summarization☆19May 17, 2019Updated 7 years ago
- Evaluating LLMs with CommonGen-Lite☆95Mar 21, 2024Updated 2 years ago
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement☆20Jul 4, 2026Updated last month
- [NeurIPS'24] Official PyTorch Implementation of Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment☆58Sep 26, 2024Updated last year
- SC-Safety: 中文大模型多轮对抗安全基准☆152Mar 15, 2024Updated 2 years ago
- 【ACL 2024】 SALAD benchmark & MD-Judge☆175Mar 8, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆11May 8, 2020Updated 6 years ago
- 指针生成网络在 中英文数据集下的应用☆16Mar 10, 2020Updated 6 years ago
- Official implementation of paper: DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers☆67Aug 25, 2024Updated 2 years ago
- ☆21Aug 19, 2024Updated 2 years ago
- PyTorch implementation of DQN☆13Sep 27, 2019Updated 6 years ago
- AIGC 系列报告 2022-2023☆13Feb 25, 2024Updated 2 years ago
- Official github repo for SafetyBench, a comprehensive benchmark to evaluate LLMs' safety. [ACL 2024]☆297Jul 28, 2025Updated last year
- Implementation of "PaLM2-VAdapter:" from the multi-modal model paper: "PaLM2-VAdapter: Progressively Aligned Language Model Makes a Stron…☆17Nov 11, 2024Updated last year
- Repository for "StrongREJECT for Empty Jailbreaks" paper☆158Nov 3, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- FlagEval is an evaluation toolkit for AI large foundation models.☆338Apr 24, 2025Updated last year
- Rationale-enhanced language models are better continual relation learners (EMNLP 2023 Main Conference)☆12Oct 11, 2023Updated 2 years ago
- Deepmark AI enables a unique testing environment for language models (LLM) assessment on task-specific metrics and on your own data so yo…☆104Nov 24, 2023Updated 2 years ago
- THOUGHTSCULPT, a general reasoning and search method for complex tasks☆13Dec 13, 2024Updated last year
- Project of llm evaluation to Japanese tasks☆94Jul 26, 2026Updated last month
- A Modular Framework for Learning on Multimodal Biomedical Knowledge Graphs☆19Jan 12, 2024Updated 2 years ago
- [ICML 2026] HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling☆30May 2, 2026Updated 4 months ago
- A package dedicated for running benchmark agreement testing☆19Sep 18, 2025Updated 11 months ago
- Code to conduct an embedding attack on LLMs☆34Jan 10, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Synthesize bio-plausible neural networks for cognitive tasks, mimicking brain architecture☆11Apr 14, 2021Updated 5 years ago
- Machine Learning Data Fairness and Bias☆15Apr 29, 2026Updated 4 months ago
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated last year
- [ICML 2023] "Robust Weight Signatures: Gaining Robustness as Easy as Patching Weights?" by Ruisi Cai, Zhenyu Zhang, Zhangyang Wang☆16May 4, 2023Updated 3 years ago
- The official repository for the experiments included in the paper titled "Patch-level Routing in Mixture-of-Experts is Provably Sample-ef…☆14Feb 12, 2026Updated 6 months ago
- This will be a command line program, reading a dxf file and saving a converted dxf file, with splines converted to biarcs.☆13Jun 19, 2020Updated 6 years ago
- Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models☆45Jun 14, 2024Updated 2 years ago