Evaluating LLMs with CommonGen-Lite
☆95Mar 21, 2024Updated 2 years ago
Alternatives and similar repositories for CommonGen-Eval
Users that are interested in CommonGen-Eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Repository for NPHardEval, a quantified-dynamic benchmark of LLMs☆64Mar 26, 2024Updated 2 years ago
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- ☆151Jan 4, 2024Updated 2 years ago
- Automatically evaluate your LLMs in Google Colab☆695May 7, 2024Updated 2 years ago
- This is a new metric that can be used to evaluate faithfulness of text generated by LLMs. The work behind this repository can be found he…☆31Aug 25, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Benchmarking LLMs with Challenging Tasks from Real Users☆255Nov 3, 2024Updated last year
- Scrape and export data from the Open LLM Leaderboard.☆48Dec 17, 2024Updated last year
- ☆10Jul 6, 2023Updated 3 years ago
- A cost estimator for OpenAI API calls in tqdm loops.☆20Nov 25, 2024Updated last year
- ☆16Jul 20, 2023Updated 3 years ago
- Query-focused summarization data☆44Feb 17, 2023Updated 3 years ago
- N/A☆19Aug 15, 2022Updated 3 years ago
- Using short models to classify long texts☆21Mar 8, 2023Updated 3 years ago
- A complete guide to evaluate LLMs and RAGs. Both theory and code based approaches covered.☆28Nov 16, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Generative AI for real world video editing☆13Oct 20, 2023Updated 2 years ago
- 📝 Reference-Free automatic summarization evaluation with potential hallucination detection☆104Jan 15, 2024Updated 2 years ago
- [ECCV 2024] M3DBench introduces a comprehensive 3D instruction-following dataset with support for interleaved multi-modal prompts.☆61Oct 1, 2024Updated last year
- Scalable Meta-Evaluation of LLMs as Evaluators☆43Feb 15, 2024Updated 2 years ago
- Evaluating LLMs with fewer examples☆181Jul 4, 2026Updated 3 weeks ago
- ☆13Jan 27, 2019Updated 7 years ago
- A simple GPT-based evaluation tool for multi-aspect, interpretable assessment of LLMs.