Latest Evaluation Toolkit (LatestEval). Assessing the language models with latest, uncontaminated materials.
☆29Feb 17, 2025Updated last year
Alternatives and similar repositories for LatestEval
Users that are interested in LatestEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Longitudinal Evaluation of LLMs via Data Compression☆32May 29, 2024Updated 2 years ago
- Public code release for the paper "Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training"☆11Oct 27, 2025Updated 9 months ago
- Source code of paper “A Novel Three-Stage Learning Framework for Low-Resource Knowledge-Grounded Dialogue Generation”☆16Nov 25, 2021Updated 4 years ago
- Benchmarking Commonsense Reasoning in Real-World Tasks☆12Dec 14, 2023Updated 2 years ago
- ☆22Dec 1, 2021Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Evaluating LLMs with Dynamic Data☆117Updated this week
- ☆11Jul 11, 2023Updated 3 years ago
- Framework for Cost-Effective Language Model Choice☆16Dec 12, 2023Updated 2 years ago
- code and dataset of EMNLP 2020 paper "PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge"☆12Nov 6, 2020Updated 5 years ago
- (EACL 2021) Discourse-Aware Unsupervised Summarization of Long Scientific Documents☆25Jun 12, 2023Updated 3 years ago
- Website for release of TellMeWhy dataset for why question answering☆14Nov 11, 2022Updated 3 years ago
- this repository contains the source code for the ACL 2019 paper "Generating Long and Informative Reviews with Aspect-Aware Coarse-to-Fine…☆37Nov 29, 2019Updated 6 years ago
- Knowledge Infused Decoding☆70Dec 31, 2023Updated 2 years ago
- Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts☆26Feb 23, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Self-Supervised Alignment with Mutual Information☆20May 24, 2024Updated 2 years ago
- Code and data setup for the paper "Are Diffusion Models Vision-and-language Reasoners?"☆33Mar 15, 2024Updated 2 years ago
- Submissions, baselines and evaluations scripts for the 2nd version of the WebNLG+ Challenge 2020☆13Feb 1, 2022Updated 4 years ago
- The Paper List on Data Contamination for Large Language Models Evaluation.☆117Jun 2, 2026Updated last month
- Training a reward model for RLHF using RWKV.☆15Jun 5, 2023Updated 3 years ago
- ☆35Jan 7, 2026Updated 6 months ago
- RWKV-7 mini☆12Mar 29, 2025Updated last year
- Direct Preference Optimization for RWKV, aiming for RWKV-5 and 6.☆11Mar 1, 2024Updated 2 years ago
- Explore what LLMs are really leanring over SFT☆28Mar 30, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- GoldFinch and other hybrid transformer components☆16Dec 9, 2025Updated 7 months ago
- Chicago Social Interaction Model (chiSIM) framework repository☆12Aug 9, 2023Updated 2 years ago
- LLM4HWDesign Starting Toolkit☆20Oct 4, 2024Updated last year
- PyTorch implementation of experiments in the paper Aligning Language Models with Human Preferences via a Bayesian Approach☆32Nov 6, 2023Updated 2 years ago
- Do Large Language Models Know What They Don’t Know?☆103Nov 8, 2024Updated last year
- Official repository for ACL 2025 paper "Model Extrapolation Expedites Alignment"☆75May 20, 2025Updated last year
- Synthetic data generation for TODs☆23Jul 17, 2024Updated 2 years ago
- Code for MERMAID : Metaphor Generation with Symbolism and Discriminative Decoding☆11May 2, 2022Updated 4 years ago
- Xlore2.0 Code[BaiduExtractor, HudongExtractor, WikiExtractor, XloreData, XloreWeb]☆12Apr 5, 2017Updated 9 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- In-Context Learning User Simulators for Task-Oriented Dialog Systems☆29Jun 2, 2023Updated 3 years ago
- RWKV Wiki website (archived, please visit official wiki)☆10Mar 26, 2023Updated 3 years ago
- The official codebase for "Experiential Reinforcement Learning" - https://arxiv.org/pdf/2602.13949v1☆76Jul 2, 2026Updated 3 weeks ago
- PLATO dialog model with pre-trained parameters in pytorch version☆29May 20, 2022Updated 4 years ago
- ☆13Dec 5, 2022Updated 3 years ago
- Repository containing the website for the EMNLP 2023 conference☆17Feb 12, 2025Updated last year
- ☆13Feb 26, 2023Updated 3 years ago