[ACL 2024 Main] NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
☆34Jun 25, 2024Updated 2 years ago
Alternatives and similar repositories for NewsBench
Users that are interested in NewsBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PGRAG☆54Jul 16, 2024Updated 2 years ago
- ☆23Jun 10, 2025Updated last year
- Grimoire is All You Need for Enhancing Large Language Models☆120Feb 29, 2024Updated 2 years ago
- [ACL 2024] User-friendly evaluation framework: Eval Suite & Benchmarks: UHGEval, HaluEval, HalluQA, etc.☆182Jun 7, 2025Updated last year
- ☆65Mar 11, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- xVerify: Efficient Answer Verifier for Reasoning Model Evaluations☆149Nov 13, 2025Updated 10 months ago
- HaluMem is the first operation level hallucination evaluation benchmark tailored to agent memory systems.☆163Sep 3, 2026Updated 3 weeks ago
- CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models☆414May 20, 2025Updated last year
- [ICLR 2025] xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation☆180Nov 14, 2025Updated 10 months ago
- EMNLP 2020: Filtering before Iteratively Referring for Knowledge-Grounded Response Selection in Retrieval-Based Chatbots☆12Dec 15, 2020Updated 5 years ago
- ☆16Sep 28, 2020Updated 6 years ago
- [ACL 2025] Are Your LLMs Capable of Stable Reasoning?☆33Aug 5, 2025Updated last year
- Subspace Representation Learning for Sparse Linear Arrays to Localize More Sources than Sensors: A Deep Learning Methodology☆21Mar 18, 2025Updated last year
- [ICDE 2024] VDTuner - Automated Performance Tuning for Vector Data Management Systems (Vector Databases)☆34Apr 21, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated 2 years ago
- ☆18Sep 10, 2025Updated last year
- ☆17Apr 27, 2026Updated 5 months ago
- get the media stream from Dahua/Haikang IPC SDK, and demux the stream to vedio and audio ES☆15Nov 15, 2015Updated 10 years ago
- ☆14Nov 12, 2024Updated last year
- Graph-based Document Structure Analysis☆19Aug 1, 2026Updated last month
- Dev and Test Data of LogicGame benchmark☆19Mar 31, 2025Updated last year
- Benchmarking LLM Inference Speeds☆14Sep 2, 2026Updated 3 weeks ago
- I don't want to maintain this project, the code probably won't compile or run. Archived.☆14Feb 25, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆22Aug 30, 2021Updated 5 years ago
- ☆19Nov 8, 2022Updated 3 years ago
- Joint learning of object and action detectors☆15Nov 5, 2019Updated 6 years ago
- Fast Memorization of Prompt Improves Context Awareness of Large Language Models (Findings of EMNLP 2024)☆22Oct 22, 2024Updated last year
- Codes for paper SoAy: A Service-oriented APIs Applying Framework of Large Language Models☆27Jul 14, 2025Updated last year
- ☆10Dec 10, 2024Updated last year
- Code accompanying our ICML 2020 paper on choice set optimization in group decision-making.☆10Jun 27, 2020Updated 6 years ago
- MATLAB code for the coarray tensor completion-based 2-D DOA estimation algorithm☆24Jul 3, 2024Updated 2 years ago
- Instruction Following Eval☆18Jan 16, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ACL 2024 | LooGLE: Long Context Evaluation for Long-Context Language Models☆200Aug 6, 2026Updated last month
- Original implementation of SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback (ICLR 2025)☆18Feb 17, 2025Updated last year
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- [ACL 2025 Main] Open-source toolkit for automatic evaluation of text-to-image generation task, including training & test datasets and a d…☆20Jul 5, 2025Updated last year
- The official site of the CVPR 2022 Affine Correspondences and Their Applications tutorial☆11Jan 17, 2023Updated 3 years ago
- Official repository for the paper Number Cookbook: Number Understanding of Language Models and How to Improve It.☆22Mar 31, 2025Updated last year
- 计算语言学22-23学年秋季学期 课程大作业baseline实现☆38Dec 8, 2022Updated 3 years ago