LinxinS97/NLPBench

Readme badge preview -

If you own this repo, copy the snippet below and add it to your README.md

[![RelatedRepos](https://img.shields.io/badge/related-repos-yellow)](https://relatedrepos.com/gh/LinxinS97/NLPBench)

LinxinS97 / NLPBench

NLPBench: Evaluating NLP-Related Problem-solving Ability in Large Language Models

☆10

Alternatives and similar repositories for NLPBench

Users that are interested in NLPBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

JackHck / SBCL
View on GitHub
[ICCV 2023] Subclass-balancing contrastive learning for long-tailed recognition
☆18Oct 30, 2023Updated 2 years ago
JackHck / MADAug
View on GitHub
[ICCV 2023] MADAug: When to Learn What: Model-Adaptive Data Augmentation Curriculum
☆20Nov 9, 2023Updated 2 years ago
weihao1115 / MMLU-ProX
View on GitHub
[EMNLP 2025 Main] The official repo of MMLU-ProX benchmark.
☆29Aug 26, 2025Updated 10 months ago
limenlp / SEA
View on GitHub
Official Implementation for the paper "Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base"
☆27Sep 2, 2025Updated 10 months ago
JieyuZ2 / TaskMeAnything
View on GitHub
[NeurIPS 2024] A task generation and model evaluation system for multimodal language models.
☆71Nov 27, 2024Updated last year
GPU virtual machines on DigitalOcean Gradient AI • Ad
Get to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
divyakraman / AerialDiffusion
View on GitHub
Codebase for the paper Aerial Diffusion: Text Guided Ground-to-Aerial View Translation from a Single Image using Diffusion Models
☆13Oct 3, 2023Updated 2 years ago
kai-wen-yang / IDAA
View on GitHub
[ICML2022] "Identity-Disentangled Adversarial Augmentation for Self-Supervised Learning"
☆10Jul 24, 2022Updated 3 years ago
tianyi-lab / RuleR
View on GitHub
[NAACL'25] RuleR: Improving LLM Controllability by Rule-based Data Recycling
☆14Sep 27, 2025Updated 9 months ago
Yibin-Lei / MetaEOL
View on GitHub
Implementation for ACL 2024 paper "Meta-Task Prompting Elicits Embeddings from Large Language Models"
☆12Jul 25, 2024Updated last year
stefanhgm / patient_summaries_with_llms
View on GitHub
Code for "A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models"
☆17Jul 20, 2025Updated last year
nliulab / ShapleyVIC
View on GitHub
ShapleyVIC: Shapley Variable Importance Cloud for Interpretable Machine Learning
☆19Jun 6, 2024Updated 2 years ago
tianyi-lab / R2-T2
View on GitHub
[ICML 2025] Code for "R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts"
☆19Mar 10, 2025Updated last year
ruiyang-medinfo / KG-Rank
View on GitHub
KG-Rank: Enhancing Large Language Models for Medical QA with Knowledge Graphs and Ranking Techniques
☆50Dec 9, 2024Updated last year
tianyi-lab / C3PO
View on GitHub
[COLM 2025] "C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing"
☆21Apr 9, 2025Updated last year
AI Agents on DigitalOcean Gradient AI Platform • Ad
Build production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
limenlp / verl
View on GitHub
AdaRFT: Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
☆56Jun 13, 2025Updated last year
shubh1905 / balanced_kmeans
View on GitHub
Balanced K-means in Pytorch with strong GPU acceleration
☆12Apr 30, 2020Updated 6 years ago
tianyi-lab / RoMA
View on GitHub
Code for "Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs"
☆19Nov 6, 2025Updated 8 months ago
RAIVNLab / sugar-crepe
View on GitHub
[NeurIPS 2023] A faithful benchmark for vision-language compositionality
☆93Feb 13, 2024Updated 2 years ago
jamestszhim / adaptive_augment
View on GitHub
☆13Mar 14, 2022Updated 4 years ago
tianyi-lab / Moltbook_Socialization
View on GitHub
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
☆18Feb 17, 2026Updated 5 months ago
xirui-li / MOSSBench
View on GitHub
An implementation for MLLM oversensitivity evaluation
☆18Nov 16, 2024Updated last year
tianyi-lab / DisCL
View on GitHub
[ICCV 2025] Diffusion Curriculum (DisCL)
☆18Sep 26, 2025Updated 9 months ago
tianyi-lab / Mosaic-IT
View on GitHub
[ACL'25] Mosaic-IT: Cost-Free Compositional Data Synthesis for Instruction Tuning
☆20Sep 27, 2025Updated 9 months ago
Managed Kubernetes at scale on DigitalOcean • Ad
DigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
Charrrrrlie / X-as-Supervision
View on GitHub
The official repository of the paper "X as Supervision: Contending with Depth Ambiguity in Unsupervised Monocular 3D Pose Estimation"
☆13Jan 22, 2025Updated last year
IIT-PAVIS / SelfGeo
View on GitHub
Repository of the ECCV 2024 paper "SelfGeo: Self-supervised and Geodesic-consistent Estimation of Keypoints on Deformable Shapes"
☆15Sep 13, 2024Updated last year
JianqiangWan / VLPT-STD
View on GitHub
Vision-Language Pre-Training for Boosting Scene Text Detectors (CVPR2022)
☆12Mar 21, 2022Updated 4 years ago
YivanZhang / lio
View on GitHub
Learning from Indirect Observations
☆11Jul 16, 2021Updated 5 years ago
UCDvision / low-budget-al
View on GitHub
PyTorch implementation of "A Simple Baseline for Low-Budget Active Learning".
☆14Dec 22, 2021Updated 4 years ago
MingLiiii / Gradient_Unified
View on GitHub
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
☆20Jun 17, 2025Updated last year
spyysalo / s800
View on GitHub
Tools for working with the S800 corpus
☆12Sep 17, 2020Updated 5 years ago
wuxiyang1996 / AutoHallusion
View on GitHub
AutoHallusion Codebase (EMNLP 2024)
☆23Dec 6, 2024Updated last year
pedropaiola / ptt5-summ
View on GitHub
☆10Nov 30, 2022Updated 3 years ago
Managed hosting for WordPress and PHP on Cloudways • Ad
Managed hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
ritikamangla / QSalience
View on GitHub
https://arxiv.org/abs/2404.10917
☆14Mar 18, 2025Updated last year
TrustGen / TrustEval-toolkit
View on GitHub
[ICLR'26, NAACL'25 Demo] Toolkit & Benchmark for evaluating the trustworthiness of generative foundation models.
☆131Aug 22, 2025Updated 10 months ago
nsi319 / Legal-Summarizer
View on GitHub
Longformer Encoder Decoder model for the legal domain, trained for long document abstractive summarization task.
☆10Feb 26, 2021Updated 5 years ago
yannick-couzinie / expander
View on GitHub
A makeshift python program which relies on nltk and Stanford Core NLP models to expand common contractions in the english language.
☆10Nov 8, 2017Updated 8 years ago
stevenyangyj / CoTASP
View on GitHub
Official code for the paper: Continual Task Allocation in Meta-Policy Network via Sparse Prompting
☆23Feb 10, 2025Updated last year
Jasonlee1995 / Gumbel_Softmax
View on GitHub
Unofficial Pytorch implementation of the paper 'Categorical Reparameterization with Gumbel-Softmax' and 'The Concrete Distribution: A Con…
☆11Apr 27, 2021Updated 5 years ago
Sean-Blank / AMRcoref
View on GitHub
☆13Oct 4, 2022Updated 3 years ago