URS Benchmark: Evaluating LLMs on User Reported Scenarios
☆31May 30, 2025Updated last year
Alternatives and similar repositories for URS
Users that are interested in URS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-throughput and memory-efficient inference and serving engine for LLMs☆17Jun 3, 2024Updated 2 years ago
- code and data associated with CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations☆11Oct 13, 2023Updated 2 years ago
- An extended project of the LLM Compiler paper, focusing on developing LLM-based Autonomous Agents.☆26Oct 22, 2024Updated last year
- Code to compute AnthroScore, a computational linguistic measure of anthropomorphism in text☆19Mar 31, 2025Updated last year
- The backup repository for FairytaleQA dataset and paper "Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset f…☆10May 30, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆11Sep 19, 2025Updated last year
- This repo is to demo the concept of lossless compression with Transformers as encoder and decoder.☆14May 2, 2024Updated 2 years ago
- 언어모델을 학습하기 위한 공개 한국어 instruction dataset들을 모아두었습니다.☆19Jul 16, 2023Updated 3 years ago
- Code of paper: Probing the Difficulty Perception Mechanism of Large Language Models☆19Mar 17, 2026Updated 6 months ago
- Baseline system for Language-based Audio Retrieval (Task 6B) in DCASE 2023 Challenge☆10Aug 8, 2023Updated 3 years ago
- Code for "Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning".☆28Nov 11, 2025Updated 10 months ago
- Automatically Generated d2l-zh TensorFlow Notebooks for Colab☆12Aug 18, 2023Updated 3 years ago
- Code and data for "KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark" (LREC-COLING…☆18Apr 15, 2025Updated last year
- ☆13Feb 11, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Code for NeurIPS 2024 Spotlight: "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations"☆94Oct 30, 2024Updated last year
- Official repo for "SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization"☆29Mar 24, 2026Updated 5 months ago
- ☆16Dec 3, 2024Updated last year
- KLUE Benchmark 1st place (2021.12) solutions. (RE, MRC, NLI, STS, TC)☆25Apr 11, 2022Updated 4 years ago
- Official implementation of "Disentangled Knowledge Transfer for OOD Intent Discovery with Unified Contrastive Learning", ACL2022 main con…☆14Jul 23, 2022Updated 4 years ago
- StrategyQA 데이터 세트 번역☆22Apr 12, 2024Updated 2 years ago
- Official code and dataset repository of KoBBQ (TACL 2024)☆23May 13, 2024Updated 2 years ago
- Suite of 500 procedurally-generated NLP tasks to study language model adaptability☆21Jul 16, 2022Updated 4 years ago
- Code for RATIONALYST: Pre-training Process-Supervision for Improving Reasoning https://arxiv.org/pdf/2410.01044☆36Oct 3, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Pre-training Multi-task Contrastive Learning Models for Scientific Literature Understanding (Findings of EMNLP'23)☆11Aug 24, 2024Updated 2 years ago
- Tools for the evaluation of audio captioning.☆19May 23, 2020Updated 6 years ago
- Official code repository for Findings of EMNLP 2022 paper: PseudoReasoner: Leveraging Pseudo Labels for Commonsense Knowledge Base Popula…☆11Oct 18, 2022Updated 3 years ago
- [ACL 2025 Findings] Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts (https://huggingface.co/papers…☆92Nov 23, 2025Updated 9 months ago
- BERT finetuned on NER downstream tasks☆15Jun 12, 2023Updated 3 years ago
- ☆98Mar 20, 2024Updated 2 years ago
- ☆116Mar 12, 2024Updated 2 years ago
- Implementation of self-certainty as an extention of ZeroEval Project☆38May 31, 2025Updated last year
- Optimizing bit-level Jaccard Index and Population Counts for large-scale quantized Vector Search via Harley-Seal CSA and Lookup Tables☆22May 18, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆16Jun 5, 2020Updated 6 years ago
- ☆14Dec 9, 2021Updated 4 years ago
- The code of CIKM 2023 (Oral Presentation) : A Multi-Task Semantic Decomposition Framework with Task-specific Pre-training for Few-Shot NE…☆14Jul 19, 2024Updated 2 years ago
- mcp wrapper for openai built-in tools☆12Mar 13, 2025Updated last year
- Repo for ACL2023 paper "Won't Get Fooled Again: Answering Questions with False Premises"☆23Jun 11, 2023Updated 3 years ago
- [EMNLP 2023] Question Answering as Programming for Solving Time-Sensitive Questions☆12Dec 18, 2023Updated 2 years ago
- Official implementation of the paper "BRUCE: Bundle Recommendation Using Contextualized item Embeddings"☆16Sep 1, 2022Updated 4 years ago