PULSE-EVAL
☆24Jan 12, 2024Updated 2 years ago
Alternatives and similar repositories for PULSE-EVAL
Users that are interested in PULSE-EVAL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆28Aug 2, 2023Updated 3 years ago
- Code for ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Context☆17Nov 15, 2024Updated last year
- PULSE: Pretrained and Unified Language Service Engine☆498Dec 26, 2023Updated 2 years ago
- Counting-Stars (★)☆84Nov 24, 2025Updated 9 months ago
- ☆10Dec 28, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official implementation of OpenTab (ICLR2024)☆14Mar 27, 2024Updated 2 years ago
- We systematically studied the influencing factors when LLM generates benchmarks,By using our code, you can generate high-quality QA datas…☆20May 20, 2025Updated last year
- ☆16Aug 23, 2023Updated 3 years ago
- [Nature Communications] The official code for "Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases".☆73Nov 7, 2025Updated 10 months ago
- ☆10Oct 11, 2022Updated 3 years ago
- CMB, A Comprehensive Medical Benchmark in Chinese☆251Mar 27, 2025Updated last year
- LogicBench is a natural language question-answering dataset consisting of 25 different reasoning patterns spanning over propositional, fi…☆40May 2, 2024Updated 2 years ago
- ☆11Nov 8, 2023Updated 2 years ago
- Very simple and crude HTTP server written in C☆14Jul 29, 2018Updated 8 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Chain of Images for Intuitively Reasoning☆10Nov 29, 2023Updated 2 years ago
- ☆30Nov 5, 2024Updated last year
- The official code of TACL 2022, "Break, Perturb, Build: Automatic Perturbation of Reasoning Paths Through Question Decomposition".☆12Oct 18, 2021Updated 4 years ago
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated last year
- Code for "Demonstration-free Autonomous Reinforcement Learning via Implicit and Bidirectional Curriculum" (ICML 2023)☆10Jul 6, 2023Updated 3 years ago
- Evaluation Pipeline for medical tasks.☆11Apr 8, 2026Updated 4 months ago
- ☆10Nov 16, 2023Updated 2 years ago
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- PromptCBLUE: a large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain in Chinese☆393Jan 23, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- get the media stream from Dahua/Haikang IPC SDK, and demux the stream to vedio and audio ES☆15Nov 15, 2015Updated 10 years ago
- ☆16Jan 15, 2021Updated 5 years ago
- The official data and code for EMNLP 2023 main conference paper: CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular D…☆13May 19, 2025Updated last year
- Bayesian low-rank adaptation for large language models☆29May 4, 2024Updated 2 years ago
- ☆11Jul 31, 2024Updated 2 years ago
- 基于LLM实现CHIP2021-Task3中文临床术语标准化任务,准确率约70%。☆16Dec 16, 2024Updated last year
- MedLSAM: Localize and Segment Anything Model for 3D Medical Images☆524Apr 30, 2024Updated 2 years ago
- A Chinese National Medical Licensing Examination dataset and large languge model benchmarks☆97Dec 2, 2023Updated 2 years ago
- LLM+RAG for QA☆23Jan 15, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The code for our NeurIPS 2021 paper "Kernelized Heterogeneous Risk Minimization".☆13Oct 13, 2021Updated 4 years ago
- Code and notebooks and data for the paper "Domain Specific Question Answering Over Knowledge Graphs Using Logical Programming and Large L…☆12Jan 23, 2024Updated 2 years ago
- ☆33Feb 9, 2025Updated last year
- 大语言模型微调的项目,包含了使用QLora微调ChatGLM和LLama☆29Jun 26, 2023Updated 3 years ago
- I don't want to maintain this project, the code probably won't compile or run. Archived.☆14Feb 25, 2024Updated 2 years ago
- ☆19Feb 20, 2025Updated last year
- MathEval is a benchmark dedicated to the holistic evaluation on mathematical capacities of LLMs.☆87Nov 15, 2024Updated last year