☆15Dec 29, 2025Updated 8 months ago
Alternatives and similar repositories for agi-eval
Users that are interested in agi-eval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A iterative feedback driven benchmark on LLM's instruction following ability☆59May 25, 2026Updated 3 months ago
- CATArena is an engineering-level tournament evaluation platform for Large Language Model-driven code agents (LLM-driven code agents), bas…☆68Dec 25, 2025Updated 8 months ago
- ☆11Nov 16, 2019Updated 6 years ago
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation☆229Updated this week
- SCUT 本科毕业论文 Word 格式化模板☆10May 12, 2015Updated 11 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR'26] R-HORIZON: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?☆26May 9, 2026Updated 3 months ago
- Feedreader is a Go package for parsing RSS 2.0 and Atom 1.0 feed.☆20Aug 17, 2015Updated 11 years ago
- IPA files available for download from ONEJailbreak.com☆14Apr 30, 2025Updated last year
- A zero-shot faithfulness evaluation metric for text summarization☆11Oct 17, 2023Updated 2 years ago
- workflows for alfred4☆14Nov 16, 2020Updated 5 years ago
- reviese pyrouge files for supporting winxp win 8.1 win10☆12Nov 21, 2017Updated 8 years ago
- Mutual attention model for matching QA pairs in dialogues☆11Sep 20, 2020Updated 5 years ago
- Commonsense Knowledge Base Reasoning☆10Sep 3, 2018Updated 8 years ago
- Original Implementation of Improving Domain-Adapted Sentiment Classification by Deep Adversarial Mutual Learning publicized in AAAI-2020☆19May 29, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A benchmark on visual perception in text strings for both LLMs and MLLMs.☆16Apr 7, 2026Updated 4 months ago
- ☆42Apr 7, 2026Updated 4 months ago
- Attention_CopyNet☆29Aug 18, 2016Updated 10 years ago
- [Neurips 2022] “ Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified Backpropogation”, Ziyu Jiang*, Xuxi Chen*, Xueqin Huan…☆19Mar 14, 2023Updated 3 years ago
- A comparison of pretraining framework for LLM☆22Feb 6, 2025Updated last year
- [NAACL 2022] "Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training", Yuanxin Liu, Fandong Meng, Zheng Lin, Pe…☆15Oct 18, 2022Updated 3 years ago
- 🐂🍺位运算技巧一览表☆30Oct 5, 2018Updated 7 years ago
- ☆20Dec 16, 2020Updated 5 years ago
- Tensorflow-LSTM-CRF tool for Named Entity Recognizer☆59Jul 23, 2017Updated 9 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Fluency ENhanced Sentence-bert Evaluation (FENSE), metric for audio caption evaluation. And Benchmark dataset AudioCaps-Eval, Clotho-Eval…☆21Feb 1, 2023Updated 3 years ago
- [EMNLP 2023]Context Compression for Auto-regressive Transformers with Sentinel Tokens☆25Nov 6, 2023Updated 2 years ago
- Code for ACL2022 publication Transkimmer: Transformer Learns to Layer-wise Skim☆22Aug 21, 2022Updated 4 years ago
- ISE (Iris Server Engine) is a C++ framework for server programming.☆36Apr 7, 2014Updated 12 years ago
- BiLSTM-CRF for sequence labeling in Dynet☆82Jun 15, 2017Updated 9 years ago
- Implementation of the cw2vec model☆29Jul 20, 2018Updated 8 years ago
- Baseline models, training scripts, and instructions on how to reproduce our results for our state-of-art grammar correction system from M…☆73May 27, 2019Updated 7 years ago
- ☆24Jun 13, 2023Updated 3 years ago
- [EMNLP'24] LongHeads: Multi-Head Attention is Secretly a Long Context Processor☆32Apr 8, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Multi-turn response selection using dialogue dependency relations☆24Sep 1, 2021Updated 5 years ago
- A flagship 560-billion-parameter open-source MoE model that advances Native Formal Reasoning in Lean4 for Mathematics Formalization and P…☆94May 9, 2026Updated 3 months ago
- [AAAI 2021]Knowledge-Driven Distractor Generation for Cloze-Style Multiple Choice Questions☆22Jul 29, 2021Updated 5 years ago
- Source code for paper on commonsense reasoning for 2020 Annual Conference of the Association for Computational Linguistics (ACL) 2020.☆29Aug 2, 2024Updated 2 years ago
- Code for the paper "Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question Answering" (AAAI 2021)☆30Feb 19, 2021Updated 5 years ago
- Using tidb-lite to create a TiDB server with mocktikv mode in your application or unit test.☆54Jan 14, 2022Updated 4 years ago
- [COLM'24] Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration☆31Oct 18, 2024Updated last year