[ACL 2025 🔥] A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
☆79Aug 10, 2026Updated 3 weeks ago
Alternatives and similar repositories for KITAB-Bench
Users that are interested in KITAB-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AIN - The First Arabic Inclusive Large Multimodal Model. It is a versatile bilingual LMM excelling in visual and contextual understanding…☆54Mar 13, 2025Updated last year
- [NAACL 2025 🔥] CAMEL-Bench is an Arabic benchmark for evaluating multimodal models across eight domains with 29,000 questions.☆38Apr 17, 2025Updated last year
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"☆57Jul 5, 2025Updated last year
- A tool for improving the output of generic Arabic OCR systems using an n-gram based post-correction approach.☆10Sep 22, 2021Updated 4 years ago
- (ICCV 2023) Generative Multiplane Neural Radiance for 3D Aware Image Generation.☆18Sep 28, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- repo for paper titled: Towards Realistic Zero-Shot Classification via Self Structural Semantic Alignment (AAAI'24 Oral)☆25May 16, 2024Updated 2 years ago
- [CVPR 2025 🔥]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos☆105Apr 14, 2025Updated last year
- [NAACL'25] Contains code and documentation for our VANE-Bench paper.☆24Aug 19, 2025Updated last year
- A new multi-task learning framework using Vision Transformers☆11Jun 19, 2024Updated 2 years ago
- This is the official repository for Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks.☆26Dec 9, 2024Updated last year
- Code, models, and data for "Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation". EMNLP 2023.☆21Aug 29, 2024Updated 2 years ago
- ☆11Oct 29, 2024Updated last year
- [ACL 2026 🔥] CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark☆37Aug 17, 2026Updated 2 weeks ago
- Reasoning DriveLMM☆16Mar 15, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code release for "Gaze-Assisted Medical Image Segmentation" [AIM-FM @ NeurIPS, 2024]☆14Oct 22, 2024Updated last year
- [ICLR 2026 🔥] Dr.LLM: Dynamic Layer Routing in LLMs☆57Apr 24, 2026Updated 4 months ago
- Bio-Medical EXpert LMM with English and Arabic Language Capabilities☆74Oct 23, 2025Updated 10 months ago
- أسئلة باللغة العربية تركز على الثقافة السعودية تم اختبارها على عدد من النماذج اللغوية الضخمة LLMs☆18Jan 22, 2025Updated last year
- [MICCAI 2024 🔥] HLSS, the first study to explore hierarchical information inherent in histopathology images and their language descripti…☆27Aug 5, 2024Updated 2 years ago
- ☆10Mar 8, 2025Updated last year
- [CVPRW-25 MMFM] Official repository of paper titled "How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite fo…☆50Aug 23, 2024Updated 2 years ago
- Reinforcement Training of Robot☆11Dec 1, 2019Updated 6 years ago
- Official repository for "Boosting Adversarial Transferability using Dynamic Cues " (ICLR 2023)☆20Aug 24, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- VideoMathQA is a benchmark designed to evaluate mathematical reasoning in real-world educational videos☆24May 7, 2026Updated 3 months ago
- Open Source tool for Arabic text readability☆25Jul 4, 2025Updated last year
- Composed Video Retrieval☆62May 2, 2024Updated 2 years ago
- ☆15Feb 26, 2024Updated 2 years ago
- Arabic Words☆12Nov 18, 2019Updated 6 years ago
- UICrit is a dataset containing human-generated natural language design critiques, corresponding bounding boxes for each critique, and des…☆27Nov 19, 2024Updated last year
- Maximize Efficiency, Elevate Accuracy: Slash GPU Hours by Half with Efficient Pre-training!☆72Mar 31, 2026Updated 5 months ago
- [CVPR -2025] GroupMamba: Parameter-Efficient and Accurate Group Visual State Space Model☆141Mar 22, 2025Updated last year
- [CVPRW 2025] Official repository of paper titled "Towards Evaluating the Robustness of Visual State Space Models"☆25Jun 8, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2025] Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding☆17Jun 16, 2025Updated last year
- [BMVC 2025] Official Implementation of the paper "PerSense: Personalized Instance Segmentation in Dense Images"☆31Dec 18, 2025Updated 8 months ago
- [ICCVW 2025 (Oral)] Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models☆29Oct 20, 2025Updated 10 months ago
- Abstract. Person search is a challenging problem with various real- world applications, that aims at joint person detection and re-identi…☆13Feb 28, 2024Updated 2 years ago
- [CVPR 2026 🔥] Time Blindness: Why Video-Language Models Can't See What Humans Can?☆68Jan 28, 2026Updated 7 months ago
- Official repository for "Stylized Adversarial Training" (TPAMI 2022)☆11Dec 30, 2022Updated 3 years ago
- Official code repository of paper titled "Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Visio…☆33May 11, 2025Updated last year