☆93May 20, 2025Updated last year
Alternatives and similar repositories for textvqa_grounding_task_qwen2.5-vl-ft
Users that are interested in textvqa_grounding_task_qwen2.5-vl-ft are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO☆27Jan 24, 2026Updated 6 months ago
- ☆13Jan 3, 2024Updated 2 years ago
- Domain Adaptation with Adversarial Training on Penultimate Activations (AAAI 2023)☆11Aug 1, 2023Updated 3 years ago
- Implements PyTorch model which updates SPD weights on Riemannian Manifold. Based on Huang, Z., & Van Gool, L. (2016). A Riemannian Netwo…☆12Mar 8, 2019Updated 7 years ago
- Add YOLOv3_tiny and data augment(clip, brighten, change saturation)☆14Jan 14, 2021Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.☆1,958Jul 25, 2026Updated 3 weeks ago
- Comprehensive benchmark for video text understanding☆29Jun 4, 2025Updated last year
- ☆28Oct 31, 2024Updated last year
- The implementation code and models of TASTR.☆10Jul 14, 2020Updated 6 years ago
- [ACCV 2024 (Oral, Best Application Paper)] Official Implementation of NT-VOT211: A Large-Scale Benchmark for Night-time Visual Object Tra…☆16Dec 30, 2025Updated 7 months ago
- ☆16May 4, 2026Updated 3 months ago
- A Large-Scale Blind Image Quality Assessment Database☆17Jul 18, 2023Updated 3 years ago
- [CVPR2024] Parameter Efficient Fine-tuning via Cross Block Orchestration for Segment Anything Model☆12Jul 31, 2024Updated 2 years ago
- Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.☆19,790Jan 30, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [Neurips 24 Spotlight] Training in Pairs + Inference on Single Image with Anchors☆52Feb 20, 2025Updated last year
- 基于LLaVA1.6微调的Xray识别的多模态大模型☆10Oct 22, 2024Updated last year
- Code and Model For SAA☆12Sep 21, 2023Updated 2 years ago
- [ECCV2024] ModTr: Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledge☆20Nov 28, 2024Updated last year
- Code for ICCV 2023 work "Generalized Few-Shot Point Cloud Segmentation Via Geometric Words"☆14Sep 26, 2023Updated 2 years ago
- Codes of RailTrack_Segmentation of TITS2021-Enhanced Few-Shot Learning for Intrusion Detection in Railway Video Surveillance☆16Jun 23, 2024Updated 2 years ago
- A general binary code optimization method☆13Oct 9, 2016Updated 9 years ago
- Fully Open Framework for Democratized Multimodal Training☆1,174Updated this week
- unofficial implementation of https://arxiv.org/pdf/2301.08871v1.pdf on pytorch☆15Apr 20, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Traffic sign recognition☆14Dec 16, 2022Updated 3 years ago
- ☆45Apr 11, 2023Updated 3 years ago
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing [ICLR 2026]☆158Jul 26, 2026Updated 3 weeks ago
- [NeurIPS 2025] Official code implementation of Perception R1: Pioneering Perception Policy with Reinforcement Learning☆290Jul 15, 2025Updated last year
- ☆12Jan 9, 2025Updated last year
- [CVPR 2026] ActivityForensics: ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos☆23Apr 27, 2026Updated 3 months ago
- A Framework for Symbolic MUsic Graph Explanations☆11Jul 30, 2025Updated last year
- This is the official implementation of paper: Landmark Localization from Medical Images with Generative Distribution Prior☆14Mar 4, 2024Updated 2 years ago
- ☆12Jun 1, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Feature extraction from speech signals based on representation learning strategies using pre-trained autoencoders☆19Jul 6, 2023Updated 3 years ago
- Non-disruptive collagen characterization in clinical histopathology using cross-modality image synthesis☆11Apr 25, 2025Updated last year
- Official implementation of "AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison"☆59Apr 25, 2026Updated 3 months ago
- Project for SNARE benchmark☆11Jun 5, 2024Updated 2 years ago
- TIER: Text-Image Encoder-based Regression for AIGC Image Quality Assessment☆10Mar 1, 2025Updated last year
- ☆15Aug 3, 2019Updated 7 years ago
- ☆11Nov 3, 2021Updated 4 years ago