☆93May 20, 2025Updated last year
Alternatives and similar repositories for textvqa_grounding_task_qwen2.5-vl-ft
Users that are interested in textvqa_grounding_task_qwen2.5-vl-ft are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆34Mar 2, 2025Updated last year
- This code calculates Abs Real Error of the struct2depth and Depth from video in the wild model. (Success)☆10Feb 25, 2021Updated 5 years ago
- EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO☆29Jan 24, 2026Updated 7 months ago
- ☆27Jul 13, 2026Updated last month
- ☆13Jan 3, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [NeurIPS-W 2025] Official Implementation of "Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning"☆73Jul 1, 2025Updated last year
- ☆89Aug 13, 2025Updated last year
- Domain Adaptation with Adversarial Training on Penultimate Activations (AAAI 2023)☆11Aug 1, 2023Updated 3 years ago
- Official code for SA-Solver: Stochastic Adams Solver for Fast Sampling of Diffusion Models (NeurIPS 2023)☆15Mar 4, 2024Updated 2 years ago
- Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost Volume☆19Nov 18, 2019Updated 6 years ago
- Pytorch implementation of 'Improving Self-supervised Lightweight Model Learning via Hard-aware Metric Distillation. In ECCV 2022'☆11Mar 22, 2023Updated 3 years ago
- Implements PyTorch model which updates SPD weights on Riemannian Manifold. Based on Huang, Z., & Van Gool, L. (2016). A Riemannian Netwo…☆12Mar 8, 2019Updated 7 years ago
- ☆16Mar 26, 2025Updated last year
- Add YOLOv3_tiny and data augment(clip, brighten, change saturation)☆14Jan 14, 2021Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.☆1,962Aug 22, 2026Updated 2 weeks ago
- This demo demonstrates the AI capabilities of the mcxn947. It displays the image captured by the camera on the LCD screen and performs fa…☆14May 18, 2026Updated 3 months ago
- 使用opencv部署yolo11表格检测,它是百度网盘AI大赛-表格检测的第2名方案,方案里包含表格框检测,表格角点检测,表格方向分类,一共三个模块。我依然是编写了C++和Python两个版本的程序☆13Dec 12, 2024Updated last year
- Code for paper "JMDC: A Joint Model and Data Compression System for Deep Neural Networks Collaborative Computing in Edge-Cloud Networks"☆26Jul 22, 2026Updated last month
- Multiple-Person Multi-Camera Tracker☆13Feb 17, 2017Updated 9 years ago
- One-Shot Unsupervised Cross Domain Detection☆13Nov 22, 2022Updated 3 years ago
- dancetrack 比赛第二名☆13Jan 29, 2023Updated 3 years ago
- ☆11Nov 8, 2022Updated 3 years ago
- RobuQ: Pushing DiTS to W1.58A2 via Robust Activation Quantization☆17Jun 28, 2026Updated 2 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆15Apr 25, 2025Updated last year
- ☆28Oct 31, 2024Updated last year
- [ACCV 2024 (Oral, Best Application Paper)] Official Implementation of NT-VOT211: A Large-Scale Benchmark for Night-time Visual Object Tra…☆16Dec 30, 2025Updated 8 months ago
- Official code for 'Transformers in Unsupervised Structure-from-Motion' and 'Transformers in Self-Supervised Monocular Depth Estimation wi…☆14Nov 12, 2023Updated 2 years ago
- [IEEE TIFS under review] TOPIC: IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision Transformer☆13Apr 9, 2024Updated 2 years ago
- ☆16May 4, 2026Updated 4 months ago
- 3D LUTs for Real Time sRGB White-Balance Correction☆14Dec 14, 2023Updated 2 years ago
- ☆13Aug 27, 2021Updated 5 years ago
- Monocular depth estimation from a single image☆10Jun 29, 2019Updated 7 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆12Mar 28, 2024Updated 2 years ago
- A Large-Scale Blind Image Quality Assessment Database☆17Jul 18, 2023Updated 3 years ago
- Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.☆19,900Jan 30, 2026Updated 7 months ago
- Source code for AAAI 2024 paper "Finding Visual Saliency in Continuous Spike Stream"☆14Aug 21, 2025Updated last year
- This code runs the experiments for 2D Compressed Sensing based Encryption-Then-Compression (2DCS-ETC) scheme.☆13Aug 8, 2020Updated 6 years ago
- Image style transfer- change any image to an artistic image by using Convolutional Neural Network.☆11Jan 30, 2020Updated 6 years ago
- Fully Open Framework for Democratized Multimodal Training☆1,197Updated this week