Real-time YOLOv5 + Intel RealSense D435 pipeline for depth-aware object detection and 3D coordinate extraction, enabling precise robotic arm grasping.
☆17Nov 29, 2025Updated 8 months ago
Alternatives and similar repositories for Real-Time-Object-Detection-with-Depth
Users that are interested in Real-Time-Object-Detection-with-Depth are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Intelligent video-learning platform that analyzes viewing behavior to generate personalized questions, track errors, and enable social, f…☆20Dec 15, 2025Updated 8 months ago
- VisionZip-enhanced LLaDA-V for DLM inference, compressing visual tokens for faster, plug-and-play vision-aware reasoning with minimal qua…☆16Nov 28, 2025Updated 8 months ago
- [ECCV2026] FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation☆60Dec 2, 2025Updated 8 months ago
- Onset-and-Offset-Aware Sound Event Detection☆21Feb 10, 2025Updated last year
- ☆17May 19, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Sound Event Detection (SED) paper collection☆15Jun 26, 2024Updated 2 years ago
- A Lightweight, Configuration-Driven, Flexible Fine-Tuning Framework for 🤗 Diffusers☆17Apr 15, 2026Updated 4 months ago
- ☆24Feb 1, 2026Updated 6 months ago
- CST-former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection (ICASSP 2024)☆40May 20, 2025Updated last year
- [ICML 2026] Official repository for the paper "Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention"☆47Aug 5, 2026Updated 2 weeks ago
- ☆49Nov 8, 2025Updated 9 months ago
- Minute-long video generation at 24FPS.☆69Mar 28, 2026Updated 4 months ago
- [NeurIPS 2025] More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models☆82May 31, 2025Updated last year
- This repository aims to collect Transformer-based sound event detection (SED) algorithms.☆105Feb 10, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 把毕业的大师兄蒸馏成一个能继续开组会、骂醒你、顺手救火的 AI Skill。☆86Apr 10, 2026Updated 4 months ago
- ☆146May 13, 2025Updated last year
- ☆91Mar 6, 2026Updated 5 months ago
- [ICLR 2025] Official implementation and benchmark evaluation repository of <PhysBench: Benchmarking and Enhancing Vision-Language Models …☆94Jan 21, 2026Updated 6 months ago
- A Minimalist, Batteries-included Repository for Advancing World Model Science.☆713Jun 15, 2026Updated 2 months ago
- [ICML 2026] d3LLM: Ultra-Fast Diffusion LLM 🚀☆149May 1, 2026Updated 3 months ago
- 4-steps distilled version of Wan2.2-TI2V-5B☆168Mar 15, 2026Updated 5 months ago
- [ICML 2026] Pytorch implementation of Self-Refining Video Sampling☆185May 1, 2026Updated 3 months ago
- ☆226Dec 9, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Interactive World Model papers organized by core research challenges.☆289Aug 4, 2026Updated 2 weeks ago
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention☆331Feb 24, 2026Updated 5 months ago
- [ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.☆381Mar 29, 2026Updated 4 months ago
- Dynamic 3D Foundation Model using Causal Transformer. [ICLR 2026]☆397May 8, 2026Updated 3 months ago
- DreamX-World: A General-Purpose Interactive World Model☆760Jul 23, 2026Updated 3 weeks ago
- A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models☆785Jun 15, 2026Updated 2 months ago
- rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale☆782Jun 25, 2026Updated last month
- [NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation☆777Apr 16, 2026Updated 4 months ago
- [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactiv…☆927Jul 23, 2026Updated 3 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.☆1,030Feb 25, 2026Updated 5 months ago
- MOVA: Towards Scalable and Synchronized Video–Audio Generation☆1,098Updated this week
- HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency☆1,578Jun 10, 2026Updated 2 months ago
- JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation☆1,858Jun 26, 2026Updated last month
- A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.☆2,477Aug 10, 2026Updated last week
- A platform for reproducible world model research and evaluation☆2,133Updated this week
- mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding☆2,409May 30, 2025Updated last year