Real-time YOLOv5 + Intel RealSense D435 pipeline for depth-aware object detection and 3D coordinate extraction, enabling precise robotic arm grasping.
☆16Nov 29, 2025Updated 10 months ago
Alternatives and similar repositories for Real-Time-Object-Detection-with-Depth
Users that are interested in Real-Time-Object-Detection-with-Depth are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [INDIN25] Multi-granularity guided Vision Transformer for efficient 2D human pose estimation, combining refined-SCConv features with ViT-…☆16Nov 29, 2025Updated 10 months ago
- Intelligent video-learning platform that analyzes viewing behavior to generate personalized questions, track errors, and enable social, f…☆19Dec 15, 2025Updated 9 months ago
- [ECCV2026] ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images☆32Aug 25, 2026Updated last month
- ☆49May 22, 2026Updated 4 months ago
- The PyTorch implementation of Cautious-Adam.☆18Jun 25, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Lightweight, Configuration-Driven, Flexible Fine-Tuning Framework for 🤗 Diffusers☆16Apr 15, 2026Updated 5 months ago
- CST-former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection (ICASSP 2024)☆40May 20, 2025Updated last year
- [ICML 2026] Official repository for the paper "Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention"☆53Aug 5, 2026Updated 2 months ago
- ☆51Nov 8, 2025Updated 11 months ago
- The official implementation of the paper "Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models" (NeurIPS 2025 Pos…☆80Sep 29, 2025Updated last year
- Minute-long video generation at 24FPS.☆70Mar 28, 2026Updated 6 months ago
- This repository aims to collect Transformer-based sound event detection (SED) algorithms.☆106Feb 10, 2026Updated 7 months ago
- ☆157May 13, 2025Updated last year
- [ICLR 2025] Official implementation and benchmark evaluation repository of <PhysBench: Benchmarking and Enhancing Vision-Language Models …☆94Jan 21, 2026Updated 8 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Awesome things about generative recommendation models.☆117Apr 28, 2025Updated last year
- Reading list for research topics in Sound AI☆202Aug 8, 2024Updated 2 years ago
- A comprehensive benchmark specifically designed to evaluate the interactive response capabilities of world models in 4D settings.☆108Mar 24, 2026Updated 6 months ago
- [ICML 2026] d3LLM: Ultra-Fast Diffusion LLM 🚀☆156May 1, 2026Updated 5 months ago
- 4-steps distilled version of Wan2.2-TI2V-5B☆174Mar 15, 2026Updated 6 months ago
- [NeurIPS 2025] Official PyTorch implementation of paper "CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up".☆219Sep 27, 2025Updated last year
- Interactive World Model papers organized by core research challenges.☆310Sep 14, 2026Updated 3 weeks ago
- Diffusion Recommender Model☆256Sep 9, 2026Updated last month
- [ICLR'25] Official code for the paper 'MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs'☆390Apr 20, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model☆2,323Mar 19, 2026Updated 6 months ago
- ☆255Nov 19, 2025Updated 10 months ago
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention☆551Feb 24, 2026Updated 7 months ago
- [CVPR 2026] [Best Paper Finalist] [Oral] Official repository of Vision Test-Time Training☆394Updated this week
- [ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.☆383Aug 30, 2026Updated last month
- Dynamic 3D Foundation Model using Causal Transformer. [ICLR 2026]☆410May 8, 2026Updated 5 months ago
- [ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers☆521Dec 6, 2025Updated 10 months ago
- DreamX-World: A General-Purpose Interactive World Model☆778Jul 23, 2026Updated 2 months ago
- [ICML2025, NeurIPS2025 Spotlight] Sparse VideoGen 1 & 2: Accelerating Video Diffusion Transformers with Sparse Attention☆709Jul 4, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models☆850Sep 10, 2026Updated 3 weeks ago
- [NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation☆788Apr 16, 2026Updated 5 months ago
- [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactiv…☆997Updated this week
- The official implementation for [NeurIPS2025 Oral] Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink…☆985Dec 20, 2025Updated 9 months ago
- 📚 数千篇 AI、LLM、NLP、CV 顶会论文解读,每篇 5 分钟读懂核心思想。☆2,092Updated this week
- Master the Toolkit of AI and Machine Learning. Mathematics for Machine Learning and Data Science is a beginner-friendly Specialization wh…☆887Jul 10, 2023Updated 3 years ago
- SoulX-FlashTalk is the first 14B model to achieve sub-second start-up latency (0.87s) while maintaining a real-time throughput of 32 FPS …☆1,534Jul 30, 2026Updated 2 months ago