video description generation vision-language model
☆22Jan 21, 2025Updated last year
Alternatives and similar repositories for SpaceTimeGPT
Users that are interested in SpaceTimeGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Apr 9, 2026Updated 4 months ago
- An official codebase for "NormLens: Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Comm…☆10May 9, 2024Updated 2 years ago
- ☆11Mar 28, 2024Updated 2 years ago
- ImageNet3D: Towards General-Purpose Object-Level 3D Understanding☆22Dec 6, 2024Updated last year
- This is the official resources for ECCV 2022 paper "Object Level Depth Reconstruction for Category Level 6D Object Pose Estimation From M…☆18Jun 15, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Project for SNARE benchmark☆11Jun 5, 2024Updated 2 years ago
- A Qwen .5B reasoning model trained on OpenR1-Math-220k☆14Updated this week
- ☆11Oct 25, 2020Updated 5 years ago
- Codebase for ACL 2023 paper "Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models' Memori…☆51Oct 8, 2023Updated 2 years ago
- ☆10Jun 19, 2023Updated 3 years ago
- Code for paper "W-RAG: Weakly Supervised Dense Retrieval in RAG for Open-domain Question Answering"☆16Oct 2, 2025Updated 10 months ago
- [CVPR 2025] UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References☆32Apr 28, 2025Updated last year
- ☆13Jul 23, 2024Updated 2 years ago
- ☆12Apr 6, 2021Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- RGB-based Category-level Object Pose Estimation via Decoupled Metric Scale Recovery☆30Apr 16, 2024Updated 2 years ago
- The correct way to resize images or tensors. For Numpy or Pytorch (differentiable).☆18May 5, 2022Updated 4 years ago
- [EMNLP 2024 Industry track] MERLIN : Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank P…☆14Mar 4, 2025Updated last year
- The CVF Open Access Downloader is a Python application designed to automate the bulk downloading of open-access papers from Computer Visi…☆11May 8, 2024Updated 2 years ago
- ☆16Oct 19, 2022Updated 3 years ago
- Doing style transfer with linguistic features using OpenAI's CLIP.☆14May 4, 2021Updated 5 years ago
- code for downloading videos from HowTo100M dataset☆18May 13, 2021Updated 5 years ago
- Official Implementation of "GMOS: Grounding Moving Object Segmentation in 3D Space and Time". Junyu Xie, Tengda Han, Weidi Xie, Andrew Zi…☆38May 29, 2026Updated 3 months ago
- Code for the paper: "Invertible CNN-Based Super Resolution with Downsampling Awareness" by Andrew Geiss and Joseph C. Hardin, Nov 2020☆12Nov 11, 2020Updated 5 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A curve-editor for Stable Diffusion prompt interpolation☆21Oct 3, 2022Updated 3 years ago
- ☆14Apr 30, 2023Updated 3 years ago
- ☆13Jun 23, 2019Updated 7 years ago
- [EMNLP 2024] TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering☆18Oct 31, 2024Updated last year
- quagga☆10Apr 7, 2020Updated 6 years ago
- ☆14Mar 11, 2024Updated 2 years ago
- RefTeacher is a strong baseline method for Semi-Supervised Referring Expression Comprehension.☆14May 26, 2023Updated 3 years ago
- RAG application to answer questions about PDF documents using LLMs.☆16Dec 1, 2023Updated 2 years ago
- Custom triton kernels for training Karpathy's nanoGPT.☆19Oct 21, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆46May 24, 2025Updated last year
- ☆74Aug 8, 2025Updated last year
- ☆15Aug 4, 2025Updated last year
- ☆15Dec 12, 2024Updated last year
- [ICCV'23] UATVR: Uncertainty-Adaptive Text-Video Retrieval☆13Nov 5, 2023Updated 2 years ago
- Official code for the paper: "Metadata Archaeology"☆19May 10, 2023Updated 3 years ago
- ☆13Jan 3, 2024Updated 2 years ago