PyTorch Implementation of Object Recognition as Next Token Prediction [CVPR'24 Highlight]
☆180May 1, 2025Updated last year
Alternatives and similar repositories for nxtp
Users that are interested in nxtp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ECCV 2024] WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation☆114Feb 6, 2025Updated last year
- This repo contains the code for our paper Towards Open-Ended Visual Recognition with Large Language Model☆102Jul 15, 2024Updated 2 years ago
- Code of the paper "Efficient Object Detection in Autonomous Driving using Spiking Neural Networks: Performance, Energy Consumption Analys…☆27Dec 13, 2023Updated 2 years ago
- (CVPR2023) CAPE: Camera View Position Embedding for Multi-View 3D Object Detection☆110May 5, 2023Updated 3 years ago
- ☆35Jan 23, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆27Aug 28, 2023Updated 2 years ago
- Adapting LLaMA Decoder to Vision Transformer☆30May 20, 2024Updated 2 years ago
- Code for this paper "HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts via HyperNetwork"☆33Nov 29, 2023Updated 2 years ago
- [ECCV 2024] official code for "Long-CLIP: Unlocking the Long-Text Capability of CLIP"☆900Aug 13, 2024Updated last year
- (ICCV 2023) Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation☆48Jul 18, 2024Updated 2 years ago
- [CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses tha…☆963Aug 5, 2025Updated 11 months ago
- [ICLR 2024 (Spotlight)] "Frozen Transformers in Language Models are Effective Visual Encoder Layers"☆244Jun 29, 2026Updated 3 weeks ago
- NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024☆1,846Nov 27, 2025Updated 7 months ago
- Official implementation of "Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models"☆38Jan 3, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"☆56Jul 5, 2025Updated last year
- High-performance Image Tokenizers for VAR and AR☆307Apr 25, 2025Updated last year
- Code release for "Language-conditioned Detection Transformer"☆86Jun 17, 2024Updated 2 years ago
- ☆12May 26, 2022Updated 4 years ago
- A detection/segmentation dataset with labels characterized by intricate and flexible expressions. "Described Object Detection: Liberating…☆138Mar 20, 2024Updated 2 years ago
- [ICLR2025] Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want☆94Dec 1, 2025Updated 7 months ago
- (ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest☆556Jun 3, 2025Updated last year
- [ECCV2024 Oral🔥] Official Implementation of "GiT: Towards Generalist Vision Transformer through Universal Language Interface"☆364Jan 14, 2025Updated last year
- AeDet: Azimuth-invariant Multi-view 3D Object Detection, CVPR2023☆75Jun 17, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [NeurIPS 2024] Official implementation of the paper "Interfacing Foundation Models' Embeddings"☆132Aug 21, 2024Updated last year
- [NeurIPS 2023] Official implementations of "Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models"☆522Jan 27, 2024Updated 2 years ago
- Code For Our Work: DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries [ECCV-2024]☆15Jul 11, 2024Updated 2 years ago
- ☆650Feb 15, 2024Updated 2 years ago
- ImageNet-12k subset of ImageNet-21k (fall11)☆23Jun 13, 2023Updated 3 years ago
- [CVPR 2024] Aligning and Prompting Everything All at Once for Universal Visual Perception☆609May 8, 2024Updated 2 years ago
- ☆38Feb 8, 2024Updated 2 years ago
- State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!☆2,324Apr 13, 2026Updated 3 months ago
- ☆34Apr 11, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official Pytorch Implementation of Self-emerging Token Labeling☆35Mar 27, 2024Updated 2 years ago
- [CVPR 2024] Official implementation of the paper "Visual In-context Learning"☆542Apr 8, 2024Updated 2 years ago
- [ECCV'24] Official PyTorch implementation of In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation☆51Sep 24, 2024Updated last year
- This repo contains the code for 1D tokenizer and generator☆1,165Mar 20, 2025Updated last year
- Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks (NeurIPS2022)☆85Nov 2, 2022Updated 3 years ago
- DALI Multi Agent System Framework☆43Mar 24, 2026Updated 3 months ago
- ☆121Jun 11, 2024Updated 2 years ago