PyTorch Implementation of Object Recognition as Next Token Prediction [CVPR'24 Highlight]
☆180May 1, 2025Updated last year
Alternatives and similar repositories for nxtp
Users that are interested in nxtp are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ECCV 2024] WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation☆114Feb 6, 2025Updated last year
- This repo contains the code for our paper Towards Open-Ended Visual Recognition with Large Language Model☆102Jul 15, 2024Updated 2 years ago
- Code of the paper "Efficient Object Detection in Autonomous Driving using Spiking Neural Networks: Performance, Energy Consumption Analys…☆27Dec 13, 2023Updated 2 years ago
- (CVPR2023) CAPE: Camera View Position Embedding for Multi-View 3D Object Detection☆110May 5, 2023Updated 3 years ago
- ☆35Jan 23, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- DALI Multi Agent System Framework☆43Mar 24, 2026Updated 4 months ago
- ☆27Aug 28, 2023Updated 2 years ago
- Adapting LLaMA Decoder to Vision Transformer☆30May 20, 2024Updated 2 years ago
- Code for this paper "HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts via HyperNetwork"☆33Nov 29, 2023Updated 2 years ago
- [ECCV 2024] official code for "Long-CLIP: Unlocking the Long-Text Capability of CLIP"☆901Aug 13, 2024Updated last year
- (ICCV 2023) Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation☆48Jul 18, 2024Updated 2 years ago
- [CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses tha…☆966Aug 5, 2025Updated last year
- [ICLR 2024 (Spotlight)] "Frozen Transformers in Language Models are Effective Visual Encoder Layers"☆245Jun 29, 2026Updated last month
- NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024☆1,851Nov 27, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official implementation of "Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models"☆38Jan 3, 2024Updated 2 years ago
- High-performance Image Tokenizers for VAR and AR☆307Apr 25, 2025Updated last year
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"☆56Jul 5, 2025Updated last year
- Code release for "Language-conditioned Detection Transformer"☆86Jun 17, 2024Updated 2 years ago
- A detection/segmentation dataset with labels characterized by intricate and flexible expressions. "Described Object Detection: Liberating…☆138Mar 20, 2024Updated 2 years ago
- (ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest☆555Jun 3, 2025Updated last year
- [ECCV2024 Oral🔥] Official Implementation of "GiT: Towards Generalist Vision Transformer through Universal Language Interface"☆364Jan 14, 2025Updated last year
- AeDet: Azimuth-invariant Multi-view 3D Object Detection, CVPR2023☆75Jun 17, 2023Updated 3 years ago
- [NeurIPS 2024] Official implementation of the paper "Interfacing Foundation Models' Embeddings"☆132Aug 21, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [NeurIPS 2023] Official implementations of "Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models"☆523Jan 27, 2024Updated 2 years ago
- Code For Our Work: DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries [ECCV-2024]☆15Jul 11, 2024Updated 2 years ago
- ☆648Feb 15, 2024Updated 2 years ago
- ImageNet-12k subset of ImageNet-21k (fall11)☆23Jun 13, 2023Updated 3 years ago
- [CVPR 2024] Aligning and Prompting Everything All at Once for Universal Visual Perception☆609May 8, 2024Updated 2 years ago
- ☆38Feb 8, 2024Updated 2 years ago
- State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!☆2,338Apr 13, 2026Updated 3 months ago
- Official Pytorch Implementation of Self-emerging Token Labeling☆34Mar 27, 2024Updated 2 years ago
- [CVPR 2024] Official implementation of the paper "Visual In-context Learning"☆543Apr 8, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [ECCV'24] Official PyTorch implementation of In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation☆51Sep 24, 2024Updated last year
- This repo contains the code for 1D tokenizer and generator☆1,172Mar 20, 2025Updated last year
- Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks (NeurIPS2022)☆85Nov 2, 2022Updated 3 years ago
- ☆121Jun 11, 2024Updated 2 years ago
- This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.☆1,424Aug 4, 2025Updated last year
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations☆22Dec 24, 2025Updated 7 months ago
- ☆110Jun 30, 2023Updated 3 years ago