π This is a repository for organizing papers, codes and other resources related to visual tokenizers.
β19Sep 27, 2026Updated last week
Alternatives and similar repositories for Awesome-Visual-Tokenizers
Users that are interested in Awesome-Visual-Tokenizers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repo for the VideoVerseβ15Mar 29, 2026Updated 6 months ago
- Transformer: PyTorch Implementation of "Attention Is All You Need"β15Dec 13, 2023Updated 2 years ago
- Official Implementation for paper "Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm"β24May 8, 2026Updated 5 months ago
- A controllable and interactive simulation framework for vision research.β17May 25, 2026Updated 4 months ago
- β10Jun 30, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [CVPR'25] MergeVQ: A Unified Framework for Visual Generation and Representation with Token Merging and Quantizationβ51Jul 22, 2025Updated last year
- Repo from the "Learning with limited labeled data" seminar @ Uni of Tuebingen. A collection of notes, notebooks and slideshows to understβ¦β17Apr 13, 2023Updated 3 years ago
- Provably (and non-vacuously) bounding test error of deep neural networks under distribution shift with unlabeled test data.β10Feb 27, 2024Updated 2 years ago
- PyTorch implementation of "PatchVAE: Learning Local Latent Codes for Recognition" to appear in CVPR 2020β14Apr 9, 2020Updated 6 years ago
- β17Mar 22, 2026Updated 6 months ago
- Code for DVD A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogueβ14Oct 12, 2021Updated 4 years ago
- The application of large pre-trained vision model DINOv2 from MetaAI for feature points matching, and a ViT decoder used for Auto Encoderβ18Apr 27, 2023Updated 3 years ago
- [NeurIPS 2025] Official repo of "Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions"β21Aug 6, 2025Updated last year
- One-Shot Learning for Pose-Guided Person Image Synthesis in the Wildβ20Apr 6, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Collect papers and codes about VQGAN in various Computer Vision tasksβ10Dec 20, 2022Updated 3 years ago
- PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokensβ18Nov 26, 2025Updated 10 months ago
- Grounding Language Models for Compositional and Spatial Reasoningβ17Oct 26, 2022Updated 3 years ago
- β28Nov 19, 2025Updated 10 months ago
- A toy text-to-image model trained from scratch.β20Jun 9, 2025Updated last year
- Med-DANet Series (ECCV 2022 & WACV 2024)β13Jan 2, 2024Updated 2 years ago
- Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"β200Feb 24, 2026Updated 7 months ago
- Code release for "Generative Modeling of Weights: Generalization or Memorization?"β23Apr 9, 2026Updated 6 months ago
- [ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasonerβ69May 29, 2026Updated 4 months ago
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Official PyTorch implementation of the paper "Enhancing Vision-Language Pre-Training with Jointly Learned Questioner and Dense Captioner"β15Aug 9, 2023Updated 3 years ago
- β22Aug 23, 2026Updated last month
- β13Oct 12, 2020Updated 5 years ago
- PyTorch implementation of "PatchGame: Learning to Signal Mid-level Patches in Referential Games" to appear in NeurIPS 2021β24Jun 4, 2021Updated 5 years ago
- Official code for the MICCAI 2025 paper "Semantically Consistent Discrete Diffusion for 3D Biological Graph Generation"β19Jul 7, 2025Updated last year
- β78Jul 3, 2024Updated 2 years ago
- [ICLR 2026] Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasksβ32Feb 5, 2026Updated 8 months ago
- Official Implementation of "Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration"β20Jul 15, 2025Updated last year
- A research-friendly PyTorch Lightning toolkit for training, fine-tuning, and benchmarking AutoencoderKL for Stable Diffusion, FLUX, and bβ¦β92Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- official implementation of the paper "Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability".β84Dec 25, 2025Updated 9 months ago
- Self-Teaching Autoencoder learning reconstructions through latent agreement, not pixel loss.β31May 25, 2026Updated 4 months ago
- Official implementation of VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latentsβ15Feb 5, 2026Updated 8 months ago
- [NAACL 2025] Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMsβ25Sep 26, 2024Updated 2 years ago
- [CVPR-2026] DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformersβ23Sep 24, 2026Updated 2 weeks ago
- Paper Reading of IMCC groups.β19Oct 22, 2025Updated 11 months ago
- Existing literature about training-data analysis.β17Dec 17, 2021Updated 4 years ago