π This is a repository for organizing papers, codes and other resources related to visual tokenizers.
β18Jul 7, 2026Updated last month
Alternatives and similar repositories for Awesome-Visual-Tokenizers
Users that are interested in Awesome-Visual-Tokenizers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repo for the VideoVerseβ15Mar 29, 2026Updated 5 months ago
- Transformer: PyTorch Implementation of "Attention Is All You Need"β15Dec 13, 2023Updated 2 years ago
- Official Implementation for paper "Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm"β23May 8, 2026Updated 3 months ago
- β10Jun 30, 2026Updated 2 months ago
- [CVPR'25] MergeVQ: A Unified Framework for Visual Generation and Representation with Token Merging and Quantizationβ51Jul 22, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- β18May 18, 2026Updated 3 months ago
- β15Mar 22, 2026Updated 5 months ago
- Repo from the "Learning with limited labeled data" seminar @ Uni of Tuebingen. A collection of notes, notebooks and slideshows to understβ¦β17Apr 13, 2023Updated 3 years ago
- β29Jun 10, 2025Updated last year
- Code for DVD A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogueβ14Oct 12, 2021Updated 4 years ago
- [ICLR'26] SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion Transformersβ16Mar 19, 2026Updated 5 months ago
- [MICCAI 2026] A longitudinal, multimodal algorithm for multi-tumor segmentation (learning from reports).β15Updated this week
- The application of large pre-trained vision model DINOv2 from MetaAI for feature points matching, and a ViT decoder used for Auto Encoderβ18Apr 27, 2023Updated 3 years ago
- Synthetic VQA data generation code for SpatialReasoner.β21Nov 25, 2025Updated 9 months ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- One-Shot Learning for Pose-Guided Person Image Synthesis in the Wildβ20Apr 6, 2025Updated last year
- ε€§ε¦Latexη辩樑ηοΌε½εε ε«ε·ε€§γεε·₯ε€§γδΈη§ε€§γβ11Jul 22, 2024Updated 2 years ago
- PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokensβ18Nov 26, 2025Updated 9 months ago
- Grounding Language Models for Compositional and Spatial Reasoningβ18Oct 26, 2022Updated 3 years ago
- β28Nov 19, 2025Updated 9 months ago
- A toy text-to-image model trained from scratch.β20Jun 9, 2025Updated last year
- β28Oct 7, 2025Updated 10 months ago
- Med-DANet Series (ECCV 2022 & WACV 2024)β13Jan 2, 2024Updated 2 years ago
- Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"β197Feb 24, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code release for "Generative Modeling of Weights: Generalization or Memorization?"β23Apr 9, 2026Updated 4 months ago
- Official PyTorch implementation of the paper "Enhancing Vision-Language Pre-Training with Jointly Learned Questioner and Dense Captioner"β15Aug 9, 2023Updated 3 years ago
- PyTorch implementation of "PatchGame: Learning to Signal Mid-level Patches in Referential Games" to appear in NeurIPS 2021β24Jun 4, 2021Updated 5 years ago
- β78Jul 3, 2024Updated 2 years ago
- [ICLR 2026] Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasksβ32Feb 5, 2026Updated 6 months ago
- Official Implementation of "Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration"β20Jul 15, 2025Updated last year
- A research-friendly PyTorch Lightning toolkit for training, fine-tuning, and evaluating AutoencoderKL for Stable Diffusion and FLUX.β91Aug 18, 2026Updated last week
- official implementation of the paper "Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability".β78Dec 25, 2025Updated 8 months ago
- Self-Teaching Autoencoder learning reconstructions through latent agreement, not pixel loss.β30May 25, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- UniVesselSeg official repositoryβ18Jan 26, 2026Updated 7 months ago
- Official implementation of VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latentsβ15Feb 5, 2026Updated 6 months ago
- [NAACL 2025] Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMsβ25Sep 26, 2024Updated last year
- The official GitHub page for the survey paper "A Survey of RWKV".β33Jan 7, 2025Updated last year
- Paper Reading of IMCC groups.β19Oct 22, 2025Updated 10 months ago
- A small project that uses Discrete Denoising Diffusion Probabilistic Models (D3PMs), a generative model for discrete data that builds upoβ¦β18Aug 10, 2024Updated 2 years ago
- [ICML 2026]β18Jul 4, 2026Updated last month