π This is a repository for organizing papers, codes and other resources related to visual tokenizers.
β18Jul 7, 2026Updated 2 months ago
Alternatives and similar repositories for Awesome-Visual-Tokenizers
Users that are interested in Awesome-Visual-Tokenizers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Transformer: PyTorch Implementation of "Attention Is All You Need"β15Dec 13, 2023Updated 2 years ago
- Official Implementation for paper "Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm"β23May 8, 2026Updated 4 months ago
- β10Jun 30, 2026Updated 2 months ago
- [CVPR'25] MergeVQ: A Unified Framework for Visual Generation and Representation with Token Merging and Quantizationβ51Jul 22, 2025Updated last year
- Repo from the "Learning with limited labeled data" seminar @ Uni of Tuebingen. A collection of notes, notebooks and slideshows to understβ¦β17Apr 13, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β29Jun 10, 2025Updated last year
- Provably (and non-vacuously) bounding test error of deep neural networks under distribution shift with unlabeled test data.β10Feb 27, 2024Updated 2 years ago
- PyTorch implementation of "PatchVAE: Learning Local Latent Codes for Recognition" to appear in CVPR 2020β14Apr 9, 2020Updated 6 years ago
- β16Mar 22, 2026Updated 5 months ago
- Code for DVD A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogueβ14Oct 12, 2021Updated 4 years ago
- [MICCAI 2026] A longitudinal, multimodal algorithm for multi-tumor segmentation (learning from reports).β15Aug 25, 2026Updated 3 weeks ago
- The application of large pre-trained vision model DINOv2 from MetaAI for feature points matching, and a ViT decoder used for Auto Encoderβ18Apr 27, 2023Updated 3 years ago
- Synthetic VQA data generation code for SpatialReasoner.β21Nov 25, 2025Updated 9 months ago
- One-Shot Learning for Pose-Guided Person Image Synthesis in the Wildβ20Apr 6, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ε€§ε¦Latexη辩樑ηοΌε½εε ε«ε·ε€§γεε·₯ε€§γδΈη§ε€§γβ11Jul 22, 2024Updated 2 years ago
- Collect papers and codes about VQGAN in various Computer Vision tasksβ10Dec 20, 2022Updated 3 years ago
- PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokensβ18Nov 26, 2025Updated 9 months ago
- Grounding Language Models for Compositional and Spatial Reasoningβ18Oct 26, 2022Updated 3 years ago
- β28Oct 7, 2025Updated 11 months ago
- Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"β200Feb 24, 2026Updated 6 months ago
- Code release for "Generative Modeling of Weights: Generalization or Memorization?"β23Apr 9, 2026Updated 5 months ago
- [ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasonerβ67May 29, 2026Updated 3 months ago
- Official PyTorch implementation of the paper "Enhancing Vision-Language Pre-Training with Jointly Learned Questioner and Dense Captioner"β15Aug 9, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β18Aug 23, 2026Updated 3 weeks ago
- [NeurIPS 2024] Official repository for "Physics-Regularized Multi-Modal Image Assimilation for Brain Tumor Localization"β22Jan 6, 2025Updated last year
- β13Oct 12, 2020Updated 5 years ago
- PyTorch implementation of "PatchGame: Learning to Signal Mid-level Patches in Referential Games" to appear in NeurIPS 2021β24Jun 4, 2021Updated 5 years ago
- β78Jul 3, 2024Updated 2 years ago
- [ICLR 2026] Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasksβ32Feb 5, 2026Updated 7 months ago
- Official Implementation of "Latent Posterior-Mean Rectified Flow for Higher-Fidelity Perceptual Face Restoration"β20Jul 15, 2025Updated last year
- A research-friendly PyTorch Lightning toolkit for training, fine-tuning, and benchmarking AutoencoderKL for Stable Diffusion, FLUX, and bβ¦β92Updated this week
- official implementation of the paper "Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability".β83Dec 25, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation of VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latentsβ15Feb 5, 2026Updated 7 months ago
- UniVesselSeg official repositoryβ18Jan 26, 2026Updated 7 months ago
- [CVPR-2026] DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformersβ23Updated this week
- A small project that uses Discrete Denoising Diffusion Probabilistic Models (D3PMs), a generative model for discrete data that builds upoβ¦β18Aug 10, 2024Updated 2 years ago
- Existing literature about training-data analysis.β17Dec 17, 2021Updated 4 years ago
- [ICML 2026]β18Jul 4, 2026Updated 2 months ago
- This repo holds the official code for the paper "FreMIM: Fourier Transform Meets Masked Image Modeling for Medical Image Segmentation".β24Jan 2, 2024Updated 2 years ago