☆34May 14, 2025Updated last year
Alternatives and similar repositories for vitok
Users that are interested in vitok are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An official code release of the paper RGB no more: Minimally Decoded JPEG Vision Transformers☆58Jul 11, 2023Updated 3 years ago
- ☆13Nov 1, 2023Updated 2 years ago
- A big_vision inspired repo that implements a generic Auto-Encoder class capable in representation learning and generative modeling.☆34Jun 26, 2024Updated 2 years ago
- This repo contains the code for the paper "Object-cropping for SSL".☆18Feb 14, 2023Updated 3 years ago
- ☆24Jun 18, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- An implementation of several unsupervised object discovery models (Slot Attention, SLATE, GNM) in PyTorch with pre-trained models.☆14May 26, 2025Updated last year
- A basic pure pytorch implementation of flash attention☆17Oct 28, 2024Updated last year
- [NeurIPS 2025] Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations☆202Sep 18, 2025Updated 11 months ago
- [NeurIPS 2022] code for "Visual Concepts Tokenization"☆23Oct 10, 2022Updated 3 years ago
- ☆61Oct 29, 2022Updated 3 years ago
- This repo contains evaluation code for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?"☆31Dec 23, 2024Updated last year
- [ICML'25] EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling.☆184Mar 18, 2026Updated 5 months ago
- ☆43Jun 6, 2025Updated last year
- Experimental GPU language with meta-programming☆32Sep 6, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- (CVPR 2025) A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning☆25Mar 11, 2025Updated last year
- [ICCV2025]Code Release of Harmonizing Visual Representations for Unified Multimodal Understanding and Generation☆193May 21, 2025Updated last year
- Code for MetaMorph Multimodal Understanding and Generation via Instruction Tuning☆236Jan 22, 2026Updated 7 months ago
- [Arxiv'25] MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization☆55Sep 16, 2025Updated 11 months ago
- MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge (ICCV 2023)☆31Sep 5, 2023Updated 2 years ago
- Slot-TTA shows that test-time adaptation using slot-centric models can improve image segmentation on out-of-distribution examples.☆26Jun 20, 2023Updated 3 years ago
- ☆20Nov 23, 2022Updated 3 years ago
- ☆21Mar 25, 2025Updated last year
- [WACV 2026]Official Code of the paper “Equivariant Sampling for Improving Diffusion Model-based Image Restoration“☆19Jan 29, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- new optimizer☆20Aug 4, 2024Updated 2 years ago
- Official Pytorch implementation for LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior (ICLR 2025 Oral).☆107Feb 11, 2025Updated last year
- ☆32Jul 29, 2024Updated 2 years ago
- Code for "Scaling Language-Free Visual Representation Learning" paper (Web-SSL).☆216Mar 20, 2026Updated 5 months ago
- Official repository for "Solving Video Inverse Problems Using Image Diffusion Models"☆12Mar 7, 2026Updated 5 months ago
- Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters☆135Dec 3, 2024Updated last year
- ☆17Mar 2, 2023Updated 3 years ago
- Official Implementation of Paper Transfer between Modalities with MetaQueries☆326Oct 12, 2025Updated 10 months ago
- 2nd place solution of ECCV 2020 workshop VIPriors Image Classification Challenge, https://arxiv.org/abs/2008.00261☆13Aug 22, 2021Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆53Jan 18, 2024Updated 2 years ago
- Benchmarking Multi-Image Understanding in Vision and Language Models☆11Jul 29, 2024Updated 2 years ago
- Official PyTorch Implementation of "Diffusion Autoencoders are Scalable Image Tokenizers"☆170Jan 31, 2025Updated last year
- Official implementation of "Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation"☆66Mar 19, 2026Updated 5 months ago
- ☆28Dec 21, 2023Updated 2 years ago
- Pytorch implementation of Twelve Labs' Video Foundation Model evaluation framework & open embeddings☆36Aug 23, 2024Updated 2 years ago
- Official repo for From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models☆34Nov 2, 2025Updated 9 months ago