Reproduction of the first step in the text-to-video model Phenaki. Code and model weights for the Transformer-based autoencoder for videos called CViViT.
☆29Aug 4, 2023Updated 3 years ago
Alternatives and similar repositories for phenaki-cvivit
Users that are interested in phenaki-cvivit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- finetune script for SDXL adapted from waifu-diffusion trainer☆11Aug 21, 2023Updated 3 years ago
- Implementation of Phenaki Video, which uses Mask GIT to produce text guided videos of up to 2 minutes in length, in Pytorch☆789Jul 29, 2024Updated 2 years ago
- Implementation of MagViT2 Tokenizer in Pytorch☆670Jan 12, 2025Updated last year
- This repo consist of some experimental results on bdd100k datasets using different object detection algorithms(Faster-RCNN, FCOS, ATSS)☆11Jun 27, 2020Updated 6 years ago
- Pytorch implementation of deep fill v2 (original by Jiayu et al.)☆10Jun 26, 2019Updated 7 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Unofficial implement of "Pix2seq: A Language Modeling Framework for Object Detection" on mmdetection☆34Apr 18, 2022Updated 4 years ago
- Official implementation of MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis☆86Jul 16, 2024Updated 2 years ago
- ☆56Oct 16, 2023Updated 2 years ago
- [ACL2026 Findings] GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning☆84Jun 23, 2025Updated last year
- [NAACL 2024] Part-based, explainable and editable fine-grained image classifier that allows users to define a species in text☆14Sep 19, 2025Updated last year
- ☆12Oct 12, 2020Updated 5 years ago
- SEED-Voken: A Series of Powerful Visual Tokenizers☆1,022Nov 25, 2025Updated 10 months ago
- This is an OCR program designed for travel document. It can now support 23 types of documents with pre-defined template. You can add what…☆10Nov 22, 2022Updated 3 years ago
- Style Transfer by Deep Learning, overview and TensorFlow implementations (UNDER CONSTRUCTION)☆14Jul 25, 2017Updated 9 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆90Jan 4, 2024Updated 2 years ago
- PyTorch Implementation of MobileDet (https://arxiv.org/abs/2004.14525v3) backbones.☆11Feb 12, 2024Updated 2 years ago
- Google MobileNets Implementation using Tensorflow☆18Jun 6, 2017Updated 9 years ago
- [IJCAI 2022 poster] PyTorch Implementation of "Universal Video Style Transfer via Crystallization, Separation, and Blending"☆17Mar 10, 2023Updated 3 years ago
- A script for spawning VSCode Remote server sessions on the TUoS HPC clusters.☆15Dec 12, 2024Updated last year
- A PyTorch re-implementation of Weakly Supervised Facial Action Unit Recognition through Adversarial Training☆10Apr 23, 2019Updated 7 years ago
- Annotated Tutorial for PerAct☆19Sep 11, 2023Updated 3 years ago
- Unofficial Pytorch Implementation of "A Simple Framework for Contrastive Learning of Visual Representations"☆10Mar 11, 2020Updated 6 years ago
- Official Pytorch code for "AesUST: Towards Aesthetic-Enhanced Universal Style Transfer" (ACM MM 2022)☆15Dec 31, 2022Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- EgoToM is an egocentric theory-of-mind benchmark built on Ego4D videos, containing multi-choice questions that evaluate multimodal large …☆18Apr 1, 2025Updated last year
- Multi-temporal Scene dataset for Scene Change Detection.☆15Apr 14, 2021Updated 5 years ago
- Framework to achieve context distillation in LLMs☆15Nov 24, 2023Updated 2 years ago
- [NeurIPS 2024] Data exporter for SS3DM: Benchmarking Street-View Surface Reconstruction with a Synthetic 3D Mesh Dataset☆16Nov 8, 2024Updated last year
- Official JAX implementation of MAGVIT: Masked Generative Video Transformer☆1,002Jan 17, 2024Updated 2 years ago
- Toolkit for VIPER benchmark☆17Aug 11, 2020Updated 6 years ago
- ☆10Jan 20, 2021Updated 5 years ago
- A list of robotics related papers accepted by ICLR'25☆27Aug 28, 2025Updated last year
- [ICCV2025] TokenBridge: Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation. https://yuqingwang1029.github.io/To…☆162Jul 24, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- This repo contains evaluation code for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?"☆31Dec 23, 2024Updated last year
- running LayoutLMv2☆11Apr 27, 2022Updated 4 years ago
- The official codebase for running the experiments described in the AVDC paper.☆21Oct 2, 2024Updated 2 years ago
- Repository for the paper CenterPoly: real-time instance segmentation using bounding polygons☆51Aug 24, 2021Updated 5 years ago
- 简书小工具集 - 探索未知☆11Jan 31, 2025Updated last year
- auto remove image backgound service by SpringBoot3+ JDK17+ AI☆15Mar 22, 2024Updated 2 years ago
- A python wrapper for ScriptHookV☆11Jun 26, 2018Updated 8 years ago