Reproduction of the first step in the text-to-video model Phenaki. Code and model weights for the Transformer-based autoencoder for videos called CViViT.
☆29Aug 4, 2023Updated 3 years ago
Alternatives and similar repositories for phenaki-cvivit
Users that are interested in phenaki-cvivit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Dec 14, 2024Updated last year
- SFT+RL boosts multimodal reasoning☆47Jun 27, 2025Updated last year
- Region Proposal generation on images using clustering in Pointcloud - Currently only for Pedestrians☆11Jul 13, 2020Updated 6 years ago
- Tutorial on how to create metrics dashboards like the THOR Dashboard☆14Mar 8, 2017Updated 9 years ago
- Implementation of MagViT2 Tokenizer in Pytorch☆668Jan 12, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository contains tools for visualization of keypoint matches over two images (ORB, SIFT, LIFT, SuperPoint, D2-Net).☆13Jul 23, 2019Updated 7 years ago
- Pseudo Labelling on MNIST dataset in Tensorflow 2.x☆10Jul 12, 2022Updated 4 years ago
- ☆16May 19, 2023Updated 3 years ago
- ☆10Jan 7, 2021Updated 5 years ago
- This repo consist of some experimental results on bdd100k datasets using different object detection algorithms(Faster-RCNN, FCOS, ATSS)☆11Jun 27, 2020Updated 6 years ago
- Code Guided Neural Style Transfer for Shape Stylization.☆11Jan 12, 2026Updated 6 months ago
- Pytorch implementation of deep fill v2 (original by Jiayu et al.)☆10Jun 26, 2019Updated 7 years ago
- Unofficial implement of "Pix2seq: A Language Modeling Framework for Object Detection" on mmdetection☆34Apr 18, 2022Updated 4 years ago
- ☆132Feb 22, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- SiT: Self-supervised vision Transformer☆21Apr 9, 2021Updated 5 years ago
- Official implementation of MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis☆86Jul 16, 2024Updated 2 years ago
- This is a toolbox repository to help evaluate various methods that perform image matching from a pair of images.☆12Jul 5, 2023Updated 3 years ago
- [ACL2026 Findings] GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning☆83Jun 23, 2025Updated last year
- [NAACL 2024] Part-based, explainable and editable fine-grained image classifier that allows users to define a species in text☆14Sep 19, 2025Updated 10 months ago
- ☆12Oct 12, 2020Updated 5 years ago
- SEED-Voken: A Series of Powerful Visual Tokenizers☆1,020Nov 25, 2025Updated 8 months ago
- This is an OCR program designed for travel document. It can now support 23 types of documents with pre-defined template. You can add what…☆10Nov 22, 2022Updated 3 years ago
- ☆89Jan 4, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- PyTorch Implementation of MobileDet (https://arxiv.org/abs/2004.14525v3) backbones.☆11Feb 12, 2024Updated 2 years ago
- Google MobileNets Implementation using Tensorflow☆18Jun 6, 2017Updated 9 years ago
- [IJCAI 2022 poster] PyTorch Implementation of "Universal Video Style Transfer via Crystallization, Separation, and Blending"☆17Mar 10, 2023Updated 3 years ago
- ☆38Dec 25, 2025Updated 7 months ago
- A script for spawning VSCode Remote server sessions on the TUoS HPC clusters.☆15Dec 12, 2024Updated last year
- ☆41Sep 21, 2023Updated 2 years ago
- Unofficial Pytorch Implementation of "A Simple Framework for Contrastive Learning of Visual Representations"☆10Mar 11, 2020Updated 6 years ago
- EgoToM is an egocentric theory-of-mind benchmark built on Ego4D videos, containing multi-choice questions that evaluate multimodal large …☆16Apr 1, 2025Updated last year
- Multi-temporal Scene dataset for Scene Change Detection.☆15Apr 14, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official JAX implementation of MAGVIT: Masked Generative Video Transformer☆1,002Jan 17, 2024Updated 2 years ago
- Toolkit for VIPER benchmark☆16Aug 11, 2020Updated 6 years ago
- ☆10Jan 20, 2021Updated 5 years ago
- Train a tiny LLaMA model from scratch to repeat your words using Reinforcement Learning from Human Feedback (RLHF)☆18May 23, 2024Updated 2 years ago
- A list of robotics related papers accepted by ICLR'25☆26Aug 28, 2025Updated 11 months ago
- ☆28Feb 7, 2024Updated 2 years ago
- [ICCV W] Contextual Convolutional Neural Networks (https://arxiv.org/pdf/2108.07387.pdf)☆14Aug 18, 2021Updated 4 years ago