Implementation of the deepmind Flamingo vision-language model, based on Hugging Face language models and ready for training
β171Apr 27, 2023Updated 3 years ago
Alternatives and similar repositories for flamingo-mini
Users that are interested in flamingo-mini are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of 𦩠Flamingo, state-of-the-art few-shot visual question answering attention net out of Deepmind, in Pytorchβ1,269Oct 18, 2022Updated 3 years ago
- An open-source framework for training large multimodal models.β4,119Aug 31, 2024Updated last year
- β11Nov 21, 2024Updated last year
- MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.β954Mar 19, 2025Updated last year
- SimVLM ---SIMPLE VISUAL LANGUAGE MODEL PRETRAINING WITH WEAK SUPERVISIONβ36Nov 7, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- π§ Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs".β484Oct 30, 2023Updated 2 years ago
- Repository for the paper "Data Efficient Masked Language Modeling for Vision and Language".β18Sep 17, 2021Updated 4 years ago
- SVIT: Scaling up Visual Instruction Tuningβ168Jun 20, 2024Updated 2 years ago
- This repository contains code and dataset splits for the paper "Classification by Attention: Scene Graph Classification with Prior Knowleβ¦β16May 27, 2022Updated 4 years ago
- DataComp: In search of the next generation of multimodal datasetsβ787Apr 28, 2025Updated last year
- A reimplementation of KOSMOS-1 from "Language Is Not All You Need: Aligning Perception with Language Models"β27Mar 3, 2023Updated 3 years ago
- 𦦠Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing impβ¦β3,433Mar 5, 2024Updated 2 years ago
- β134Dec 22, 2023Updated 2 years ago
- Implementation of LaTr: Layout-aware transformer for scene-text VQA,a novel multimodal architecture for Scene Text Visual Question Answerβ¦β56Jul 22, 2026Updated 3 weeks ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Official implementation of SEED-LLaMA (ICLR 2024).β642Sep 21, 2024Updated last year
- Chain of Images for Intuitively Reasoningβ10Nov 29, 2023Updated 2 years ago
- This code provides a PyTorch implementation for OTTER (Optimal Transport distillation for Efficient zero-shot Recognition), as described β¦β71Dec 20, 2021Updated 4 years ago
- β25Jun 5, 2023Updated 3 years ago
- β203May 10, 2023Updated 3 years ago
- [ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parametersβ5,913Mar 14, 2024Updated 2 years ago
- Official code for the ICLR2023 paper Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detectionβ43Jun 4, 2024Updated 2 years ago
- GRiT: A Generative Region-to-text Transformer for Object Understanding (ECCV2024)β342Jan 8, 2024Updated 2 years ago
- Deep Learning for Video Retrieval by Natural Languageβ11Oct 20, 2019Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- GIT: A Generative Image-to-text Transformer for Vision and Languageβ582Dec 2, 2023Updated 2 years ago
- COYO-700M: Large-scale Image-Text Pair Datasetβ1,255Nov 30, 2022Updated 3 years ago
- [ICLR 2024] Analyzing and Mitigating Object Hallucination in Large Vision-Language Modelsβ159Apr 30, 2024Updated 2 years ago
- [ICCV 2021 Oral + TPAMI] Just Ask: Learning to Answer Questions from Millions of Narrated Videosβ127Sep 29, 2023Updated 2 years ago
- [CVPR 2023] Learning Visual Representations via Language-Guided Samplingβ150Apr 13, 2023Updated 3 years ago
- Code used for the creation of OBELICS, an open, massive and curated collection of interleaved image-text web documents, containing 141M dβ¦β217Aug 28, 2024Updated last year
- A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models!β141Dec 31, 2023Updated 2 years ago
- Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Lβ¦β2,559Apr 24, 2024Updated 2 years ago
- An open source implementation of CLIP.β14,075Aug 10, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACL 2024 Findings & ICLR 2024 WS] An Evaluator VLM that is open-source, offers reproducible evaluation, and inexpensive to use. Specificβ¦β86Sep 13, 2024Updated last year
- β231Dec 18, 2023Updated 2 years ago
- β20May 30, 2024Updated 2 years ago
- Teacher - student distillation using DeepSpeedβ20Oct 7, 2022Updated 3 years ago
- β12Jun 29, 2024Updated 2 years ago
- [ICML2023] Instant Soup Cheap Pruning Ensembles in A Single Pass Can Draw Lottery Tickets from Large Models. Ajay Jaiswal, Shiwei Liu, Tiβ¦β11Nov 28, 2023Updated 2 years ago
- [ICML 2023] "Robust Weight Signatures: Gaining Robustness as Easy as Patching Weights?" by Ruisi Cai, Zhenyu Zhang, Zhangyang Wangβ16May 4, 2023Updated 3 years ago