Code for "Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models"
☆20Feb 16, 2026Updated 7 months ago
Alternatives and similar repositories for linear-mech-vlms
Users that are interested in linear-mech-vlms are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Mar 5, 2024Updated 2 years ago
- [CVPR 2025 Highlight] Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding☆86Aug 31, 2025Updated last year
- ☆14Apr 10, 2025Updated last year
- HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models☆16Mar 6, 2025Updated last year
- Masked Autoencoders for Unsupervised Anomaly Detection in Medical Images☆21Aug 15, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository contains the code used for the experiments in the paper "Language Models use Lookbacks to Track Beliefs".☆17Mar 14, 2026Updated 6 months ago
- The official code and model for ACL 2023 paper 'mCLIP: Multilingual CLIP via Cross-lingual Transfer'☆10Jan 23, 2024Updated 2 years ago
- up-to-date curated list of state-of-the-art Large vision language models hallucinations research work, papers & resources☆331Aug 28, 2026Updated 3 weeks ago
- [原理解析] 大模型基本功(手撕Transformer模型、手撕PPO、GRPO、DPO训练器)☆33Jul 8, 2025Updated last year
- 🚲 Code and benchmark for our COLM 2025 paper - "Thought Tracing: Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models"☆15Aug 8, 2025Updated last year
- (NeurIPS 2025) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation☆78May 21, 2026Updated 4 months ago
- ☆87Nov 5, 2024Updated last year
- ☆10Nov 18, 2024Updated last year
- Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation☆15Aug 11, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [CVPR 2025] Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding☆17Oct 4, 2025Updated 11 months ago
- [NeurIPS 2025] Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs☆37Sep 21, 2025Updated last year
- Mamba-Spike——CGI2024☆14Dec 3, 2025Updated 9 months ago
- Lyrics Generator based on GPT-2☆10Jun 20, 2023Updated 3 years ago
- [EMNLP 2025 Main] Official implementation of VRoPE: Rotary Position Embedding for Video Large Language Models.☆28Nov 18, 2025Updated 10 months ago
- Official PyTorch implementation for "Where You Edit is What You Get: Text-Guided Image Editing with Region-Based Attention" (Pattern Reco…☆10Oct 1, 2024Updated last year
- Official repository for Robust Multimodal Large Language Models Against Modality Conflict☆22Jul 9, 2025Updated last year
- [ICLR 2026] Official implementation of "ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents"☆24Mar 23, 2026Updated 5 months ago
- ANDROID APP that can RECOGNIZE VLC LIVE AUDIO/VIDEO STREAMING (using free Android Developers Speech Recognition API) then TRANSLATE (usin…☆15Aug 27, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICCV 2025] Identity Preserving 3D Head Stylization with Multiview Score Distillation☆16Jul 15, 2026Updated 2 months ago
- [CVPR Findings 2026] "Circuit Tracing in Vision-Language Models"☆31Aug 18, 2026Updated last month
- ☆14Jan 22, 2025Updated last year
- EasyTTS是一个便捷的工具,旨在方便地使用第三方API服务来调用OpenAI的文本转语音(TTS)功能。 EasyTTS允许用户输入文本,并选择不同的模型、音色、格式来生成音频文件。☆10Nov 26, 2023Updated 2 years ago
- Text-guided 3D texture generation using training-free multi-diffusion in UV space.☆13Apr 7, 2025Updated last year
- ☆13Jun 13, 2024Updated 2 years ago
- ☆15Dec 11, 2024Updated last year
- Official PyTorch Implementation of Guarding Barlow Twins Against Overfitting with Mixed Samples☆19Jan 19, 2024Updated 2 years ago
- Implementation of followinf estimation algorithms in python: Kalman Filter, Extended Kalman Filter, Unscented Kalman Filter, Cubature Kal…☆12Aug 15, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A python implementation of PSNR that takes the Human visual system into account.☆13Jul 17, 2026Updated 2 months ago
- [CVPR2025] BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding☆56Feb 5, 2026Updated 7 months ago
- Repository for "Training Language Models To Explain Their Own Computations"☆37Jul 7, 2026Updated 2 months ago
- ☆19Sep 11, 2026Updated last week
- [CVPR 2025] Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention☆69Jul 16, 2024Updated 2 years ago
- ☆64Mar 3, 2025Updated last year
- ☆12Apr 18, 2025Updated last year