jiaquan301 / simple_vlmView on GitHub
A minimal Vision Language Model (VLM) implementation built from scratch in PyTorch, extending LLM capabilities with visual understanding. Features cross-modal attention, vision encoder, and complete training pipeline.
15Jul 13, 2025Updated last year

Alternatives and similar repositories for simple_vlm

Users that are interested in simple_vlm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?