#vision-language-model (19 Repositories)
Ranked open-source repositories tagged with #vision-language-model, scored by pull request acceptance likelihood and maintainer engagement velocity.
19.9%
35.1h
19 repositories tagged #vision-language-model
EvolvingLMMs-Lab/LLaVA-OneVision-2
Fully Open Framework for Democratized Multimodal Training
2U1/Qwen-VL-Series-Finetune
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
Blaizzy/mlx-vlm
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
EvolvingLMMs-Lab/lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
waybarrios/vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
Mengqi-Lei/count-anything
Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/
mll-lab-nu/VAGEN
World model reinforcement learning for multi-turn VLM agents. RL for vision framework (NeurIPS 2025).
jolibrain/colette
Multimodal RAG to search and interact locally with technical documents of any kind
PKU-Alignment/align-anything
Align Anything: Training All-modality Model with Feedback
InternLM/InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
zhengli97/PromptKD
[CVPR 2024] Official PyTorch Code for "PromptKD: Unsupervised Prompt Distillation for Vision-Language Models"
XiaomiMiMo/MiMo-VL
MiMo-VL
OpenMOSS/MOSS-VL
MOSS-VL is the core multimodal model series within the OpenMOSS ecosystem, dedicated to visual understanding.
sh4den/Montscan
🖨️ Automated scanner document processor with AI-powered naming and WebDav integration. Receives scans via FTP, extracts text using Vision AI, generates intelligent filenames with Ollama AI, and uploads to your cloud storage.
meituan/EvoCUA
EvoCUA: Evolving Computer Use Agent
OpenDriveLab/RISE
[RSS 2026] Code for RISE: Self-Improving Robot Policy with Compositional World Model
MrGiovanni/R-Super
[MICCAI 2025 Best Paper Award] Learning Segmentation from Radiology Reports
zhengli97/Awesome-Prompt-Adapter-Learning-for-VLMs-CLIP
A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.
tongjingqi/Game-RL
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning