Back to Topics Directory
Topic Hub

#vision-language-model (19 Repositories)

Ranked open-source repositories tagged with #vision-language-model, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

19.9%

Avg Review Latency

35.1h

Filter by language

19 repositories tagged #vision-language-model

B TierPython 1.2k

EvolvingLMMs-Lab/LLaVA-OneVision-2

Fully Open Framework for Democratized Multimodal Training

83.6%
Merge Rate
3d
First Review
88%
1st-Timers
0
Maintainers
B TierPython 2.0k

2U1/Qwen-VL-Series-Finetune

An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

100.0%
Merge Rate
8d
First Review
100%
1st-Timers
1
Maintainers
B TierPython 5.4k

Blaizzy/mlx-vlm

MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

62.1%
Merge Rate
12h
First Review
47%
1st-Timers
60
Maintainers
B TierPython 4.4k

EvolvingLMMs-Lab/lmms-eval

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

71.9%
Merge Rate
4d
First Review
67%
1st-Timers
18
Maintainers
B TierPython 1.5k 1 GFIs

waybarrios/vllm-mlx

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

61.5%
Merge Rate
12d
First Review
34%
1st-Timers
20
Maintainers
C TierPython 425

Mengqi-Lei/count-anything

Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/

0.0%
Merge Rate
-
First Review
0%
1st-Timers
1
Maintainers
D TierPython 493

mll-lab-nu/VAGEN

World model reinforcement learning for multi-turn VLM agents. RL for vision framework (NeurIPS 2025).

0.0%
Merge Rate
-
First Review
0%
1st-Timers
1
Maintainers
D TierHTML 301

jolibrain/colette

Multimodal RAG to search and interact locally with technical documents of any kind

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 4.7k

PKU-Alignment/align-anything

Align Anything: Training All-modality Model with Feedback

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 2.9k

InternLM/InternLM-XComposer

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 347

zhengli97/PromptKD

[CVPR 2024] Official PyTorch Code for "PromptKD: Unsupervised Prompt Distillation for Vision-Language Models"

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMulti-language 642

XiaomiMiMo/MiMo-VL

MiMo-VL

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 486

OpenMOSS/MOSS-VL

MOSS-VL is the core multimodal model series within the OpenMOSS ecosystem, dedicated to visual understanding.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierGo 443

sh4den/Montscan

🖨️ Automated scanner document processor with AI-powered naming and WebDav integration. Receives scans via FTP, extracts text using Vision AI, generates intelligent filenames with Ollama AI, and uploads to your cloud storage.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 343

meituan/EvoCUA

EvoCUA: Evolving Computer Use Agent

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 344

OpenDriveLab/RISE

[RSS 2026] Code for RISE: Self-Improving Robot Policy with Compositional World Model

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 131

MrGiovanni/R-Super

[MICCAI 2025 Best Paper Award] Learning Segmentation from Radiology Reports

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMulti-language 796

zhengli97/Awesome-Prompt-Adapter-Learning-for-VLMs-CLIP

A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 159

tongjingqi/Game-RL

Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Vision-language-model Open Source Repositories & C-Rank™ | GetMerged