Back to Topics Directory
Topic Hub

#multimodal (30 Repositories)

Ranked open-source repositories tagged with #multimodal, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

55.6%

Avg Review Latency

58.2h

Filter by language

30 repositories tagged #multimodal

A TierZig 426

antflydb/antfly

85.4%
Merge Rate
17h
First Review
89%
1st-Timers
9
Maintainers
A TierPython 873 18 GFIs

verl-project/verl-omni

Multimodal RL training framework for diffusion & omni models

73.7%
Merge Rate
2d
First Review
68%
1st-Timers
42
Maintainers
A TierC# 487

clawdotnet/openclaw.net

Self-hosted Personal AI + agent runtime in .NET (NativeAOT-friendly)

95.8%
Merge Rate
4d
First Review
80%
1st-Timers
2
Maintainers
B TierTypeScript 992

tong-io/tongflow

TongFlow — Multimodal GenAI Studio

92.1%
Merge Rate
19d
First Review
100%
1st-Timers
0
Maintainers
A TierRust 3.2k 1 GFIs

vortex-data/vortex

An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.

67.9%
Merge Rate
3d
First Review
67%
1st-Timers
23
Maintainers
A TierPython 1.6k

pixeltable/pixeltable

Unified multimodal backend for AI data apps

73.6%
Merge Rate
19h
First Review
43%
1st-Timers
10
Maintainers
A TierPython 15.4k

modelscope/ms-swift

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).

77.9%
Merge Rate
3d
First Review
51%
1st-Timers
61
Maintainers
A TierPython 6.3k 153 GFIs

vllm-project/vllm-omni

A framework for efficient model inference with omni-modality models

56.0%
Merge Rate
23h
First Review
37%
1st-Timers
235
Maintainers
A TierC++ 10.3k

RunanywhereAI/runanywhere-sdks

Production ready toolkit to run AI locally

70.9%
Merge Rate
13h
First Review
55%
1st-Timers
10
Maintainers
A TierRust 21.2k

screenpipe/screenpipe

YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)

67.9%
Merge Rate
1d
First Review
34%
1st-Timers
22
Maintainers
B TierPython 946 4 GFIs

sgl-project/sglang-omni

SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

54.8%
Merge Rate
13h
First Review
40%
1st-Timers
90
Maintainers
B TierC++ 1.6k

UbiquitousLearning/mllm

Fast Multimodal LLM on Mobile Devices

75.0%
Merge Rate
17h
First Review
50%
1st-Timers
2
Maintainers
B TierPython 2.8k

datachain-ai/datachain

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

40.0%
Merge Rate
2h
First Review
67%
1st-Timers
9
Maintainers
B TierRust 5.7k 9 GFIs

Eventual-Inc/Daft

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

66.8%
Merge Rate
6d
First Review
60%
1st-Timers
37
Maintainers
B TierC# 569

microsoft/psi

Platform for Situated Intelligence

100.0%
Merge Rate
1d
First Review
0%
1st-Timers
1
Maintainers
B TierPython 4.4k

EvolvingLMMs-Lab/lmms-eval

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

71.9%
Merge Rate
4d
First Review
67%
1st-Timers
18
Maintainers
B TierPython 270

Anionex/codex-vision-proxy

让纯文本模型在 Codex 中无障碍看图(view_image)的更优方案,附为纯文本 LLM 设计的视觉工具包&skill | A superior approach for enabling text-only models to seamlessly use Codex’s built-in view_image, plus a vision toolkit & skill designed for pure-text LLMs.

66.7%
Merge Rate
3h
First Review
67%
1st-Timers
1
Maintainers
B TierJavaScript 980

ysr666/dsh-vision-router

Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

12.8%
Merge Rate
3h
First Review
0%
1st-Timers
5
Maintainers
B TierTypeScript 3.0k

TanStack/ai

🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.

70.4%
Merge Rate
3d
First Review
65%
1st-Timers
11
Maintainers
B TierPython 5.2k 1 GFIs

InternLM/xtuner

A Next-Generation Training Engine Built for Ultra-Large MoE Models

38.9%
Merge Rate
3h
First Review
25%
1st-Timers
12
Maintainers
B TierJavaScript 65.4k

Mintplex-Labs/anything-llm

Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience

61.9%
Merge Rate
2d
First Review
34%
1st-Timers
57
Maintainers
B TierRust 11.3k

rerun-io/rerun

Visualize, query, and stream to train on multimodal robotics data.

40.5%
Merge Rate
23h
First Review
57%
1st-Timers
48
Maintainers
B TierTypeScript 6.3k

genkit-ai/genkit

Open-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google

57.8%
Merge Rate
3d
First Review
28%
1st-Timers
21
Maintainers
B TierJavaScript 100

shixinnt/codex-image-context-runtime

A Codex plugin and local MCP runtime for context-bounded image generation and inspection.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
1
Maintainers
B TierPython 3.1k

xlang-ai/OSWorld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

49.2%
Merge Rate
1d
First Review
43%
1st-Timers
8
Maintainers
C TierPython 486 1 GFIs

TeleAI-UAGI/telemem

TeleMem is a high-performance drop-in replacement for Mem0, featuring semantic deduplication, long-term dialogue memory, and multimodal video reasoning.

100.0%
Merge Rate
-
First Review
100%
1st-Timers
0
Maintainers
D TierPython 162

xiincs/claude-code-vision-skill

为 Claude Code 赋能多模态视觉能力,支持豆包、通义千问、GPT-4o 等模型,用于截图 / UI / 图表分析;适配 DeepSeek 等无视觉底座,搭配 browser-harness 可做前端布局自动化检查。

0.0%
Merge Rate
15d
First Review
0%
1st-Timers
1
Maintainers
D TierPython 153

isLinXu/paper-list

autoupdate paper list

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJavaScript 100

ZSeven-W/dsh-crew

DeepSeek Harness (DSH) plugin: dispatch work to DSH agents from Claude Code / Codex — native subagent progress, in-host worker sessions with per-tier presets, and a multimodal bridge that lends the text-only harness vision and image generation.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 344

RoffyS/MarkEverythingDown

Convert files (PDF, image, Word, PPT, Excel, notebooks, code snippets) to markdown using powerful multimodal LLM

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Multimodal Open Source Repositories & C-Rank™ | GetMerged