#multimodal (30 Repositories)
Ranked open-source repositories tagged with #multimodal, scored by pull request acceptance likelihood and maintainer engagement velocity.
55.6%
58.2h
30 repositories tagged #multimodal
verl-project/verl-omni
Multimodal RL training framework for diffusion & omni models
clawdotnet/openclaw.net
Self-hosted Personal AI + agent runtime in .NET (NativeAOT-friendly)
tong-io/tongflow
TongFlow — Multimodal GenAI Studio
vortex-data/vortex
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.
pixeltable/pixeltable
Unified multimodal backend for AI data apps
modelscope/ms-swift
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
vllm-project/vllm-omni
A framework for efficient model inference with omni-modality models
RunanywhereAI/runanywhere-sdks
Production ready toolkit to run AI locally
screenpipe/screenpipe
YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)
sgl-project/sglang-omni
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
UbiquitousLearning/mllm
Fast Multimodal LLM on Mobile Devices
datachain-ai/datachain
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Eventual-Inc/Daft
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
microsoft/psi
Platform for Situated Intelligence
EvolvingLMMs-Lab/lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Anionex/codex-vision-proxy
让纯文本模型在 Codex 中无障碍看图(view_image)的更优方案,附为纯文本 LLM 设计的视觉工具包&skill | A superior approach for enabling text-only models to seamlessly use Codex’s built-in view_image, plus a vision toolkit & skill designed for pure-text LLMs.
ysr666/dsh-vision-router
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
TanStack/ai
🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.
InternLM/xtuner
A Next-Generation Training Engine Built for Ultra-Large MoE Models
Mintplex-Labs/anything-llm
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
rerun-io/rerun
Visualize, query, and stream to train on multimodal robotics data.
genkit-ai/genkit
Open-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google
shixinnt/codex-image-context-runtime
A Codex plugin and local MCP runtime for context-bounded image generation and inspection.
xlang-ai/OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
TeleAI-UAGI/telemem
TeleMem is a high-performance drop-in replacement for Mem0, featuring semantic deduplication, long-term dialogue memory, and multimodal video reasoning.
xiincs/claude-code-vision-skill
为 Claude Code 赋能多模态视觉能力,支持豆包、通义千问、GPT-4o 等模型,用于截图 / UI / 图表分析;适配 DeepSeek 等无视觉底座,搭配 browser-harness 可做前端布局自动化检查。
isLinXu/paper-list
autoupdate paper list
ZSeven-W/dsh-crew
DeepSeek Harness (DSH) plugin: dispatch work to DSH agents from Claude Code / Codex — native subagent progress, in-host worker sessions with per-tier presets, and a multimodal bridge that lends the text-only harness vision and image generation.
RoffyS/MarkEverythingDown
Convert files (PDF, image, Word, PPT, Excel, notebooks, code snippets) to markdown using powerful multimodal LLM