Back to Topics Directory
Topic Hub

#vllm (30 Repositories)

Ranked open-source repositories tagged with #vllm, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

48.1%

Avg Review Latency

53.3h

Filter by language

30 repositories tagged #vllm

S TierPython 100

Blackwellboy/model-serving-minefield

Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.

90.0%
Merge Rate
5h
First Review
100%
1st-Timers
2
Maintainers
A TierGo 115 11 GFIs

pmady/keda-gpu-scaler

KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required

84.8%
Merge Rate
16h
First Review
67%
1st-Timers
3
Maintainers
A TierRust 192 5 GFIs

novitalabs/pegaflow

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

71.4%
Merge Rate
16h
First Review
33%
1st-Timers
4
Maintainers
A TierGo 5.4k 4 GFIs

vllm-project/semantic-router

A programmable Mixture-of-Models router for heterogeneous LLM inference

73.8%
Merge Rate
1d
First Review
68%
1st-Timers
56
Maintainers
A TierPython 873 18 GFIs

verl-project/verl-omni

Multimodal RL training framework for diffusion & omni models

73.7%
Merge Rate
2d
First Review
68%
1st-Timers
42
Maintainers
A TierPython 1.2k

ModelCloud/GPTQModel

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

94.5%
Merge Rate
14d
First Review
100%
1st-Timers
2
Maintainers
A TierRust 7.9k 6 GFIs

ai-dynamo/dynamo

A Datacenter Scale Distributed Inference Serving Framework

62.1%
Merge Rate
8h
First Review
48%
1st-Timers
163
Maintainers
A TierPython 1.6k 3 GFIs

intel/auto-round

A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers|简洁且高效的量化工具包

77.2%
Merge Rate
2d
First Review
71%
1st-Timers
19
Maintainers
A TierC++ 6.3k

kvcache-ai/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

70.2%
Merge Rate
20h
First Review
71%
1st-Timers
82
Maintainers
A TierC++ 2.7k 76 GFIs

vllm-project/vllm-ascend

Community maintained hardware plugin for vLLM on Ascend

50.9%
Merge Rate
20h
First Review
44%
1st-Timers
308
Maintainers
A TierPython 1.5k

SemiAnalysisAI/InferenceX

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

61.4%
Merge Rate
1d
First Review
52%
1st-Timers
35
Maintainers
A TierPython 903

Tencent-Hunyuan/UniRL

UniRL is a Framework for Unified Multimodal Model Reinforcement Learning

65.7%
Merge Rate
6h
First Review
56%
1st-Timers
23
Maintainers
A TierPython 1.1k 2 GFIs

ovg-project/kvcached

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

40.0%
Merge Rate
<1h
First Review
17%
1st-Timers
12
Maintainers
A TierPython 195

syv-ai/qwen38-27b-rtx3090

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

62.5%
Merge Rate
3h
First Review
50%
1st-Timers
20
Maintainers
B TierPython 230

lightseekorg/TorchSpec

A PyTorch native library for training speculative decoding models

81.4%
Merge Rate
11d
First Review
79%
1st-Timers
2
Maintainers
B TierPython 3.0k 2 GFIs

containers/ramalama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

50.0%
Merge Rate
<1h
First Review
67%
1st-Timers
16
Maintainers
B TierPython 11.4k 2 GFIs

LMCache/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

53.9%
Merge Rate
20h
First Review
43%
1st-Timers
121
Maintainers
B TierRust 489

smg-project/smg

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

73.5%
Merge Rate
5d
First Review
48%
1st-Timers
19
Maintainers
B TierC++ 314

mudler/vllm.cpp

a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (for example, RadixAttention, Cache-aware scheduling)

32.9%
Merge Rate
8h
First Review
14%
1st-Timers
13
Maintainers
B TierTypeScript 1.7k

sybil-solutions/local-studio

Control panel for VLLM, Sglang, llama.cpp, exllamav3

52.4%
Merge Rate
3d
First Review
17%
1st-Timers
9
Maintainers
B TierGo 5.5k

mostlygeek/llama-swap

Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

60.0%
Merge Rate
2d
First Review
50%
1st-Timers
59
Maintainers
B TierPython 1.5k 1 GFIs

waybarrios/vllm-mlx

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

61.5%
Merge Rate
12d
First Review
34%
1st-Timers
20
Maintainers
D TierPython 293

AstraNetLab/CacheRoute

CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
1
Maintainers
D TierPython 117

5p00kyy/club-5060ti

Practical local LLM recipes and benchmarks for RTX 5060 Ti setups

0.0%
Merge Rate
7d
First Review
0%
1st-Timers
1
Maintainers
D TierC++ 1.1k

jmaczan/tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 420

FujitsuResearch/OneCompression

Python package for LLM compression

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJUJupyter Notebook 18.6k

meta-llama/llama-cookbook

Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJavaScript 348

jasonacox/TinyLLM

Setup and run a local LLM and Chatbot using consumer grade hardware.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 465

runpod-workers/worker-vllm

The Runpod worker template for serving our large language model endpoints. Powered by vLLM.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierGo 309

InftyAI/llmaz

☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Vllm Open Source Repositories & C-Rank™ | GetMerged