#vllm (30 Repositories)
Ranked open-source repositories tagged with #vllm, scored by pull request acceptance likelihood and maintainer engagement velocity.
48.1%
53.3h
30 repositories tagged #vllm
Blackwellboy/model-serving-minefield
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
pmady/keda-gpu-scaler
KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required
novitalabs/pegaflow
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
vllm-project/semantic-router
A programmable Mixture-of-Models router for heterogeneous LLM inference
verl-project/verl-omni
Multimodal RL training framework for diffusion & omni models
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
intel/auto-round
A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers|简洁且高效的量化工具包
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
vllm-project/vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
SemiAnalysisAI/InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
Tencent-Hunyuan/UniRL
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
ovg-project/kvcached
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
syv-ai/qwen38-27b-rtx3090
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
lightseekorg/TorchSpec
A PyTorch native library for training speculative decoding models
containers/ramalama
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
smg-project/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
mudler/vllm.cpp
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (for example, RadixAttention, Cache-aware scheduling)
sybil-solutions/local-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
mostlygeek/llama-swap
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
waybarrios/vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
AstraNetLab/CacheRoute
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
5p00kyy/club-5060ti
Practical local LLM recipes and benchmarks for RTX 5060 Ti setups
jmaczan/tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
FujitsuResearch/OneCompression
Python package for LLM compression
meta-llama/llama-cookbook
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
jasonacox/TinyLLM
Setup and run a local LLM and Chatbot using consumer grade hardware.
runpod-workers/worker-vllm
The Runpod worker template for serving our large language model endpoints. Powered by vLLM.
InftyAI/llmaz
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!