#sglang (15 Repositories)
Ranked open-source repositories tagged with #sglang, scored by pull request acceptance likelihood and maintainer engagement velocity.
49.3%
63.8h
15 repositories tagged #sglang
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
intel/auto-round
A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers|简洁且高效的量化工具包
ModelCloud/GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Tencent-Hunyuan/UniRL
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
SemiAnalysisAI/InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
sgl-project/sglang-omni
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
smg-project/smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
sybil-solutions/local-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
sgl-project/SpecForge
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
sgl-project/rbg
A workload for deploying LLM inference services on Kubernetes
ovg-project/kvcached
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
InftyAI/llmaz
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
OpenMOSS/MOVA
MOVA: Towards Scalable and Synchronized Video–Audio Generation