Back to Topics Directory
Topic Hub

#llm-serving (10 Repositories)

Ranked open-source repositories tagged with #llm-serving, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

38.9%

Avg Review Latency

21.4h

Filter by language

10 repositories tagged #llm-serving

S TierPython 100

Blackwellboy/model-serving-minefield

Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.

90.0%
Merge Rate
5h
First Review
100%
1st-Timers
2
Maintainers
A TierGo 803

helixml/helix

♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude, Codex, Gemini and open models on a full private AI Stack ♾️

93.9%
Merge Rate
6d
First Review
80%
1st-Timers
6
Maintainers
A TierPython 14.5k

NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

59.7%
Merge Rate
3h
First Review
55%
1st-Timers
222
Maintainers
A TierC++ 2.7k 76 GFIs

vllm-project/vllm-ascend

Community maintained hardware plugin for vLLM on Ascend

50.9%
Merge Rate
20h
First Review
44%
1st-Timers
308
Maintainers
A TierRust 657 4 GFIs

pegainfer-project/pegainfer

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

7.1%
Merge Rate
12h
First Review
0%
1st-Timers
4
Maintainers
B TierPython 90.5k 25 GFIs

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

40.9%
Merge Rate
10h
First Review
30%
1st-Timers
939
Maintainers
B TierCUCuda 1.3k

alibaba/rtp-llm

RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

29.1%
Merge Rate
4h
First Review
36%
1st-Timers
11
Maintainers
B TierPython 8.8k

bentoml/BentoML

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

17.2%
Merge Rate
16h
First Review
33%
1st-Timers
9
Maintainers
D TierPython 476

hpcaitech/SwiftInfer

Efficient AI Inference & Serving

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 3.0k

thu-pacman/chitu

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Llm-serving Open Source Repositories & C-Rank™ | GetMerged