Back to Topics Directory
Topic Hub

#inference-server (5 Repositories)

Ranked open-source repositories tagged with #inference-server, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

30.3%

Avg Review Latency

37.6h

Filter by language

5 repositories tagged #inference-server

B TierPython 3.0k 2 GFIs

containers/ramalama

RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

50.0%
Merge Rate
<1h
First Review
67%
1st-Timers
16
Maintainers
B TierPython 20.7k

jundot/omlx

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

56.4%
Merge Rate
1d
First Review
53%
1st-Timers
123
Maintainers
C TierGo 267

raketenkater/ggrun

llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

20.0%
Merge Rate
3d
First Review
0%
1st-Timers
0
Maintainers
C TierRust 5.8k

Michael-A-Kuykendall/shimmy

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

25.0%
Merge Rate
4d
First Review
100%
1st-Timers
2
Maintainers
D TierScala 166

autodeployai/ai-serving

Serving AI/ML models in the open standard formats PMML and ONNX with both HTTP (REST API) and gRPC endpoints

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Inference-server Open Source Repositories & C-Rank™ | GetMerged