Back to Topics Directory
Topic Hub

#inference (30 Repositories)

Ranked open-source repositories tagged with #inference, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

83.5%

Avg Review Latency

23.2h

Filter by language

30 repositories tagged #inference

S TierC# 380

zhongkaifu/TensorSharp

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

91.8%
Merge Rate
1d
First Review
100%
1st-Timers
5
Maintainers
S TierPython 341

ultralytics/xview-yolov3

YOLOv3 training, preprocessing, validation, and inference for object detection in xView satellite imagery and the xView detection challenge.

90.0%
Merge Rate
<1h
First Review
100%
1st-Timers
2
Maintainers
B TierPython 10.6k

ultralytics/yolov3

PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.

94.1%
Merge Rate
<1h
First Review
100%
1st-Timers
1
Maintainers
S TierGo 199 1 GFIs

defilantech/LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

95.3%
Merge Rate
4d
First Review
88%
1st-Timers
4
Maintainers
S TierJUJulia 118

ReactiveBayes/ReactiveMP.jl

High-performance reactive message-passing based Bayesian inference engine

92.9%
Merge Rate
<1h
First Review
83%
1st-Timers
2
Maintainers
S TierRust 191

timtoole02/Camelid

Camelid: a Rust-native local inference backend with evidence-gated model compatibility.

86.7%
Merge Rate
23h
First Review
100%
1st-Timers
3
Maintainers
A TierTypeScript 232

mohitsoni48/TurboLLM

Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.

91.4%
Merge Rate
1d
First Review
75%
1st-Timers
5
Maintainers
A TierC# 57 2 GFIs

orcasound/orcahello

Real-time AI-assisted killer whale notification system (model and moderator portal) :star:

84.8%
Merge Rate
1h
First Review
60%
1st-Timers
9
Maintainers
A TierPython 3.5k 1 GFIs

raullenchai/Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

88.8%
Merge Rate
3d
First Review
53%
1st-Timers
15
Maintainers
A TierGo 115 11 GFIs

pmady/keda-gpu-scaler

KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required

84.8%
Merge Rate
16h
First Review
67%
1st-Timers
3
Maintainers
B TierPython 198

autonomous-ai/autonomous-grid

Your AI intranet: network the computers you already own for inference and training.

97.1%
Merge Rate
-
First Review
100%
1st-Timers
0
Maintainers
A TierC++ 217

ROCm/MIVisionX

AMD MIVisionX is a computer vision toolkit built around a highly optimized, conformant open-source implementation of the Khronos OpenVX™ 1.3.2 specification. As of the 4.0.0 release, MIVisionX ships three components: the AMD OpenVX™ engine, the AMD RPP OpenVX extension, and the RunVX graph executor — across CPU, HIP, and OpenCL backends.

93.2%
Merge Rate
21h
First Review
100%
1st-Timers
5
Maintainers
A TierPython 5.6k

gpustack/gpustack

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

88.5%
Merge Rate
2d
First Review
65%
1st-Timers
25
Maintainers
A TierPython 169

inclusionAI/Awex

A high-performance RL training-inference weight synchronization framework, designed to enable second-level parameter updates from training to inference in RL workflows

100.0%
Merge Rate
1h
First Review
100%
1st-Timers
2
Maintainers
A TierRust 192 5 GFIs

novitalabs/pegaflow

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

71.4%
Merge Rate
16h
First Review
33%
1st-Timers
4
Maintainers
A TierGo 5.4k 4 GFIs

vllm-project/semantic-router

A programmable Mixture-of-Models router for heterogeneous LLM inference

73.8%
Merge Rate
1d
First Review
68%
1st-Timers
56
Maintainers
A TierGo 312 10 GFIs

llm-d/llm-d-router

llm-d Router: The intelligent entry point for inference requests

71.6%
Merge Rate
20h
First Review
63%
1st-Timers
84
Maintainers
A TierSwift 742

SharpAI/SwiftLM

⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.

86.0%
Merge Rate
6h
First Review
67%
1st-Timers
3
Maintainers
A TierPython 9.5k

xorbitsai/inference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

86.0%
Merge Rate
1d
First Review
86%
1st-Timers
32
Maintainers
A TierC++ 125

NVIDIA-ISAAC-ROS/isaac_ros_image_segmentation

NVIDIA-accelerated, deep learned semantic image segmentation

100.0%
Merge Rate
<1h
First Review
100%
1st-Timers
2
Maintainers
A TierRust 1.7k

trymirai/uzu

A high-performance inference engine for AI models

79.7%
Merge Rate
20h
First Review
41%
1st-Timers
13
Maintainers
A TierC++ 1.5k

xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

76.7%
Merge Rate
16h
First Review
60%
1st-Timers
47
Maintainers
A TierZig 762

ddalcu/mlx-serve

Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.

70.1%
Merge Rate
7h
First Review
50%
1st-Timers
29
Maintainers
A TierGo 202 3 GFIs

NVIDIA/nvcf

Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.

70.8%
Merge Rate
2d
First Review
50%
1st-Timers
33
Maintainers
A TierC++ 227

nvidia-holoscan/holohub

Central repository for the Holoscan Ecosystem

85.5%
Merge Rate
4d
First Review
81%
1st-Timers
15
Maintainers
A TierPython 32.9k 12 GFIs

sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

49.1%
Merge Rate
5h
First Review
29%
1st-Timers
593
Maintainers
B TierKotlin 178

ferranpons/Llamatik

True on-device AI for Kotlin Multiplatform (Android, iOS, Desktop, JVM, WASM). LLM, Speech-to-Text and Image Generation — powered by llama.cpp, whisper.cpp and stable-diffusion.cpp.

100.0%
Merge Rate
17h
First Review
100%
1st-Timers
1
Maintainers
A TierShell 4.1k 12 GFIs

llm-d/llm-d

Achieve state of the art inference performance with modern accelerators on Kubernetes

71.0%
Merge Rate
1d
First Review
62%
1st-Timers
75
Maintainers
A TierC++ 6.3k

kvcache-ai/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

70.2%
Merge Rate
20h
First Review
71%
1st-Timers
82
Maintainers
A TierC++ 10.7k 3 GFIs

openvinotoolkit/openvino

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

62.5%
Merge Rate
19h
First Review
52%
1st-Timers
155
Maintainers
Best Inference Open Source Repositories & C-Rank™ | GetMerged