Back to Topics Directory
Topic Hub

#inference-engine (18 Repositories)

Ranked open-source repositories tagged with #inference-engine, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

26.2%

Avg Review Latency

16.3h

Filter by language

18 repositories tagged #inference-engine

S TierC# 380

zhongkaifu/TensorSharp

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

91.8%
Merge Rate
1d
First Review
100%
1st-Timers
5
Maintainers
A TierC++ 217

ROCm/MIVisionX

AMD MIVisionX is a computer vision toolkit built around a highly optimized, conformant open-source implementation of the Khronos OpenVX™ 1.3.2 specification. As of the 4.0.0 release, MIVisionX ships three components: the AMD OpenVX™ engine, the AMD RPP OpenVX extension, and the RunVX graph executor — across CPU, HIP, and OpenCL backends.

93.2%
Merge Rate
21h
First Review
100%
1st-Timers
5
Maintainers
A TierRust 1.1k

nobodywho-ooo/nobodywho

NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.

80.3%
Merge Rate
24h
First Review
83%
1st-Timers
12
Maintainers
A TierC++ 1.5k

xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

76.7%
Merge Rate
16h
First Review
60%
1st-Timers
47
Maintainers
A TierRust 657 4 GFIs

pegainfer-project/pegainfer

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

7.1%
Merge Rate
12h
First Review
0%
1st-Timers
4
Maintainers
A TierPython 1.1k 2 GFIs

ovg-project/kvcached

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

40.0%
Merge Rate
<1h
First Review
17%
1st-Timers
12
Maintainers
A TierRust 207

giannisanni/pulsar

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.

83.3%
Merge Rate
10h
First Review
0%
1st-Timers
3
Maintainers
D TierC++ 512

brontoguana/krasis

Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

0.0%
Merge Rate
5d
First Review
0%
1st-Timers
0
Maintainers
D TierJava 449

qualcomm/ai-hub-apps

The Qualcomm® AI Hub apps are a collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.

0.0%
Merge Rate
3d
First Review
0%
1st-Timers
1
Maintainers
D TierC++ 1.8k

dphnAI/sonar

Large-scale LLM inference engine

0.0%
Merge Rate
-
First Review
0%
1st-Timers
1
Maintainers
D TierPython 1.2k

qualcomm/ai-hub-models

Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
16
Maintainers
D TierPython 352

EfficientMoE/MoE-Infinity

PyTorch library for cost-effective, fast and easy serving of MoE models.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRuby 489

ulfurinn/wongi-engine

A rule engine written in Ruby.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC++ 341

kigner/audio.cpp-webui

audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 189

ferrumox/fox

A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJUJupyter Notebook 2.0k

siliconflow/onediff

OneDiff: An out-of-the-box acceleration library for diffusion models.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 848

insight-platform/Savant

Python Computer Vision & Video Analytics Framework With Batteries Included

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 494

buguroo/pyknow

PyKnow: Expert Systems for Python

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Inference-engine Open Source Repositories & C-Rank™ | GetMerged