#inference-engine (18 Repositories)
Ranked open-source repositories tagged with #inference-engine, scored by pull request acceptance likelihood and maintainer engagement velocity.
26.2%
16.3h
18 repositories tagged #inference-engine
zhongkaifu/TensorSharp
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability
ROCm/MIVisionX
AMD MIVisionX is a computer vision toolkit built around a highly optimized, conformant open-source implementation of the Khronos OpenVX™ 1.3.2 specification. As of the 4.0.0 release, MIVisionX ships three components: the AMD OpenVX™ engine, the AMD RPP OpenVX extension, and the RunVX graph executor — across CPU, HIP, and OpenCL backends.
nobodywho-ooo/nobodywho
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
pegainfer-project/pegainfer
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
ovg-project/kvcached
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
giannisanni/pulsar
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
brontoguana/krasis
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
qualcomm/ai-hub-apps
The Qualcomm® AI Hub apps are a collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
dphnAI/sonar
Large-scale LLM inference engine
qualcomm/ai-hub-models
Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
EfficientMoE/MoE-Infinity
PyTorch library for cost-effective, fast and easy serving of MoE models.
ulfurinn/wongi-engine
A rule engine written in Ruby.
kigner/audio.cpp-webui
audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.
ferrumox/fox
A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.
siliconflow/onediff
OneDiff: An out-of-the-box acceleration library for diffusion models.
insight-platform/Savant
Python Computer Vision & Video Analytics Framework With Batteries Included
buguroo/pyknow
PyKnow: Expert Systems for Python