#inference (30 Repositories)
Ranked open-source repositories tagged with #inference, scored by pull request acceptance likelihood and maintainer engagement velocity.
83.5%
23.2h
30 repositories tagged #inference
zhongkaifu/TensorSharp
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability
ultralytics/xview-yolov3
YOLOv3 training, preprocessing, validation, and inference for object detection in xView satellite imagery and the xView detection challenge.
ultralytics/yolov3
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
ReactiveBayes/ReactiveMP.jl
High-performance reactive message-passing based Bayesian inference engine
timtoole02/Camelid
Camelid: a Rust-native local inference backend with evidence-gated model compatibility.
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
orcasound/orcahello
Real-time AI-assisted killer whale notification system (model and moderator portal) :star:
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
pmady/keda-gpu-scaler
KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required
autonomous-ai/autonomous-grid
Your AI intranet: network the computers you already own for inference and training.
ROCm/MIVisionX
AMD MIVisionX is a computer vision toolkit built around a highly optimized, conformant open-source implementation of the Khronos OpenVX™ 1.3.2 specification. As of the 4.0.0 release, MIVisionX ships three components: the AMD OpenVX™ engine, the AMD RPP OpenVX extension, and the RunVX graph executor — across CPU, HIP, and OpenCL backends.
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
inclusionAI/Awex
A high-performance RL training-inference weight synchronization framework, designed to enable second-level parameter updates from training to inference in RL workflows
novitalabs/pegaflow
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
vllm-project/semantic-router
A programmable Mixture-of-Models router for heterogeneous LLM inference
llm-d/llm-d-router
llm-d Router: The intelligent entry point for inference requests
SharpAI/SwiftLM
⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.
xorbitsai/inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
NVIDIA-ISAAC-ROS/isaac_ros_image_segmentation
NVIDIA-accelerated, deep learned semantic image segmentation
trymirai/uzu
A high-performance inference engine for AI models
xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
ddalcu/mlx-serve
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
NVIDIA/nvcf
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
nvidia-holoscan/holohub
Central repository for the Holoscan Ecosystem
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
ferranpons/Llamatik
True on-device AI for Kotlin Multiplatform (Android, iOS, Desktop, JVM, WASM). LLM, Speech-to-Text and Image Generation — powered by llama.cpp, whisper.cpp and stable-diffusion.cpp.
llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
kvcache-ai/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
openvinotoolkit/openvino
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference