Back to Topics Directory
Topic Hub

#llm-inference (30 Repositories)

Ranked open-source repositories tagged with #llm-inference, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

58.7%

Avg Review Latency

77.5h

Filter by language

30 repositories tagged #llm-inference

B TierPython 188

harleyszhang/lite_llama

A light llama-like llm inference framework based on the triton kernel.

96.0%
Merge Rate
2d
First Review
100%
1st-Timers
0
Maintainers
B TierRust 186

feichai0017/loom-infer

Rust-native GPU operator library for LLM inference, built with cuda-oxide

98.9%
Merge Rate
14h
First Review
100%
1st-Timers
1
Maintainers
A TierRust 3.1k

spiceai/spiceai

Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.

84.8%
Merge Rate
19h
First Review
86%
1st-Timers
18
Maintainers
A TierC++ 1.7k

software-mansion/react-native-executorch

Declarative way to run AI models in React Native on device, powered by ExecuTorch.

88.5%
Merge Rate
4d
First Review
83%
1st-Timers
11
Maintainers
A TierTypeScript 581

felladrin/MiniSearch

Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space

83.3%
Merge Rate
20h
First Review
100%
1st-Timers
3
Maintainers
B TierTypeScript 468

NPC-Worldwide/incognide

Explore the unknown, build the future, own your data.

90.0%
Merge Rate
4d
First Review
100%
1st-Timers
0
Maintainers
A TierRust 7.9k 6 GFIs

ai-dynamo/dynamo

A Datacenter Scale Distributed Inference Serving Framework

62.1%
Merge Rate
8h
First Review
48%
1st-Timers
163
Maintainers
A TierC++ 1.5k

xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

76.7%
Merge Rate
16h
First Review
60%
1st-Timers
47
Maintainers
A TierTypeScript 452

tetherto/qvac

Open-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, and more. Cross-platform: Linux, macOS, Windows, Android, iOS.

71.4%
Merge Rate
1d
First Review
66%
1st-Timers
41
Maintainers
A TierC++ 10.7k 3 GFIs

openvinotoolkit/openvino

OpenVINO™ is an open source toolkit for optimizing and deploying AI inference

62.5%
Merge Rate
19h
First Review
52%
1st-Timers
155
Maintainers
A TierRust 554 4 GFIs

warpfront/hipfire

RDNA-native LLM inference engine in Rust.

70.2%
Merge Rate
6d
First Review
61%
1st-Timers
12
Maintainers
A TierC++ 5.4k

lemonade-sdk/lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

71.1%
Merge Rate
20h
First Review
67%
1st-Timers
63
Maintainers
A TierGo 1.7k

beam-cloud/beta9

Ultrafast serverless GPU inference, sandboxes, and background jobs

91.4%
Merge Rate
12d
First Review
25%
1st-Timers
4
Maintainers
A TierRust 669 1 GFIs

Avarok-Cybersecurity/atlas

Pure Rust Inference Engine

51.9%
Merge Rate
20h
First Review
38%
1st-Timers
27
Maintainers
A TierPython 8.0k

InternLM/lmdeploy

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

70.3%
Merge Rate
9h
First Review
50%
1st-Timers
34
Maintainers
A TierPython 6.2k 5 GFIs

flashinfer-ai/flashinfer

FlashInfer: Kernel Library for LLM Serving

57.3%
Merge Rate
21h
First Review
51%
1st-Timers
137
Maintainers
A TierPython 195

syv-ai/qwen38-27b-rtx3090

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

62.5%
Merge Rate
3h
First Review
50%
1st-Timers
20
Maintainers
A TierPHP 2.1k

neuron-core/neuron-ai

The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that interact with your data and UI.

69.4%
Merge Rate
6h
First Review
58%
1st-Timers
17
Maintainers
B TierC++ 5.9k

cactus-compute/cactus

Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.

46.2%
Merge Rate
3h
First Review
67%
1st-Timers
14
Maintainers
B TierGo 498

ome-projects/ome

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

70.3%
Merge Rate
9d
First Review
41%
1st-Timers
10
Maintainers
B TierGo 795 4 GFIs

kubernetes-sigs/lws

LeaderWorkerSet: An API for deploying a group of pods as a unit of replication

58.5%
Merge Rate
2d
First Review
44%
1st-Timers
26
Maintainers
B TierRust 7.0k 2 GFIs

katanemo/plano

Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.

65.9%
Merge Rate
13d
First Review
44%
1st-Timers
10
Maintainers
B TierC++ 915

foldl/chatllm.cpp

Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)

100.0%
Merge Rate
24d
First Review
0%
1st-Timers
0
Maintainers
B TierPython 8.8k

bentoml/BentoML

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

17.2%
Merge Rate
16h
First Review
33%
1st-Timers
9
Maintainers
B TierKotlin 153

NightMean/OlliteRT

Turn your Android phone into an OpenAI-compatible LLM inference server — Fully local, private and Open Source

20.0%
Merge Rate
2d
First Review
0%
1st-Timers
2
Maintainers
C TierRust 5.8k

Michael-A-Kuykendall/shimmy

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

25.0%
Merge Rate
4d
First Review
100%
1st-Timers
2
Maintainers
D TierC++ 512

brontoguana/krasis

Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

0.0%
Merge Rate
5d
First Review
0%
1st-Timers
0
Maintainers
D TierC++ 118

Mobile-Artificial-Intelligence/llama_sdk

lcpp is a dart implementation of llama.cpp used by the mobile artificial intelligence distribution (maid)

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierScala 137

mit-han-lab/spatten

[HPCA'21] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMAMarkdown 209

amitshekhariitbhu/llm-inference-engineering

Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Llm-inference Open Source Repositories & C-Rank™ | GetMerged