Back to Topics Directory
Topic Hub

#llama-cpp (29 Repositories)

Ranked open-source repositories tagged with #llama-cpp, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

38.9%

Avg Review Latency

46.9h

Filter by language

29 repositories tagged #llama-cpp

S TierGo 199 1 GFIs

defilantech/LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

95.3%
Merge Rate
4d
First Review
88%
1st-Timers
4
Maintainers
S TierPython 100

Blackwellboy/model-serving-minefield

Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.

90.0%
Merge Rate
5h
First Review
100%
1st-Timers
2
Maintainers
A TierTypeScript 232

mohitsoni48/TurboLLM

Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.

91.4%
Merge Rate
1d
First Review
75%
1st-Timers
5
Maintainers
A TierTypeScript 2.5k

AtomicBot-ai/atomic-agent

Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.

68.8%
Merge Rate
23h
First Review
17%
1st-Timers
9
Maintainers
B TierC++ 163

VinRobotics/vla.cpp

A unified inference runtime for VLA models.

100.0%
Merge Rate
6d
First Review
100%
1st-Timers
2
Maintainers
B TierSwift 123

engeldlgado/toshllm

Run large language models locally on Intel Macs with AMD GPUs - native macOS app with Metal acceleration

40.0%
Merge Rate
16h
First Review
50%
1st-Timers
2
Maintainers
B TierC 248

OleksandrChekhovskyi/hax

A minimalist, terminal-native coding agent written in C.

12.5%
Merge Rate
<1h
First Review
100%
1st-Timers
8
Maintainers
B TierTypeScript 3.0k

off-grid-ai/OGAM

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.

73.8%
Merge Rate
10d
First Review
71%
1st-Timers
9
Maintainers
B TierSwift 842

kennss/SiliconScope

Sudoless Apple Silicon system monitor (native SwiftUI GUI) with ANE / Media Engine / memory-bandwidth tracking

76.5%
Merge Rate
3d
First Review
40%
1st-Timers
1
Maintainers
B TierC++ 25.6k

mozilla-ai/llamafile

Distribute and run LLMs with a single file.

84.6%
Merge Rate
4d
First Review
80%
1st-Timers
11
Maintainers
B TierC++ 2.8k

Luce-Org/lucebox

LLM speculative inference server for consumer & heterogeneous hardware

68.5%
Merge Rate
5d
First Review
69%
1st-Timers
17
Maintainers
B TierRust 223

morganlinton/Albatross

Open source, terminal-first AI coding agent with fully transparent multi-model routing. Local (Ollama, LM Studio, MLX, llama.cpp) or cloud, your keys, one TUI. No black box.

78.9%
Merge Rate
3d
First Review
75%
1st-Timers
2
Maintainers
B TierSwift 190

Eric-Terminal/ETOS-LLM-Studio

A native LLM client for iOS & Apple Watch. Run local GGUF models offline via llama.cpp, or connect to OpenAI/Claude/Gemini. Features local RAG, Model Context Protocol (MCP) tools, Siri Shortcuts, and cross-device sync. Built with Swift.

50.0%
Merge Rate
2d
First Review
100%
1st-Timers
2
Maintainers
B TierGo 286

gpustack/gguf-parser-go

Review/Check GGUF files and estimate the memory usage and maximum tokens per second.

54.5%
Merge Rate
10h
First Review
0%
1st-Timers
2
Maintainers
B TierSwift 10.8k

altic-dev/FluidVoice

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X for an easter egg 😉 - https://x.com/fluidvoiceapp

52.2%
Merge Rate
3d
First Review
40%
1st-Timers
68
Maintainers
C TierGo 267

raketenkater/ggrun

llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

20.0%
Merge Rate
3d
First Review
0%
1st-Timers
0
Maintainers
C TierTypeScript 195

yoloshii/ClawMem

On-device memory layer for AI agents. Claude Code, Hermes and OpenClaw. Hooks + MCP server + hybrid RAG search.

71.4%
Merge Rate
2d
First Review
0%
1st-Timers
1
Maintainers
D TierSwift 1.5k

ggml-org/Llama-macOS

A cosy home for your LLMs.

0.0%
Merge Rate
9h
First Review
0%
1st-Timers
2
Maintainers
D TierPython 117

5p00kyy/club-5060ti

Practical local LLM recipes and benchmarks for RTX 5060 Ti setups

0.0%
Merge Rate
7d
First Review
0%
1st-Timers
1
Maintainers
D TierHTML 115

slimeglitch/gryffin-calorai-ventus

Top AI Calorie Tracker GitHub 2026

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 291

hogeheer499-commits/strix-halo-guide

Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 104

Nathanw1014/strix-halo-llamacpp

Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 139

ThinkOffApp/CarWatch

Your car as a chat-room agent: Raspberry Pi 5 + dashcam + local AI. CodeWatch's sibling for the garage.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 160

Scottcjn/ram-coffers

LLM infrastructure cost reduction via NUMA-aware weight banking: 147 t/s (8.8x stock llama.cpp) on refurbished enterprise POWER8. Self-hosted inference, no cloud APIs. Part of the Proof of Physical AI stack.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 499

withcatai/catai

Run AI ✨ assistant locally! with simple API for Node.js 🚀

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 424

mdrokz/rust-llama.cpp

LLama.cpp rust bindings

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 136

julianmb/q38rocm

Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 189

ferrumox/fox

A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 498

AudarAI/Audar-ASR-V1

Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Llama-cpp Open Source Repositories & C-Rank™ | GetMerged