#llama-cpp (29 Repositories)
Ranked open-source repositories tagged with #llama-cpp, scored by pull request acceptance likelihood and maintainer engagement velocity.
38.9%
46.9h
29 repositories tagged #llama-cpp
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Blackwellboy/model-serving-minefield
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
mohitsoni48/TurboLLM
Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.
AtomicBot-ai/atomic-agent
Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.
VinRobotics/vla.cpp
A unified inference runtime for VLA models.
engeldlgado/toshllm
Run large language models locally on Intel Macs with AMD GPUs - native macOS app with Metal acceleration
OleksandrChekhovskyi/hax
A minimalist, terminal-native coding agent written in C.
off-grid-ai/OGAM
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.
kennss/SiliconScope
Sudoless Apple Silicon system monitor (native SwiftUI GUI) with ANE / Media Engine / memory-bandwidth tracking
mozilla-ai/llamafile
Distribute and run LLMs with a single file.
Luce-Org/lucebox
LLM speculative inference server for consumer & heterogeneous hardware
morganlinton/Albatross
Open source, terminal-first AI coding agent with fully transparent multi-model routing. Local (Ollama, LM Studio, MLX, llama.cpp) or cloud, your keys, one TUI. No black box.
Eric-Terminal/ETOS-LLM-Studio
A native LLM client for iOS & Apple Watch. Run local GGUF models offline via llama.cpp, or connect to OpenAI/Claude/Gemini. Features local RAG, Model Context Protocol (MCP) tools, Siri Shortcuts, and cross-device sync. Built with Swift.
gpustack/gguf-parser-go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
altic-dev/FluidVoice
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative. DM us on X for an easter egg 😉 - https://x.com/fluidvoiceapp
raketenkater/ggrun
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
yoloshii/ClawMem
On-device memory layer for AI agents. Claude Code, Hermes and OpenClaw. Hooks + MCP server + hybrid RAG search.
ggml-org/Llama-macOS
A cosy home for your LLMs.
5p00kyy/club-5060ti
Practical local LLM recipes and benchmarks for RTX 5060 Ti setups
slimeglitch/gryffin-calorai-ventus
Top AI Calorie Tracker GitHub 2026
hogeheer499-commits/strix-halo-guide
Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.
Nathanw1014/strix-halo-llamacpp
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
ThinkOffApp/CarWatch
Your car as a chat-room agent: Raspberry Pi 5 + dashcam + local AI. CodeWatch's sibling for the garage.
Scottcjn/ram-coffers
LLM infrastructure cost reduction via NUMA-aware weight banking: 147 t/s (8.8x stock llama.cpp) on refurbished enterprise POWER8. Self-hosted inference, no cloud APIs. Part of the Proof of Physical AI stack.
withcatai/catai
Run AI ✨ assistant locally! with simple API for Node.js 🚀
mdrokz/rust-llama.cpp
LLama.cpp rust bindings
julianmb/q38rocm
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
ferrumox/fox
A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.
AudarAI/Audar-ASR-V1
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.