Back to Topics Directory
Topic Hub

#gguf (28 Repositories)

Ranked open-source repositories tagged with #gguf, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

50.6%

Avg Review Latency

38.8h

Filter by language

28 repositories tagged #gguf

S TierC# 380

zhongkaifu/TensorSharp

A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/Linux with full GPU capability

91.8%
Merge Rate
1d
First Review
100%
1st-Timers
5
Maintainers
S TierGo 199 1 GFIs

defilantech/LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

95.3%
Merge Rate
4d
First Review
88%
1st-Timers
4
Maintainers
S TierRust 191

timtoole02/Camelid

Camelid: a Rust-native local inference backend with evidence-gated model compatibility.

86.7%
Merge Rate
23h
First Review
100%
1st-Timers
3
Maintainers
A TierTypeScript 232

mohitsoni48/TurboLLM

Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.

91.4%
Merge Rate
1d
First Review
75%
1st-Timers
5
Maintainers
A TierJava 270 1 GFIs

beehive-lab/GPULlama3.java

GPU-accelerated Llama3.java inference in pure Java using TornadoVM.

85.7%
Merge Rate
<1h
First Review
50%
1st-Timers
4
Maintainers
B TierRust 328

avifenesh/memra

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

73.1%
Merge Rate
2h
First Review
100%
1st-Timers
0
Maintainers
A TierRust 142 1 GFIs

llamastash/llamastash

Zero-overhead, terminal-native local-LLM runtime manager. Launches, supervises, and routes local models behind one OpenAI-compatible endpoint.

84.2%
Merge Rate
13h
First Review
100%
1st-Timers
11
Maintainers
A TierPython 1.6k 3 GFIs

intel/auto-round

A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers|简洁且高效的量化工具包

77.2%
Merge Rate
2d
First Review
71%
1st-Timers
19
Maintainers
A TierGo 1.4k 1 GFIs

kitops-ml/kitops

An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.

77.4%
Merge Rate
19h
First Review
70%
1st-Timers
14
Maintainers
A TierRust 207

giannisanni/pulsar

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.

83.3%
Merge Rate
10h
First Review
0%
1st-Timers
3
Maintainers
A TierPython 2.5k 24 GFIs

MakazhanAlpamys/Soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

37.3%
Merge Rate
3h
First Review
38%
1st-Timers
13
Maintainers
A TierTypeScript 2.5k

AtomicBot-ai/atomic-agent

Local First Ai Agent. Optimized for Local Ai models. Long context window. Proper tools callings. Runs privately on your device.

68.8%
Merge Rate
23h
First Review
17%
1st-Timers
9
Maintainers
A TierRust 34.5k

AlexsJones/llmfit

Hundreds of models & providers. One command to find what runs on your hardware.

62.1%
Merge Rate
20h
First Review
58%
1st-Timers
40
Maintainers
B TierC++ 163

VinRobotics/vla.cpp

A unified inference runtime for VLA models.

100.0%
Merge Rate
6d
First Review
100%
1st-Timers
2
Maintainers
B TierC++ 1.8k

handy-computer/transcribe.cpp

ggml speech-to-text inference for 16+ model families

78.6%
Merge Rate
7d
First Review
40%
1st-Timers
23
Maintainers
B TierTypeScript 3.0k

off-grid-ai/OGAM

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.

73.8%
Merge Rate
10d
First Review
71%
1st-Timers
9
Maintainers
B TierSwift 190

Eric-Terminal/ETOS-LLM-Studio

A native LLM client for iOS & Apple Watch. Run local GGUF models offline via llama.cpp, or connect to OpenAI/Claude/Gemini. Features local RAG, Model Context Protocol (MCP) tools, Siri Shortcuts, and cross-device sync. Built with Swift.

50.0%
Merge Rate
2d
First Review
100%
1st-Timers
2
Maintainers
B TierGo 286

gpustack/gguf-parser-go

Review/Check GGUF files and estimate the memory usage and maximum tokens per second.

54.5%
Merge Rate
10h
First Review
0%
1st-Timers
2
Maintainers
C TierGo 267

raketenkater/ggrun

llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

20.0%
Merge Rate
3d
First Review
0%
1st-Timers
0
Maintainers
C TierRust 5.8k

Michael-A-Kuykendall/shimmy

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

25.0%
Merge Rate
4d
First Review
100%
1st-Timers
2
Maintainers
D TierShell 384

Entrpi/ds4-on-spark

Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC++ 118

Mobile-Artificial-Intelligence/llama_sdk

lcpp is a dart implementation of llama.cpp used by the mobile artificial intelligence distribution (maid)

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 498

AudarAI/Audar-ASR-V1

Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 136

julianmb/q38rocm

Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 114

duckyshell/ComfyUI-MiniMaxH3-Prompt-Writer

Multimodal MiniMax H3 prompt writer for ComfyUI with local and API providers.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC 399

quantumaikr/quant.cpp

LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 189

ferrumox/fox

A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 472

lone-cloud/gerbil

A desktop app for running Large Language Models locally.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Gguf Open Source Repositories & C-Rank™ | GetMerged