#llamacpp (21 Repositories)
Ranked open-source repositories tagged with #llamacpp, scored by pull request acceptance likelihood and maintainer engagement velocity.
38.6%
28.2h
21 repositories tagged #llamacpp
xybrid-ai/xybrid
Cross-platform on-device AI toolkit
gptme/gptme
Your agent in your terminal, equipped with local tools: writes code, uses the terminal, browses the web. Make your own persistent autonomous agent on top!
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
llamastash/llamastash
Zero-overhead, terminal-native local-LLM runtime manager. Launches, supervises, and routes local models behind one OpenAI-compatible endpoint.
xorbitsai/inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
janhq/jan
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
containers/ramalama
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
cactus-compute/cactus
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
sybil-solutions/local-studio
Control panel for VLLM, Sglang, llama.cpp, exllamav3
mostlygeek/llama-swap
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
floneum/kalosm
Instant, controllable, local pre-trained AI models in Rust
raketenkater/ggrun
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
Michael-A-Kuykendall/shimmy
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
bloodworks-io/phlox
Open source, local first AI medical agent for desktop and web.
nekomeowww/ollama-operator
🚢 Yet another operator for running large language models on Kubernetes with ease. Powered by Ollama! 🐫
Mobile-Artificial-Intelligence/llama_sdk
lcpp is a dart implementation of llama.cpp used by the mobile artificial intelligence distribution (maid)
staghado/vit.cpp
Inference Vision Transformer (ViT) in plain C/C++ with ggml
InftyAI/llmaz
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
BrutalCoding/aub.ai
AubAI brings you on-device gen-AI capabilities, including offline text generation and more, directly within your app.
nerve-sparks/iris_android
IRIS is an android app for interfacing with GGUF / llama.cpp models locally.
lxe/llavavision
A simple "Be My Eyes" web app with a llama.cpp/llava backend