Back to Topics Directory
Topic Hub

#kv-cache (8 Repositories)

Ranked open-source repositories tagged with #kv-cache, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

44.6%

Avg Review Latency

17.0h

Filter by language

8 repositories tagged #kv-cache

B TierRust 186

feichai0017/loom-infer

Rust-native GPU operator library for LLM inference, built with cuda-oxide

98.9%
Merge Rate
14h
First Review
100%
1st-Timers
1
Maintainers
A TierRust 192 5 GFIs

novitalabs/pegaflow

High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.

71.4%
Merge Rate
16h
First Review
33%
1st-Timers
4
Maintainers
A TierPython 195

syv-ai/qwen38-27b-rtx3090

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

62.5%
Merge Rate
3h
First Review
50%
1st-Timers
20
Maintainers
B TierC++ 233

alibaba/tair-kvcache

Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSim), and more.

70.3%
Merge Rate
3d
First Review
75%
1st-Timers
17
Maintainers
B TierPython 11.4k 2 GFIs

LMCache/LMCache

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

53.9%
Merge Rate
20h
First Review
43%
1st-Timers
121
Maintainers
D TierPython 479

huawei-csl/KVarN

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 185

feichai0017/orbitkv

OrbitKV: a Rust attention-state compiler and lifetime-safe KV block manager.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC 399

quantumaikr/quant.cpp

LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Kv-cache Open Source Repositories & C-Rank™ | GetMerged