Back to Topics Directory
Topic Hub

#quantization (26 Repositories)

Ranked open-source repositories tagged with #quantization, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

41.5%

Avg Review Latency

55.1h

Filter by language

26 repositories tagged #quantization

S TierRust 191

timtoole02/Camelid

Camelid: a Rust-native local inference backend with evidence-gated model compatibility.

86.7%
Merge Rate
23h
First Review
100%
1st-Timers
3
Maintainers
A TierPython 1.2k

ModelCloud/GPTQModel

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

94.5%
Merge Rate
14d
First Review
100%
1st-Timers
2
Maintainers
A TierPython 1.6k 3 GFIs

intel/auto-round

A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers|简洁且高效的量化工具包

77.2%
Merge Rate
2d
First Review
71%
1st-Timers
19
Maintainers
B TierC++ 258

VectorDB-NTU/RaBitQ-Library

An official lightweight library for the RaBitQ algorithm and its applications in vector search.

100.0%
Merge Rate
-
First Review
100%
1st-Timers
0
Maintainers
A TierPython 1.3k

open-edge-platform/geti

Build computer vision models in a fraction of the time and with less data.

89.6%
Merge Rate
8d
First Review
84%
1st-Timers
20
Maintainers
A TierRust 16.5k

RyanCodrai/turbovec

A vector index built on TurboQuant, written in Rust with Python bindings

85.9%
Merge Rate
2d
First Review
67%
1st-Timers
6
Maintainers
A TierRust 554 4 GFIs

warpfront/hipfire

RDNA-native LLM inference engine in Rust.

70.2%
Merge Rate
6d
First Review
61%
1st-Timers
12
Maintainers
A TierRust 207

giannisanni/pulsar

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.

83.3%
Merge Rate
10h
First Review
0%
1st-Timers
3
Maintainers
A TierPython 195

syv-ai/qwen38-27b-rtx3090

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

62.5%
Merge Rate
3h
First Review
50%
1st-Timers
20
Maintainers
B TierC++ 163

VinRobotics/vla.cpp

A unified inference runtime for VLA models.

100.0%
Merge Rate
6d
First Review
100%
1st-Timers
2
Maintainers
B TierPython 2.7k

intel/neural-compressor

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

75.9%
Merge Rate
9d
First Review
91%
1st-Timers
8
Maintainers
B TierPython 3.0k 2 GFIs

pytorch/ao

PyTorch native quantization for training and inference

53.3%
Merge Rate
1d
First Review
46%
1st-Timers
51
Maintainers
B TierPython 74.4k

hiyouga/LlamaFactory

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

38.8%
Merge Rate
2d
First Review
47%
1st-Timers
51
Maintainers
B TierPython 3.7k 9 GFIs

vllm-project/llm-compressor

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

38.0%
Merge Rate
7d
First Review
38%
1st-Timers
25
Maintainers
B TierPython 3.5k

huggingface/optimum

🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools

23.1%
Merge Rate
23h
First Review
0%
1st-Timers
7
Maintainers
D TierPython 2.7k

qualcomm/aimet

AIMET is a library that provides advanced quantization and compression techniques for trained neural network models.

0.0%
Merge Rate
<1h
First Review
0%
1st-Timers
0
Maintainers
D TierPython 18.9k

ymcui/Chinese-LLaMA-Alpaca

中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 312

Dicklesworthstone/franken_ocr

Pure-Rust, CPU-only OCR engine for Baidu Unlimited-OCR (a DeepSeek-OCR-derived 3B MoE VLM). Five-model zoo, custom int8 kernels, no ML framework, no Python, no GPU.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 2.9k

nunchaku-ai/ComfyUI-nunchaku

ComfyUI Plugin of Nunchaku

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC 399

quantumaikr/quant.cpp

LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression + single-header library.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJUJupyter Notebook 18.4k

UFund-Me/Qbot

[🔥updating ...] AI 自动量化交易机器人(完全本地部署) AI-powered Quantitative Investment Research Platform. 📃 online docs: https://ufund-me.github.io/Qbot ✨ :news: qbot-mini: https://github.com/Charmve/iQuant

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 420

FujitsuResearch/OneCompression

Python package for LLM compression

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 343

inisis/brocolli

Everything in Torch Fx

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 340

Mininglamp-AI/cider

W8A8/W4A8 inference + optimized SDPA on Apple Silicon — unlocking unused INT8 TensorOps in M5 for 1.2–1.9× faster LLM prefill, plus FlashInfer-inspired GQA decode attention for up to 1.6× SDPA speedup, built as MLX custom primitives.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 479

huawei-csl/KVarN

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC 339

cpldcpu/BitNetMCU

Neural Networks with low bit weights on low end 32 bit microcontrollers such as the CH32V003 RISC-V Microcontroller and others

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Quantization Open Source Repositories & C-Rank™ | GetMerged