Back to Topics Directory
Topic Hub

#moe (16 Repositories)

Ranked open-source repositories tagged with #moe, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

45.3%

Avg Review Latency

7.9h

Filter by language

16 repositories tagged #moe

B TierRust 328

avifenesh/memra

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

73.1%
Merge Rate
2h
First Review
100%
1st-Timers
0
Maintainers
A TierSwift 742

SharpAI/SwiftLM

⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.

86.0%
Merge Rate
6h
First Review
100%
1st-Timers
3
Maintainers
A TierPython 169

inclusionAI/Awex

A high-performance RL training-inference weight synchronization framework, designed to enable second-level parameter updates from training to inference in RL workflows

100.0%
Merge Rate
1h
First Review
100%
1st-Timers
2
Maintainers
A TierPython 694 5 GFIs

Tencent/YOLO-Master

[CVPR2026]🚀🚀🚀Official code for the paper "YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection." *(YOLO = You Only Look Once)* 🔥🔥🔥

73.4%
Merge Rate
22h
First Review
60%
1st-Timers
55
Maintainers
A TierRust 207

giannisanni/pulsar

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.

83.3%
Merge Rate
10h
First Review
100%
1st-Timers
3
Maintainers
A TierPython 915

NVIDIA/cudnn-frontend

cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

63.9%
Merge Rate
11h
First Review
70%
1st-Timers
26
Maintainers
A TierPython 32.9k 12 GFIs

sgl-project/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

49.1%
Merge Rate
5h
First Review
28%
1st-Timers
582
Maintainers
A TierPython 14.5k

NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

59.8%
Merge Rate
3h
First Review
49%
1st-Timers
222
Maintainers
A TierPython 6.2k 5 GFIs

flashinfer-ai/flashinfer

FlashInfer: Kernel Library for LLM Serving

57.4%
Merge Rate
21h
First Review
51%
1st-Timers
136
Maintainers
B TierKotlin 317

LISTEN-moe/android-app

Official LISTEN.moe Android app

58.3%
Merge Rate
2d
First Review
0%
1st-Timers
0
Maintainers
C TierGo 267

raketenkater/ggrun

llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

20.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierShell 384

Entrpi/ds4-on-spark

Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 111

guqiong96/Lsglang

Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parallel architecture, suitable for MOE model hybrid inference.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 354

zeraix/zeraix

Open-source local AI workspace — advancing on-device inference.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 437

lucienhuangfu/eLLM

eLLM: Run Long-Horizon Inference Faster on CPUs Than on GPUs

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierRust 312

Dicklesworthstone/franken_ocr

Pure-Rust, CPU-only OCR engine for Baidu Unlimited-OCR (a DeepSeek-OCR-derived 3B MoE VLM). Five-model zoo, custom int8 kernels, no ML framework, no Python, no GPU.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers