#flash-attention (5 Repositories)
Ranked open-source repositories tagged with #flash-attention, scored by pull request acceptance likelihood and maintainer engagement velocity.
32.0%
47.6h
5 repositories tagged #flash-attention
xlite-dev/ffpa-attn
Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.
NVIDIA/cudnn-frontend
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
NVlabs/rcm
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
Nathanw1014/strix-halo-llamacpp
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
HKUSTDial/flash-sparse-attention
Trainable fast and memory-efficient sparse attention