#cuda (30 Repositories)
Ranked open-source repositories tagged with #cuda, scored by pull request acceptance likelihood and maintainer engagement velocity.
88.4%
44.1h
30 repositories tagged #cuda
glotzerlab/fresnel
Publication quality path tracing in real time.
utensils/comfyui-nix
A slightly opinionated Nix flake for ComfyUI with curated custom nodes. Supports macOS (Apple Silicon) and Linux with CUDA.
brucefan1983/GPUMD
Graphics Processing Units Molecular Dynamics
mathiasbourgoin/Sarek
SIMT Abstractions for Runtime Extensible Kernels (GPGPU programing with OCaml)
Zaneham/Booth
Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.
roflcoopter/viseron
Self-hosted, local only NVR and AI Computer Vision software. With features such as object detection, motion detection, face recognition and more, it gives you the power to keep an eye on your home, office or any other place you want to monitor.
Blackwellboy/model-serving-minefield
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
unilabsim/UniLab
UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
QMCPACK/qmcpack
Main repository for QMCPACK, an open-source production level many-body ab initio Quantum Monte Carlo code for computing the electronic structure of atoms, molecules, and solids with full performance portable GPU support
ultralytics/inference
High-performance Ultralytics YOLO inference in Rust with ONNX Runtime, GPU backends, CLI, and WebGPU/WASM.
NVIDIA/skills
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
MrNeRF/LichtFeld-Studio
Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.
LuisaGroup/LuisaCompute
High-Performance Rendering Framework on Stream Architectures
NVIDIA/cccl
CUDA Core Compute Libraries
gpustack/gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
luigifcruz/CyberEther
High-performance GPU-accelerated signal processing and visualization framework that runs anywhere.
inclusionAI/Awex
A high-performance RL training-inference weight synchronization framework, designed to enable second-level parameter updates from training to inference in RL workflows
avifenesh/memra
Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai
hybridgroup/yzma
Go with your own intelligence - Go applications that directly integrate llama.cpp for local inference using hardware acceleration.
tracel-ai/cubek
CubeK: high-performance multi-platform kernels in CubeCL
invergent-ai/surogate
Training/Fine-tuning at the speed of light
debpalash/VoiceStudio
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
uccl-project/uccl
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
llnl/blt
A streamlined CMake build system foundation for developing HPC software
celeritas-project/celeritas
Celeritas is a new Monte Carlo transport code designed to accelerate scientific discovery in high energy physics by improving detector simulation throughput and energy efficiency using GPUs.
crazyguitar/cppcheatsheet
C/C++ Cheat Sheet
rapidsai/cugraph
cuGraph - RAPIDS Graph Analytics Library
shader-slang/slang
Making it easier to work with shaders
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.