#asr (30 Repositories)
Ranked open-source repositories tagged with #asr, scored by pull request acceptance likelihood and maintainer engagement velocity.
53.5%
53.5h
30 repositories tagged #asr
QuintinShaw/openasr
Local-first speech-to-text: no cloud, no telemetry, fail-closed by design. One CLI, seven model families, signed model catalog, OpenAI-compatible local API.
modelscope/FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
soniqo/speech-android
On-device speech SDK for Android — ASR, TTS, VAD, and noise cancellation powered by ONNX Runtime with Qualcomm NNAPI acceleration
Picovoice/cheetah
On-device streaming speech-to-text engine powered by deep learning
drakulavich/kesha-voice-kit
Give your tools a voice — speech to text and back, 25 languages, up to ~19× faster than Whisper. On your machine.
n0an/VivaDicta
iOS & watchOS speech-to-text app with AI voice keyboard, on-device RAG, and chat with your notes - powered by Apple Foundation Models, WhisperKit, NVIDIA Parakeet, and 20+ AI providers
jim60105/docker-whisperX
Dockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization (Dockerfile, CI image build and test)
soniqo/speech-swift
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
hehehai/voxt
🎙️ An intelligent voice productivity assistant that turns speech into clean text, useful actions, and structured knowledge. It helps users capture ideas, communicate naturally, automate repetitive tasks, and stay productive across different apps and workflows.
Open-Less/openless
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)
izwi-ai/izwi
Voice AI runtime. Local first transcription, speaker diarization, TTS, and voice cloning with an OpenAI compatible API.
Kieirra/murmure
Fully local, private and cross platform Speech-to-Text with LLM Post-processing
sgl-project/sglang-omni
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
NVIDIA-NeMo/Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
handy-computer/transcribe.cpp
ggml speech-to-text inference for 16+ model families
OpenBMB/UltraEval-Audio
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation
Quantatirsk/qwen3-asr
All in one Qwen3-ASR Server, compatible with OpenAI API
istupakov/onnx-asr
A lightweight Python package for Automatic Speech Recognition using ONNX models
TheDeathDragon/LiveTranslate
Real-time audio translation, captures system audio + mic, runs ASR (Whisper/SenseVoice), translates via LLM API with streaming display. Perfect for VTubers, livestreamers, and watching foreign content. Windows 实时音频翻译,ASR 语音识别后 LLM 流式翻译显示,适合 VTuber、主播和外语视频观看。
mkiol/dsnote
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
QwenAudio/SenseVoice
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
CheshireCC/faster-whisper-GUI
faster_whisper GUI with PySide6
AudarAI/Audar-ASR-V1
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
PiSugar/whisplay-ai-chatbot
Pocket-sized AI chatbot built using a RPI Zero 2w / 5
amicalhq/amical
🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate.
zenstory-ai/video-recap-skills
Clip any video into a narration recap with claude code skill|用 claude code skill 把任何视频剪辑成中文解说视频,支持剪映导出
alphacep/vosk-api
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
Picovoice/leopard
On-device speech-to-text engine powered by deep learning
lihaiya/freeipcc
Call Center,Contact Center,AI,呼叫中心,客服系统,工单系统,智能外呼,大模型呼叫中心,智能呼叫中心,FreeSWITCH
SergeyShk/Speech-to-Text-Russian
Проект для распознавания речи на русском языке на основе pykaldi.