#speech (29 Repositories)
Ranked open-source repositories tagged with #speech, scored by pull request acceptance likelihood and maintainer engagement velocity.
10.8%
11.7h
29 repositories tagged #speech
praat/praat.github.io
Praat: Doing Phonetics By Computer
kadirnar/voicehub
VoiceHub: A Unified Inference Interface for TTS Models
jim60105/docker-whisperX
Dockerfile for WhisperX: Automatic Speech Recognition with Word-Level Timestamps and Speaker Diarization (Dockerfile, CI image build and test)
modelscope/modelscope
ModelScope: bring the notion of Model-as-a-Service to life.
microsoft/UniSpeech
UniSpeech - Large Scale Self-Supervised Learning for Speech
OpenBMB/VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
speechbrain/speechbrain.github.io
The SpeechBrain project aims to build a novel speech toolkit fully based on PyTorch. With SpeechBrain users can easily create speech processing systems, ranging from speech recognition (both HMM/DNN and end-to-end), speaker recognition, speech enhancement, speech separation, multi-microphone speech processing, and many others.
Rikorose/DeepFilterNet
Noise supression using deep filtering
AlphaAvatar/AlphaAvatar
A real-time interactive Omni Avatar built on LiveKit, which allows you to seamlessly integrate with any open source Avatar components (real-time model, visual, voice, memory, search, etc.).
pytorch/audio
Data manipulation and transformation for audio signal processing, powered by PyTorch
daniilrobnikov/vits2
VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design
Kyubyong/css10
CSS10: A Collection of Single Speaker Speech Datasets for 10 Languages
jianchang512/stt
Voice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
coqui-ai/TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
SuperKogito/SER-datasets
A collection of datasets for the purpose of emotion recognition/detection in speech.
snakers4/silero-vad
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
babysor/MockingBird
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
CokoIya/MioVRC_Translator
Mio VRC 语音转文字翻译转语音插件
Migushthe2nd/MsEdgeTTS
A simple Azure Speech Service module that uses the Microsoft Edge Read Aloud API. https://www.npmjs.com/package/msedge-tts
primaryobjects/voice-gender
Gender recognition by voice and speech analysis
khanhuitse05/speech-and-text-unity-ios-android
Speed to text in Unity iOS use Native Speech Recognition
HG-ha/MTools
MTools 是一个功能强大的多功能桌面应用程序,集成了音视频处理、图片编辑、文本操作和编码工具,内置AI增强功能。旨在简化您的工作流程,提升生产效率
DeutscheKI/tevr-asr-tool
State-of-the-art (ranked #1 Aug 2022) German Speech Recognition in 284 lines of C++. This is a 100% private 100% offline 100% free CLI tool.
Baidu-AIP/speech-vad-demo
集成Webrtc的VAD,用于切分音频文件
SWHL/AI-Competition-Collections
AI比赛经验帖子 & 训练和测试技巧帖子 集锦(收集整理各种人工智能比赛经验帖)
Graphi07/room-impulse-responses
A list of publicly available room impulse response datasets and scripts to download them.
yongxuUSTC/sednn
deep learning based speech enhancement using keras or pytorch, make it easy to use
echogarden-project/echogarden
Cross-platform speech toolset, used from the command-line or as a Node.js library. Includes a variety of engines for speech synthesis, speech recognition, forced alignment, speech translation, voice isolation, language detection and more.
IDEA-Research/Grounded-Segment-Anything
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything