A PyTorch-based Speech Toolkit
-
Updated
Aug 27, 2026 - Python
A PyTorch-based Speech Toolkit
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
Foundation Architecture for (M)LLMs
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
AI powered speech denoising and enhancement
WaveNet vocoder
Controllable and fast Text-to-Speech for over 7000 languages!
PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
General Speech Restoration
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
SincNet is a neural architecture for efficiently processing raw audio samples.
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
Tensorflow 2.x implementation of the DTLN real time speech denoising model. With TF-lite, ONNX and real-time audio processing support.
A neural network for end-to-end speech denoising
PyTorch implementation of "FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement."
UniSpeech - Large Scale Self-Supervised Learning for Speech
🔉 spafe: Simplified Python Audio Features Extraction
A python wrapper for Speech Signal Processing Toolkit (SPTK).
Problem Agnostic Speech Encoder
To associate your repository with the speech-processing topic, visit your repo's landing page and select "manage topics."