Search

Full catalogue 155 resources

Page 10 of 11

Abstracts

EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation

Julius Richter, Yi-Chiao Wu, Steven Krenn + 5 others

We release the EARS (Expressive Anechoic Recordings of Speech) dataset, a high-quality speech dataset comprising 107 speakers from diverse backgrounds, totalling in 100 hours of clean, anechoic speech data. The dataset covers a large range of different speaking styles, including emotional speech, different reading styles, non-verbal sounds, and conversational freeform speech. We benchmark...

View on www.isca-archive.org
EARS Dataset

Expressive Anechoic Recordings of Speech (EARS). Highlights: - 100 h of speech data from 107 speakers - high-quality recordings at 48 kHz in an anechoic chamber - high speaker diversity with speakers from different ethnicities and age range from 18 to 75 years - full dynamic range of human speech, ranging from whispering to yelling - 18 minutes of freeform monologues per speaker - sentence...

View on github.com
Pyannote

Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding

View on github.com
Silero-vad

Silero VAD: pre-trained enterprise-grade Voice Activity Detector

View on github.com
SpeechBrain

Mirco Ravanelli, Titouan Parcollet, Peter Plantinga + 18 others

SpeechBrain is an open-source PyTorch toolkit that accelerates Conversational AI development, i.e., the technology behind speech assistants, chatbots, and large language models. It is crafted for fast and easy creation of advanced technologies for Speech and Text Processing.

View on github.com
Opensmile

openSMILE (open-source Speech and Music Interpretation by Large-space Extraction) is a complete and open-source toolkit for audio analysis, processing and classification especially targeted at speech and music applications, e.g. automatic speech recognition, speaker identification, emotion recognition, or beat tracking and chord detection.

View on github.com
Whisper

Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification.

View on github.com
WhisperX

Max Bain

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

View on github.com
CrisperWhisper

Laurin Wagner,

Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection

View on github.com
VoiceSauce

Y.-L. Shue

VoiceSauce is an application, implemented in Matlab, which provides automated voice measurements over time from audio recordings. Inputs are standard wave (*.wav) files and the measures currently computed are: F0 Formants F1-F4 H1(*) H2(*) H4(*) A1(*) A2(*) A3(*) 2K(*) 5K H1(*)-H2(*) H2(*)-H4(*) H1(*)-A1(*) H1(*)-A2(*) H1(*)-A3(*) H4(*)-2K(*) 2K(*)-5K Energy Cepstral Peak Prominence Harmonic...

View on phonetics.ucla.edu
SoX - Sound eXchange

SoX is the Swiss Army Knife of sound processing utilities. It can convert audio files to other popular audio file types and also apply sound effects and filters during the conversion.

View on sourceforge.net
FFmpeg

A complete, cross-platform solution to record, convert and stream audio and video.

View on www.ffmpeg.org
ALICE: An open-source tool for automatic measurement of phoneme, syllable, and word counts from child-centered daylong recordings

Okko Räsänen, Shreyas Seshadri, Marvin Lavechin + 2 others

Recordings captured by wearable microphones are a standard method for investigating young children’s language environments. A key measure to quantify from such data is the amount of speech present in children’s home environments. To this end, the LENA recorder and software—a popular system for measuring linguistic input—estimates the number of adult words that children may hear over the course...

View on doi.org
ALICE

orasanen

Automatic LInguistic Unit Count Estimator (ALICE). ALICE is a tool for estimating the number of adult-spoken linguistic units from child-centered audio recordings, as captured by microphones worn by children. It is meant as an open-source alternative for LENA adult word count (AWC) estimator [1]. ALICE uses SylNet [2] for feature extraction and voice type classifier [3] for broad-class...

View on github.com
Praat: doing Phonetics by Computer

Paul Boersma, David Weenink

View on www.fon.hum.uva.nl

Page 10 of 11

Main feed

Last update from database: 05/07/2026, 04:10 (UTC)

Search

Full catalogue 155 resources

Explore

Audio Data

Derived & Measured Data

Speech Perception Data

Speech Production Data

Teaching Resources

Software, Processing & Utilities

Tags

Resource type