Results | YorVoice Catalogue

The Fake-or-Real (FoR) dataset is a collection of more than 195,000 utterances from real humans and computer generated speech. The dataset can be used to train classifiers to detect synthetic speech. The dataset aggregates data from the latest TTS solutions (such as Deep Voice 3 and Google Wavenet TTS) as well as a variety of real human speech, including the Arctic Dataset...

View on bil.eecs.yorku.ca

The M-AILABS Speech Dataset

I. Celeste Aurora Solak, Dima Naumov

The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis. Most of the data is based on LibriVox and Project Gutenberg. The training data consist of nearly thousand hours of audio and the text-files in prepared format. A transcription is provided for each clip. Clips vary in length...

View on github.com

ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

Ashi Garg, Zexin Cai, Henry Li Xinyuan + 5 others

This repository introduces: 🌀 ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts 🔥 Key Features 3000+ hours of synthetic speech Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including: 📖 Reading Style 🎙️ Podcast 🎥 YouTube 🗣️ Languages (Three different languages) 🌎 Demographics (including variations in age, accent, and gender) Multiple...

View on huggingface.co

CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems

Haibin Wu, Yuan Tseng, Hung-yi Lee

CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems TL;DR: We show that better detection of deepfake speech from codec-based TTS systems can be achieved by training models on speech re-synthesized with neural audio codecs. This dataset is released for this purpose. See our paper and Github for more details on using our...

View on huggingface.co

In-the-Wild: A Deepfake Detection Dataset

Nicolas M. Müller, Pavel Czempin, Franziska Dieckmann + 2 others

The In-the-Wild dataset contains real and synthetic speech recordings of 58 celebrities and politicians, collected from online videos. It provides a realistic benchmark for testing how well audio deepfake detection models generalize beyond laboratory data such as ASVspoof. Task: Audio Classification (Deepfake / Genuine) Languages: English Modality: Audio Size: 37.9 hours total 17.2 hours fake 20.7 hours real

View on huggingface.co

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

Wen Huang, Yanmei Gu, Zhiming Wang + 2 others

SpeechFake is a large-scale multilingual dataset for speech deepfake detection, featuring over 3 million fake samples across 46 languages. Generated using 30 diverse open-source models* spanning text-to-speech (TTS), voice conversion or clone (VC), and neural vocoder (NV) methods, it offers rich metadata and strong coverage of modern generation techniques, enabling robust and generalizable detection research.

View on github.com

LifeLUCID Corpus: Recordings of Speakers Aged 8 to 85 Years Engaged in Interactive Task in the Presence of Energetic and Informational Masking, 2017-2020

Outi Tuomainen, Linda Taschenberger, Valerie Hazan

This speech corpus contains recordings for 104 monolingual native southern British English speakers aged between 8 and 85 years old while they engaged in a problem-solving picture-based ‘spot the difference’ task (Diapix) with a conversational partner in four listening conditions. In NORM (quiet, no masking), participants heard each other normally. In SPSN (speech-shaped noise), participants...

View on reshare.ukdataservice.ac.uk

elderLUCID: London UCL Older Adults' clear speech in interaction database

Valerie Hazan, Outi Tuomainen, Jeesun Kim + 1 others

This collection contains the quantitative data resulting from the analysis of the elderLUCID audio corpus – a set of speech recordings collected for 83 adults aged 19 to 84 years inclusive. Recordings were made while participants carried out two types of collaborative tasks with a conversational partner who was a young adult of the same sex: (1) a ‘spot the difference’ picture task (‘diapix’)...

View on reshare.ukdataservice.ac.uk

kidLUCID

Fully-annotated corpus of spontaneous speech dialogues for children. Diapix task recorded as a stereo wav files with one speaker per channel. 96 children aged between 9 to 14 years old Non-bilingual native Southern British English speakers

View on speechbox.linguistics.northwestern.edu

Nijmegen Corpus of Casual Czech

The Nijmegen Corpus of Casual Czech contains 30 hours of high-quality recordings featuring 60 Czech speakers conversing among friends. The speech has been orthographically transcribed.

View on mirjamernestus.nl

Nijmegen Corpus of Casual French

The Nijmegen Corpus of Casual French contains 35 hours of high-quality recordings featuring 46 French speakers conversing among friends. The speech has been orthographically annotated by professional transcribers.

View on mirjamernestus.nl

Nijmegen Corpus of Casual Spanish

The Nijmegen Corpus of Casual Spanish contains around 30 hours of high-quality recordings featuring 52 Spanish speakers from Madrid conversing among friends. The speech has been orthographically annotated by professional transcribers.

View on mirjamernestus.nl

Nijmegen Corpus of Spanish English

The Nijmegen Corpus of Spanish English (NCSE) contains 38.5 hours of high-quality recordings of English speech produced by 34 native Spanish speakers in interaction with two native Dutch confederates. The NCSE contains a formal and an informal recording for each Spanish speaker. The speech has been orthographically transcribed.

View on mirjamernestus.nl

High quality TTS data for Bengali languages

Multi-speaker TTS data for Bangladesh Bengali (bn-BD) and Indian Bengali (bn-IN).

View on openslr.org

High quality TTS data for four South African languages

Multi-speaker TTS data for four South African languages, Afrikaans, Sesotho, Setswana and isiXhosa. This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In each folder, the file line_index.tsv contains a FileID, which in turn contains the UserID and the...

View on openslr.org

Your search

Results 78 resources

Explore

Audio Data

Derived & Measured Data

Speech Production Data

Teaching Resources

Tags

Resource type