Register
Login
Resources
Docs Blog Datasets Glossary Case Studies Tutorials & Webinars
Product
Data Engine LLMs Platform Enterprise
Pricing Explore
Connect to our Discord channel
55900c0e19
Initial commit
1 year ago
aa8bd3996b
update readme automation
1 year ago
Storage Buckets

README.md

You have to be logged in to leave a comment. Sign In

Voices Obscured in Complex Environmental Settings (VOiCES)

Stream data with DDA:

from dagshub.streaming import DagsHubFilesystem

fs = DagsHubFilesystem(".", repo_url="https://dagshub.com/DagsHub-Datasets/lab41-sri-voices-dataset")

fs.listdir("s3://lab41openaudiocorpus")

Description:

VOiCES is a speech corpus recorded in acoustically challenging settings, using distant microphone recording. Speech was recorded in real rooms with various acoustic features (reverb, echo, HVAC systems, outside noise, etc.). Adversarial noise, either television, music, or babble, was concurrently played with clean speech. Data was recorded using multiple microphones strategically placed throughout the room. The corpus includes audio recordings, orthographic transcriptions, and speaker labels.

Contact:

VOiCES is a speech corpus recorded in acoustically challenging settings, using distant microphone recording. Speech was recorded in real rooms with various acoustic features (reverb, echo, HVAC systems, outside noise, etc.). Adversarial noise, either television, music, or babble, was concurrently played with clean speech. Data was recorded using multiple microphones strategically placed throughout the room. The corpus includes audio recordings, orthographic transcriptions, and speaker labels.

Update Frequency:

Data from two additional rooms will be added to the corpus Fall 2018.

Managed By:

https://www.iqt.org/

Resources:

  1. resource:
    • Description: wav audio files, orthographic transcriptions, and speaker ID
    • ARN: arn:aws:s3:::lab41openaudiocorpus
    • Region: us-east-1
    • Type: S3 Bucket

Tags:

aws-pds, machine learning, automatic speech recognition, speaker identification, denoising, speech processing

Tutorials:

  1. tutorial:
Tip!

Press p or to see the previous file or, n or to see the next file

About

lab41-sri-voices-dataset is originate from the Registry of Open Data on AWS

Collaborators 5

Comments

Loading...