Register
Login
Resources
Docs Blog Datasets Glossary Case Studies Tutorials & Webinars
Product
Data Engine LLMs Platform Enterprise
Pricing Explore
Connect to our Discord channel
33d6184ec5
Initial commit
1 year ago
2866d3039c
update readme automation
1 year ago
Storage Buckets
s3://aws-roda-ml-datalake-us-east-1/
s3://aws-roda-ml-datalake/yt8m/
s3://aws-roda-ml-datalake/yt8m_ods/

README.md

You have to be logged in to leave a comment. Sign In

YouTube 8 Million - Data Lakehouse Ready

Stream data with DDA:

from dagshub.streaming import DagsHubFilesystem

fs = DagsHubFilesystem(".", repo_url="https://dagshub.com/DagsHub-Datasets/yt8m-dataset")

fs.listdir("s3://aws-roda-ml-datalake/yt8m/")

Description:

This both the original .tfrecords and a Parquet representation of the YouTube 8 Million dataset. YouTube-8M is a large-scale labeled video dataset that consists of millions of YouTube video IDs, with high-quality machine-generated annotations from a diverse vocabulary of 3,800+ visual entities. It comes with precomputed audio-visual features from billions of frames and audio segments, designed to fit on a single hard disk. This dataset also includes the YouTube-8M Segments data from June 2019. This dataset is 'Lakehouse Ready'. Meaning, you can query this data in-place straight out of the Registry of Open Data S3 bucket. Deploy this dataset's corresponding CloudFormation template to create the AWS Glue Catalog entries into your account in about 30 seconds. That one step will enable you to interact with the data with AWS Athena, AWS SageMaker, AWS EMR, or join into your AWS Redshift clusters. More detail in (the documentation)[https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/README.md.

Contact:

This both the original .tfrecords and a Parquet representation of the YouTube 8 Million dataset. YouTube-8M is a large-scale labeled video dataset that consists of millions of YouTube video IDs, with high-quality machine-generated annotations from a diverse vocabulary of 3,800+ visual entities. It comes with precomputed audio-visual features from billions of frames and audio segments, designed to fit on a single hard disk. This dataset also includes the YouTube-8M Segments data from June 2019. This dataset is 'Lakehouse Ready'. Meaning, you can query this data in-place straight out of the Registry of Open Data S3 bucket. Deploy this dataset's corresponding CloudFormation template to create the AWS Glue Catalog entries into your account in about 30 seconds. That one step will enable you to interact with the data with AWS Athena, AWS SageMaker, AWS EMR, or join into your AWS Redshift clusters. More detail in (the documentation)[https://github.com/aws-samples/data-lake-as-code/blob/roda-ml/README.md.

Update Frequency:

Google Research has not updated the dataset since 2019.

Managed By:

https://aws.amazon.com/

Resources:

  1. resource:

  2. resource:

    • Description: Lakehouse ready YT8M as Glue Parquet files. Install instructions here.
    • ARN: arn:aws:s3:::aws-roda-ml-datalake/yt8m_ods/
    • Region: us-west-2
    • Type: S3 Bucket
  3. resource:

    • Description: Replica of the two locations above in us-east-1.
    • ARN: arn:aws:s3:::aws-roda-ml-datalake-us-east-1/
    • Region: us-east-1
    • Type: S3 Bucket

Tags:

computer vision, machine learning, labeled, parquet, video

Tutorials:

  1. tutorial:

Publication:

  1. publication:
Tip!

Press p or to see the previous file or, n or to see the next file

About

yt8m-dataset is originate from the Registry of Open Data on AWS

Collaborators 5

Comments

Loading...