1000 Genomes Dataset for Machine Learning

Install DagsHub:

pip install dagshub

Click on copy button to copy content

To stream this data directly on DagsHub

from dagshub.streaming import DagsHubFilesystem

fs = DagsHubFilesystem(".", repo_url="https://dagshub.com/DagsHub-Datasets/1000-genomes-dataset")

fs.listdir("s3://1000genomes")

Click on copy button to copy content

Description

The 1000 Genomes Project is an international collaboration which has established the most detailed catalogue of human genetic variation, including SNPs, structural variants, and their haplotype context. The final phase of the project sequenced more than 2500 individuals from 26 different populations around the world and produced an integrated set of phased haplotypes with more than 80 million variants for these individuals.

Explore this dataset on DagsHub

Additional information

Documentation

https://github.com/awslabs/open-data-docs/tree/main/docs/1000genomes

Update frequency

Not updated

Managed by

National Institutes of Health

License

Data from the 1000 Genomes Project is now available without embargo, following the final publication from the project. Use of the data should be cited in the usual way, with current details available at http://www.internationalgenome.org/faq/how-do-i-cite-1000-genomes-project.

Explore this dataset on DagsHub

1000 Genomes Dataset for Machine Learning

Install DagsHub:

To stream this data directly on DagsHub

Description

Additional information

Documentation

Update frequency

Managed by

License

Related datasets

Allen Brain Observatory – Visual Coding AWS Public Data Set

Allen Cell Imaging Collections

Biological and Physical Sciences (BPS) Microscopy Benchmark Training Dataset

Cancer Cell Line Encyclopedia (CCLE)

Launch your ML development to new heights with DagsHub

Take control of your multimodal data

ML Newsletter

1000 Genomes Dataset for Machine Learning

Install DagsHub:

To stream this data directly on DagsHub

Description

Additional information

Documentation

Update frequency

Managed by

License

Tags

Related datasets

Allen Brain Observatory – Visual Coding AWS Public Data Set

Allen Cell Imaging Collections

Biological and Physical Sciences (BPS) Microscopy Benchmark Training Dataset

Cancer Cell Line Encyclopedia (CCLE)

Launch your ML development to new heights with DagsHub

Take control of your multimodal data

ML Newsletter