Register
Login
Resources
Docs Blog Datasets Glossary Case Studies Tutorials & Webinars
Product
Data Engine LLMs Platform Enterprise
Pricing Explore
Connect to our Discord channel
9f1b2928d6
Initial commit
1 year ago
3ad690d70f
update readme automation
1 year ago
Storage Buckets

README.md

You have to be logged in to leave a comment. Sign In

1000 Genomes Phase 3 Reanalysis with DRAGEN 3.5 - Data Lakehouse Ready

Stream data with DDA:

from dagshub.streaming import DagsHubFilesystem

fs = DagsHubFilesystem(".", repo_url="https://dagshub.com/DagsHub-Datasets/1000-genomes-data-lakehouse-ready-dataset")

fs.listdir("s3://aws-roda-hcls-datalake/thousandgenomes_dragen")

Description:

The 1000 Genomes Project is an international collaboration which has established the most detailed catalogue of human genetic variation, including SNPs, structural variants, and their haplotype context. There were a total of 3202 individuals sequenced as part of Phase 3 of this project. The high coverage samples were processed using the Illumina DRAGEN v3.5.7b pipeline and are available at s3://1000genomes-dragen/. This dataset contains the VCFs transformed to Parquet/ORC in 3 different schemas - partitioned by samples, partitioned by chromosome and a nested data format. These representations of the 1000 Genomes DRAGEN data are stored in Parquet/ORC format and can be queried through Amazon Athena. To add these tables to your Glue Data Catalog and for sample queries on this dataset, please refer to the link in our Documentation.

Contact:

The 1000 Genomes Project is an international collaboration which has established the most detailed catalogue of human genetic variation, including SNPs, structural variants, and their haplotype context. There were a total of 3202 individuals sequenced as part of Phase 3 of this project. The high coverage samples were processed using the Illumina DRAGEN v3.5.7b pipeline and are available at s3://1000genomes-dragen/. This dataset contains the VCFs transformed to Parquet/ORC in 3 different schemas - partitioned by samples, partitioned by chromosome and a nested data format. These representations of the 1000 Genomes DRAGEN data are stored in Parquet/ORC format and can be queried through Amazon Athena. To add these tables to your Glue Data Catalog and for sample queries on this dataset, please refer to the link in our Documentation.

Update Frequency:

Not updated

Managed By:

https://aws.amazon.com/

Resources:

  1. resource:
    • Description: Parquet representations of 1000 Genomes VCF outputs from DRAGEN, ready for enrollment into Data Lake as Code.
    • ARN: arn:aws:s3:::aws-roda-hcls-datalake/thousandgenomes_dragen
    • Region: us-east-1
    • Type: S3 Bucket

Tags:

biology, bioinformatics, genetic, genomic, Homo sapiens, life sciences, parquet, population genetics, vcf

Tutorials:

  1. tutorial:
Tip!

Press p or to see the previous file or, n or to see the next file

About

1000-genomes-data-lakehouse-ready-dataset is originate from the Registry of Open Data on AWS

Collaborators 5

Comments

Loading...