Key takeaway: Downloading your raw DNA data opens up a world of third-party analysis tools — but these are for research and education, not medical diagnosis. Approximately 40% of rare clinically-actionable variants flagged by DTC genotyping chips are false positives (Genetics in Medicine, 2020). Always confirm significant findings with a CLIA-certified clinical lab before acting on them.
Step 1: Download Your Raw Data
Every major DNA testing company allows you to download your raw genotype data, though the process varies. Here's how to get your raw data from the most common providers:
| Provider | Download Location | File Format | Approximate Variants |
|---|---|---|---|
| 23andMe | Settings → 23andMe Data → Download | .txt (TSV) or .zip | ~600K-700K SNPs |
| AncestryDNA | Settings → DNA Settings → Download | .txt (TSV) or .zip | ~700K SNPs |
| FamilyTreeDNA | MyDNA → Data Download | .csv | ~700K SNPs |
| MyHeritage | DNA → Manage DNA Kits → Download | .csv | ~700K SNPs |
| WGS (Nebula, Dante) | Provider portal → Download raw data | VCF, FASTQ, or BAM | ~3M+ variants |
File size warning: Raw data from genotyping chips (23andMe, AncestryDNA) is typically 10-25 MB. Whole genome sequencing (WGS) files can be 50-100 GB uncompressed for FASTQ and 30-50 GB for BAM files. Ensure you have adequate storage and a stable internet connection before downloading WGS data. Many WGS providers ship data on physical hard drives for this reason.
Step 2: Understand Your File Format
The file you download contains your genetic variants in a structured format:
- Genotype .txt/.csv files (23andMe, AncestryDNA): Each line contains an rsID (the variant identifier), the chromosome position, and your genotype (two alleles — e.g., A/A, A/G, or G/G). A typical line looks like:
rs429358 19 45411941 A G— meaning at position 45411941 on chromosome 19, your genotype is A/G (one copy of each variant). 23andMe files include roughly 600,000 to 700,000 SNPs. - VCF (Variant Call Format) files: The standard format for whole genome sequencing data, containing not just your genotype at each position but also quality metrics, read depth, and variant annotation. A single VCF file from a 30x WGS contains approximately 3 million single nucleotide variants, plus structural variants, insertions, and deletions. VCF files are significantly larger and more complex than genotype files.
- FASTQ files: Raw sequencing reads — the actual nucleotide sequences output by the sequencer before any analysis or variant calling. These are the largest files (often 50-100 GB per sample for WGS) and require bioinformatics tools to process. Most users will never work with FASTQ files directly.
Step 3: Choose Your Analysis Tools
Once you have your raw data file, these are the most useful third-party tools for exploration:
| Tool | Price | Best For | Limitations |
|---|---|---|---|
| Promethease | $12 | Literature-backed variant reports with SNPedia links | Can be overwhelming; categorizes every SNP as good/bad, causing unnecessary anxiety for benign variants |
| Genetic Genie | Free | Methylation and detox pathway analysis (MTHFR, COMT, etc.) | Narrow focus — only analyzes ~50-100 genes in specific pathways |
| Codegen | Free (basic) | Variant browser with ClinVar, dbSNP, gnomAD integration | Requires some genetic literacy to interpret results |
| OpenSNP | Free | Open-source data sharing and community analysis | Data is public — privacy implications of sharing |
| Gene.iobio | Free | Interactive visual variant analysis (supports VCF) | Designed for clinical geneticists; steep learning curve |
Promethease is the most popular third-party DNA analysis tool. For $12, it generates a detailed report that compares every variant in your raw data against the SNPedia database — a community-curated wiki of genetic variant research. Each variant receives a magnitude score (0-10) reflecting the strength of evidence linking it to a trait or disease. Promethease is excellent for exploratory learning but has a well-known flaw: it reports every statistically significant research finding for every SNP, including weak associations and conflicting results, which can create unnecessary alarm.
The 40% False Positive Problem — And How to Protect Yourself
A landmark study published in Genetics in Medicine (Tandy-Connor et al., 2020) sent 49 DTC test samples to a clinical lab for confirmation of raw data findings. The result: 40% of variants flagged as clinically significant by DTC raw data were false positives — meaning the variant wasn't actually present in the individual's DNA. This finding has been replicated in subsequent studies and is widely cited in genetic counseling guidelines.
The false positives arise because DTC genotyping chips use SNP probes — short DNA sequences designed to bind to specific positions in the genome. At rare variants, these probes can bind imperfectly or fail to distinguish between closely related sequences, producing incorrect calls. The overall genotyping error rate is low (~0.03%), but for rare variants — the very ones most likely to be medically significant — the false positive rate is dramatically higher.
To protect yourself:
- Never change medications or make medical decisions based on DTC raw data analysis alone
- If you find a potentially significant variant, have it confirmed by a CLIA-certified clinical lab
- Consult a genetic counselor (find one at nsgc.org) who can help interpret results and order appropriate confirmatory testing
- Consider that WGS from a clinical-grade provider offers significantly lower error rates and includes variant quality scores that let you assess reliability
How to Look Up an rsID
Every SNP in your raw data has an rsID (Reference SNP cluster ID) — a unique identifier maintained by the NCBI's dbSNP database. Here's how to research any rsID you find:
- SNPedia (snpedia.com) — Search your rsID for a plain-language summary of what studies have found about this variant. SNPedia is community-maintained and cites its sources. Example: searching rs429358 tells you it's the APOE ε4 allele, associated with increased Alzheimer's risk.
- ClinVar (ncbi.nlm.nih.gov/clinvar) — Check the clinical significance classification. ClinVar aggregates variant interpretations from clinical labs and expert panels. Look for variants classified as "Pathogenic" or "Likely Pathogenic" with a review status of at least "reviewed by expert panel."
- gnomAD (gnomad.broadinstitute.org) — Check the population frequency. If a variant labeled "pathogenic" appears in 10% of the population, it's almost certainly benign — truly pathogenic variants are rare in healthy populations.
Pro tip — population frequency as a sanity check: Truly devastating genetic variants rarely appear at high frequency in the general population because they reduce fitness. If ClinVar says "pathogenic" but gnomAD shows the variant at 5% population frequency, the classification is likely incorrect. This is one of the most useful heuristics for spot-checking variant interpretations.
Limitations of Third-Party DNA Analysis
Even with proper tools, raw data analysis has fundamental limitations:
- Coverage gaps: 23andMe's ~600K SNPs represent roughly 0.02% of the 3 billion base pairs in your genome. You're looking at a tiny slice of your DNA. If a variant isn't on the chip, you won't see it — even if it's medically critical.
- No structural variants: Genotyping chips cannot detect large deletions, duplications, or chromosomal rearrangements. A pathogenic BRCA1 deletion, for example, would be completely invisible.
- Ancestry bias: Most SNP associations were discovered in European populations. A variant that's pathogenic in one population may be benign in another, and vice versa. Third-party tools rarely account for this.
- Polygenic traits: Most common diseases are influenced by thousands of variants, each with tiny effects. Looking up individual SNPs cannot meaningfully predict your risk for conditions like heart disease, diabetes, or depression.
Ready for the full picture? Compare WGS providers
Whole genome sequencing reads all 3 billion base pairs — no coverage gaps, no false positives from probe errors.
Compare WGS Providers Browse DirectoryFrequently Asked Questions
Can I upload my raw DNA data to multiple sites safely?
Technically yes — your raw data is just a file, and you can upload copies to as many services as you want. The privacy risk is cumulative: each service you upload to becomes another entity that stores your genetic data, with its own security practices and data-sharing policies. Read each site's privacy policy and terms of service carefully. Some research-oriented sites (like OpenSNP) make your data publicly available by design. If you're concerned about privacy, prefer tools that process your data locally in your browser rather than uploading it to a remote server.
What's the difference between a VCF, FASTQ, and BAM file?
FASTQ files contain raw sequencing reads — the actual A/C/T/G sequences output by the sequencing machine, along with quality scores for each base. BAM files are aligned FASTQ reads mapped to a reference genome, showing where each read sits in the genome. VCF files are variant call files — the output of variant calling software that identifies where your DNA differs from the reference. For most users, the VCF is the most useful file: it contains the list of variants (SNPs, insertions, deletions) that differ from the reference genome, typically 3-5 million for a WGS sample.
Can I use my raw data to find unknown relatives?
Third-party tools like GEDmatch allow you to upload raw DNA data from any provider and compare it against others who have uploaded to the same database. GEDmatch gained widespread attention for its role in forensic genealogy (identifying the Golden State Killer), but it is primarily used for genealogy research. Important caveat: GEDmatch and similar services maintain databases that can be accessed by law enforcement in some circumstances, depending on your privacy settings. The default GEDmatch setting now excludes law enforcement matching — you must explicitly opt in.