Functional Annotation of Genetic Variants with ENCODE Data
When to Use
- User wants to annotate genetic variants with ENCODE regulatory element overlap and functional evidence
- User asks about "variant annotation", "regulatory variants", "non-coding variants", or "variant prioritization"
- User needs to assess whether a variant falls in an active enhancer, promoter, or TF binding site
- User wants to build a multi-evidence variant interpretation combining ENCODE + ClinVar + gnomAD
- Example queries: "annotate my GWAS hits with ENCODE regulatory data", "does this variant disrupt a TF binding site?", "prioritize non-coding variants by regulatory impact"
Interpret non-coding genetic variation by layering ENCODE functional genomics annotations to identify causal regulatory variants and link them to target genes.
Scientific Rationale
The question: "Which of my GWAS/eQTL variants actually disrupt regulatory elements, and what genes do they affect?"
Over 90% of disease-associated variants from GWAS fall in non-coding regions of the genome. Without functional annotation, a GWAS locus is just a genomic coordinate — it does not tell you which variant is causal, what regulatory element it disrupts, or which gene it affects. ENCODE provides the richest catalog of functional elements for interpreting these variants.
The Core Challenge
A typical GWAS locus contains dozens to hundreds of variants in linkage disequilibrium (LD) with the lead SNP. The causal variant(s) may not be the one with the strongest association. Functional annotation helps distinguish causal from tag variants by asking: does this variant overlap a regulatory element that is active in disease-relevant tissue?
The ENCODE Solution: Candidate cis-Regulatory Elements (cCREs)
The ENCODE Phase 3 project (ENCODE Project Consortium 2020, Nature, ~1,656 citations) established a registry of 926,535 human candidate cis-regulatory elements (cCREs) covering 7.9% of the genome. These are classified into:
| cCRE Class | Abbreviation | Definition | Count (human) |
|---|---|---|---|
| Promoter-like | PLS | DNase + H3K4me3 ± H3K27ac near TSS | ~34,000 |
| Proximal enhancer-like | pELS | DNase + H3K27ac within 2kb of TSS | ~46,000 |
| Distal enhancer-like | dELS | DNase + H3K27ac >2kb from TSS | ~670,000 |
| CTCF-only | CTCF-only | DNase + CTCF, no H3K4me3/H3K27ac | ~83,000 |
| DNase-H3K4me3 | DNase-H3K4me3 | DNase + H3K4me3, not near TSS | ~93,000 |
These cCREs are accessible via the SCREEN web interface and provide the foundation for variant annotation.
Literature Support
- ENCODE Project Consortium 2020 (Nature, ~1,656 citations): Registry of 926,535 human cCREs. The primary resource for variant-to-regulatory-element annotation. DOI
- Boyle et al. 2012 (Genome Research, ~2,623 citations): RegulomeDB — integrates ENCODE ChIP-seq, DNase-seq, eQTLs to score non-coding variants on a 1–6 scale. DOI
- Kircher et al. 2014 (Nature Genetics, ~5,719 citations): CADD — Combined Annotation-Dependent Depletion. SVM trained on 14.7M simulated vs. high-frequency alleles, pre-computes scores for all 8.6B possible human SNVs. DOI
- Rentzsch et al. 2019 (Nucleic Acids Research, ~1,500 citations): CADD v1.4/v1.6 update with GRCh38 support. DOI
- Finucane et al. 2015 (Nature Genetics, ~2,253 citations): Stratified LD Score Regression (S-LDSC) — partitions SNP heritability into functional categories using GWAS summary statistics. Foundational for quantifying how much disease heritability is attributable to ENCODE-defined regulatory elements. DOI
- Iotchkova et al. 2019 (Nature Genetics, ~157 citations): GARFIELD — GWAS enrichment in regulatory annotations with LD correction, allele frequency, and TSS distance confounders. DOI
- Wang et al. 2020 (JRSS-B, ~500 citations): SuSiE — Sum of Single Effects model for Bayesian fine-mapping, outputs credible sets per independent signal. DOI
- Weissbrod et al. 2020 (Nature Genetics, ~200 citations): PolyFun — leverages ENCODE/Roadmap functional annotations as prior causal probabilities for fine-mapping. PolyFun+SuSiE identified 32% more causal variants than SuSiE alone. DOI
- Nasser et al. 2021 (Nature, ~468 citations): ABC model enhancer-gene maps in 131 cell types. Links 5,036 GWAS signals to 2,249 genes. >20-fold enrichment of causal variants in cell-type-specific predicted enhancers. DOI
- Broekema et al. 2020 (Open Biology, ~107 citations): Practical review of post-GWAS fine-mapping and gene prioritization strategies. DOI
- Schaub et al. 2012 (Genome Research, ~675 citations): Early framework linking ENCODE functional data with GWAS disease associations. Functional annotations for up to 80% of reported associations. DOI
- Maurano et al. 2012 (Science, ~2,800 citations): Systematic demonstration that disease-associated variants concentrate in DNase I hypersensitive sites. 88% of variant-containing DHSs are active during fetal development. Enabled de novo identification of pathogenic cell types. DOI
- Amemiya et al. 2019 (Scientific Reports, ~1,372 citations): ENCODE Blacklist — regions producing artifactual signal in functional genomics. Must filter before annotation. DOI
Step 1: Define the Variant Set and Disease Context
Before any annotation, clarify the input:
Variant Set Types
| Input Type | Description | LD Consideration |
|---|---|---|
| GWAS lead SNPs | Top associations per locus | Must expand to LD proxies (r² > 0.8) |
| Fine-mapped credible sets | Post-fine-mapping variants with posterior probabilities | Already LD-aware — annotate directly |
| eQTL variants | Expression-associated SNPs | May need LD expansion depending on source |
| Rare variants (ClinVar) | Pathogenic/likely pathogenic | No LD concern — annotate directly |
| Candidate variants | User-curated list | Clarify whether LD expansion is needed |
LD Awareness is Critical
A GWAS lead SNP is NOT necessarily causal. It is the variant with the strongest statistical signal, but the true causal variant may be any of the dozens to hundreds of SNPs in LD. Before annotation:
- If the user provides lead SNPs only: recommend expanding to LD proxies (r² ≥ 0.8 in the relevant population) using LDlink, LDproxy, or 1000 Genomes
- If the user provides credible sets from fine-mapping: no expansion needed — these already account for LD
- If the user provides a single variant of known interest: proceed directly
Step 2: Map Disease to Relevant Tissues
ENCODE regulatory elements are highly tissue-specific. An enhancer active in liver may be completely silent in brain. The disease context determines which ENCODE data to query.
Common Disease-Tissue Mappings
| Disease Category | Primary Tissues | Key ENCODE Biosamples |
|---|---|---|
| Type 2 diabetes | Pancreas, liver, adipose, muscle | pancreas tissue, HepG2, adipose tissue |
| Alzheimer's disease | Brain (hippocampus, cortex) | brain tissue, astrocytes, neurons |
| Cardiovascular | Heart, blood vessels | heart tissue, HUVEC, aorta |
| Blood disorders | Blood, bone marrow | K562, GM12878, CD34+ cells |
| Autoimmune | Immune cells, thymus | GM12878, Treg, Th17, monocytes |
| Cancer | Tissue of origin | K562 (CML), HepG2 (liver), MCF-7 (breast), A549 (lung) |
| Inflammatory bowel | Intestine, colon | intestin |