Methods · Validation
Every inference has a method
The public ledger for Haeckel analyzers: named methods, reference panels, confidence gates, and known limits.
1000 Genomes + HGDP reference cohorts · Pipeline v7
The ledger
What runs, and what backs it
Research enters here only as executable method, with a named analyzer, a defined scope, and a stated uncertainty model.
Ancestry inference
26,000 independent AIMs across 62 sub-populations from 1000 Genomes and HGDP. Signal below threshold returns Unassigned.
Polygenic risk scoring
The live pipeline uses effect-allele harmonization and build-aware weighted scoring for six traits. PRS-CS, LDpred2, Lassosum, SBayesR, and C+T are implemented research engines, not live runtime methods.
Recursive phylogenetic traversal + back-mutation handling + Bayesian confidence
Y-DNA + mtDNA haplogroups (165 nodes)
Lineage calls are annotated with historical figures, migration context, and confidence against random expectation.
Archaic introgression
The live pipeline uses a coverage-gated Neanderthal and Denisovan marker panel. S*, ABBA-BABA, and tract-HMM modules exist as research code but are not wired into the active upload pipeline.
HIrisPlex-S · GIANT height · per-SNP contribution tracking[7]
Phenotype prediction
Eye, hair, skin, and height predictions with coverage gates, MC1R missingness protection, and per-SNP explainability.
CPIC star-allele calling · activity scores · FDA Black Box alerts[11]
Pharmacogenomics & safety
Star-allele calls and CPIC-guided interpretation across 12 pharmacogenes, with HLA and G6PD safety signals and boxed-warning context surfaced when the required variants are present.
Withdrawn legacy selected-variant projection · no active compute
Offspring and embryo modeling
Partner compatibility and offspring trait distributions across complex polygenic phenotypes. Sibling comparison preserves the spread and does not produce a ranking; embryo-level prediction is unvalidated and Haeckel does not offer embryo selection.
Quality control
Every upload is screened before analysis; low-confidence data is surfaced as missing or blocked, not guessed.
KING-robust kinship · ROH detection[8]
Relationship inference
Kinship, consanguinity, and runs of homozygosity for family structure and inherited-risk context.
pgvector 1536-d embeddings
Networks and candidate discovery
Vectorized genomic, phenotype, and profile embeddings for kinship-aware search and matching.
Evo 2 DNA foundation model (in integration)[13]
DNA foundation-model interpretation
Research integration for variants of uncertain significance. It is not part of the active upload pipeline or a live clinical interpretation path.
Rigor
Validation & confidence
Haeckel applies each module's defined coverage and confidence rules, and keeps unsupported results unavailable rather than estimating them.
Ancestry
Significance-gated
Bootstrap Wald; signal below the threshold is returned as Unassigned
AMR deconvolution
Admixed reference orthogonalized to its Native American vertex
Reduces spurious cross-population signal in unadmixed individuals
Phenotype
Coverage-gated (HIrisPlex-S)
A 67% coverage threshold protects against systematic MC1R missingness bias
Haplogroups
Tree integrity gate at module load
Back-mutations and unauthored child nodes are handled before calls ship
Offspring modeling
Probability distributions, not certainties
Sibling comparisons preserve uncertainty instead of collapsing to a rank alone
Upload QC
Call rate · heterozygosity · Ti/Tv · chromosome stats
Analyzer eligibility depends on measured data quality
Pipeline
Versioned (v7)
Existing genomes are flagged for re-analysis on every version bump
Citations
References
The peer-reviewed foundation behind the analyzers Haeckel runs.
- [1]The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature 526, 68–74 (2015).
- [2]Bergström A, et al. Insights into human genetic variation and population history from 929 diverse genomes. Science 367 (2020).
- [3]Ge T, Chen C-Y, Ni Y, et al. Polygenic prediction via Bayesian regression and continuous shrinkage priors (PRS-CS). Nat Commun 10, 1776 (2019).
- [4]Privé F, Arbel J, Vilhjálmsson BJ. LDpred2: better, faster, stronger. Bioinformatics 36, 5424–5431 (2020).
- [5]Mak TSH, et al. Polygenic scores via penalized regression on summary statistics (lassosum). Genet Epidemiol 41, 469–480 (2017).
- [6]Lloyd-Jones LR, et al. Improved polygenic prediction by Bayesian multiple regression on summary statistics (SBayesR). Nat Commun 10, 5086 (2019).
- [7]Walsh S, et al. The HIrisPlex system for simultaneous prediction of hair and eye colour. Forensic Sci Int Genet 7, 98–115 (2013).
- [8]Manichaikul A, et al. Robust relationship inference in genome-wide association studies (KING). Bioinformatics 26, 2867–2873 (2010).
- [9]Green RE, et al. A draft sequence of the Neandertal genome. Science 328, 710–722 (2010).
- [10]Vernot B, Akey JM. Resurrecting surviving Neandertal lineages from modern human genomes (S*). Science 343, 1017–1021 (2014).
- [11]Clinical Pharmacogenetics Implementation Consortium (CPIC) guidelines. cpicpgx.org.
- [12]Lambert SA, et al. The Polygenic Score Catalog as an open database for reproducibility. Nat Genet 53, 420–425 (2021).
- [13]Brixi G, et al. Genome modeling and design across all domains of life with Evo 2. Arc Institute (2025).