Methods · Validation

Every inference has a method.

The public ledger for Haeckel analyzers: named methods, reference panels, confidence gates, and known limits.

1000 Genomes + HGDP validation · Pipeline v7

The ledger

What runs, and what backs it

Research belongs here as executable method: one analyzer, one scope, one uncertainty model.

MLE-EM + spatial thinning + Bootstrap Wald + AMR deconvolution[1][2]
Ancestry inference
26,000 independent AIMs across 62 sub-populations from 1000 Genomes and HGDP. Signal below threshold returns Unassigned.
Basic PRS · PRS-CS · C+T · LDpred2 · Lassosum · SBayesR · ensemble R-hat[3][4][5][6][12]
Polygenic risk scoring
Six production methods with palindromic-SNP handling, LD-aware models, ensemble weighting, and convergence diagnostics.
Recursive phylogenetic traversal + back-mutation handling + Bayesian confidence
Y-DNA + mtDNA haplogroups (165 nodes)
Lineage calls are annotated with historical figures, migration context, and confidence against random expectation.
S* statistic · ABBA-BABA D-stats · block jackknife · 4-state Viterbi HMM[9][10]
Archaic introgression
Deep coalescence, Neanderthal and Denisovan tract detection, coalescent dating, and 53 archaic markers.
HIrisPlex-S · GIANT height · per-SNP contribution tracking[7]
Phenotype prediction
Eye, hair, skin, and height predictions with coverage gates, MC1R missingness protection, and per-SNP explainability.
CPIC star-allele calling · activity scores · FDA Black Box alerts[11]
Pharmacogenomics & safety
12 CPIC Level-A genes including CYP2D6, CYP2C19, DPYD, SLCO1B1, and HLA-B; 10 critical FDA Black Box interactions.
Whole-genome inheritance calculator · Mendelian simulation · PRS comparison[3][4][5][6][12]
Offspring and embryo modeling
Partner compatibility, offspring trait simulation, and sibling/embryo comparison across complex polygenic phenotypes.
Call rate · heterozygosity · Ti/Tv · per-chromosome statistics[1][2]
Quality control
Every upload is screened before analysis; low-confidence data is surfaced as missing or blocked, not guessed.
KING-robust kinship · ROH detection[8]
Relationship inference
Kinship, consanguinity, and runs of homozygosity for family structure and inherited-risk context.
pgvector 1536-d embeddings
Networks and candidate discovery
Vectorized genomic, phenotype, and profile embeddings for kinship-aware search and matching.
Evo 2 DNA foundation model (in integration)[13]
DNA foundation-model interpretation
Evo 2 integration for variants of uncertain significance and first-principles sequence reasoning.
Rigor

Validation & confidence

A method is only as good as the honesty of its uncertainty. Haeckel gates on confidence and refuses to guess past its coverage.

Ancestry
Significance-gated
Bootstrap Wald; signal below the threshold is returned as Unassigned
AMR deconvolution
Admixed reference orthogonalized to its Native American vertex
Reduces spurious cross-population signal in unadmixed individuals
Phenotype
Coverage-gated (HIrisPlex-S)
A 67% coverage threshold protects against systematic MC1R missingness bias
Haplogroups
Tree integrity gate at module load
Back-mutations and unauthored child nodes are handled before calls ship
Offspring modeling
Probability distributions, not certainties
Sibling comparisons preserve uncertainty instead of collapsing to a rank alone
Upload QC
Call rate · heterozygosity · Ti/Tv · chromosome stats
Analyzer eligibility depends on measured data quality
Pipeline
Versioned (v7)
Existing genomes are flagged for re-analysis on every version bump
Citations

References

The peer-reviewed foundation behind the analyzers Haeckel runs.

  1. [1]The 1000 Genomes Project Consortium. A global reference for human genetic variation. Nature 526, 68–74 (2015).
  2. [2]Bergström A, et al. Insights into human genetic variation and population history from 929 diverse genomes. Science 367 (2020).
  3. [3]Ge T, Chen C-Y, Ni Y, et al. Polygenic prediction via Bayesian regression and continuous shrinkage priors (PRS-CS). Nat Commun 10, 1776 (2019).
  4. [4]Privé F, Arbel J, Vilhjálmsson BJ. LDpred2: better, faster, stronger. Bioinformatics 36, 5424–5431 (2020).
  5. [5]Mak TSH, et al. Polygenic scores via penalized regression on summary statistics (lassosum). Genet Epidemiol 41, 469–480 (2017).
  6. [6]Lloyd-Jones LR, et al. Improved polygenic prediction by Bayesian multiple regression on summary statistics (SBayesR). Nat Commun 10, 5086 (2019).
  7. [7]Walsh S, et al. The HIrisPlex system for simultaneous prediction of hair and eye colour. Forensic Sci Int Genet 7, 98–115 (2013).
  8. [8]Manichaikul A, et al. Robust relationship inference in genome-wide association studies (KING). Bioinformatics 26, 2867–2873 (2010).
  9. [9]Green RE, et al. A draft sequence of the Neandertal genome. Science 328, 710–722 (2010).
  10. [10]Vernot B, Akey JM. Resurrecting surviving Neandertal lineages from modern human genomes (S*). Science 343, 1017–1021 (2014).
  11. [11]Clinical Pharmacogenetics Implementation Consortium (CPIC) guidelines. cpicpgx.org.
  12. [12]Lambert SA, et al. The Polygenic Score Catalog as an open database for reproducibility. Nat Genet 53, 420–425 (2021).
  13. [13]Brixi G, et al. Genome modeling and design across all domains of life with Evo 2. Arc Institute (2025).