Indexed metadata

A proof-of-principle read-count-based probabilistic genotyping framework for mixture interpretation using a 6,267-locus hybridization-capture SNP-NGS panel

Tingyun Hou, Yuting Wang, Tiantian Shan, Haoyu Wang, Chun Yang, Yufang Wang, Qiang Zhu, Ji Zhang

Source record

Source: Crossref

Published: Jan 1, 2026

DOI: 10.2139/ssrn.7537006

Open original source ↗

Source abstract

Next-generation sequencing (NGS) enables genotyping of large-scale single nucleotide polymorphisms (SNPs), but interpretation of SNP-NGS mixtures remains challenging because most SNPs are biallelic, locus amplification efficiency affects allele read counts, sequencing noise occurs, and true alleles may drop out. In this proof-of-principle study, we developed a read-count-based fully continuous probabilistic genotyping (PG) framework for PCR-based hybridization-capture SNP-NGS mixtures. The framework incorporates allele read counts, sequencing noise, allele dropout, locus amplification efficiency (LAE), and a co-ancestry coefficient for population substructure. Noise and LAE parameters were calibrated using 35 single-source samples. We then examined single-source profiles and known-truth in silico mixtures, and assessed laboratory mixtures (15 two-person and 9 three-person). Performance was assessed using true-contributor likelihood ratios (LRs), non-contributor LRs, and genotype deconvolution. No-reference deconvolution of the laboratory mixtures was also compared with EuroForMix MPS. LAE correction improved agreement between allele read counts and relative DNA template from R² = 0.502 ± 0.103 to 0.800 ± 0.102. Most true-contributor tests yielded LRs greater than 1. LRs below 1 were observed mainly for low-template contributors in highly unbalanced mixtures. Non-contributor testing using 203 individuals from the 1000 Genomes CHS and CEU populations generated 4,872 computational H2-true tests, and no likelihood ratio greater than 1 was observed. These tests are not a casework false-positive rate. Genotype deconvolution was most reliable for highest-template contributors and for balanced to moderately unbalanced mixtures. These findings support read-count-based modelling as a calibration-dependent approach to PCR-based hybridization-capture SNP-NGS mixture interpretation and identify conditions under which the model is currently informative or limited.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.