DReichLab / DReichLab/AdmixTools
Question about extremely high Z-scores (Z = 100) in qpDstat results
Nobody has claimed this yet.
- Dominant language
- C
- Stars
- 235
- Forks
- 76
- PR merge metrics
- No merged PRs in 30d
Description
Dear Author,
I am currently using qpDstat from the Admixtools package to calculate D-statistics among several rodent species. However, I have noticed that many of my results show Z-scores = 100, and most tests have |Z| > 3. I would like to ask whether this pattern is expected or indicates a problem in my setup.
Here are some details about my data and parameters:
Reference genome size: ~2.3 Gb
Number of SNPs used: 278518514,quite large
Parameter file: blgsize: 0.01
Outgroup: a relatively distant species
Many tests were run for all possible quartets (W, X, Y, Z), not only for topologies consistent with the species tree.
My questions are:
Is it normal to get Z = 100 for many tests, or does this indicate numerical saturation (e.g., SE(D) too small)?
Should I increase the block size (e.g., blgsize: 0.05 or larger) to avoid unrealistically small SE values?
Would it be more appropriate to limit the tests to quartets consistent with the species tree, instead of testing all possible combinations?
Could the high Z-scores result from using too distant an outgroup or from excessive divergence among species? In that case, would you recommend restricting the D-statistic tests within clades and choosing the nearest outgroup for each trio?
Any guidance on how to interpret these large Z-scores and how to adjust parameters or filtering strategies would be greatly appreciated.
Thank you very much for your time and for maintaining this excellent tool.
Best regards,
Na Wan
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the qpDstat entry point and its handling of the blgsize parameter, then examine how Z-scores and block-based standard errors are calculated. Done would require determining whether the reported Z = 100 values reflect expected input or a numerical problem and documenting evidence-based guidance on block size, quartet selection, and outgroup choice.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- c
- Domain
- bioinformatics
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100