DReichLab / DReichLab/AdmixTools

Question about extremely high Z-scores (Z = 100) in qpDstat results

Open
#122 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C
Stars
235
Forks
76
PR merge metrics
No merged PRs in 30d

Description

Dear Author,

I am currently using qpDstat from the Admixtools package to calculate D-statistics among several rodent species. However, I have noticed that many of my results show Z-scores = 100, and most tests have |Z| > 3. I would like to ask whether this pattern is expected or indicates a problem in my setup.

Image

Here are some details about my data and parameters:

Reference genome size: ~2.3 Gb

Number of SNPs used: 278518514,quite large

Parameter file: blgsize: 0.01

Outgroup: a relatively distant species

Many tests were run for all possible quartets (W, X, Y, Z), not only for topologies consistent with the species tree.

My questions are:

Is it normal to get Z = 100 for many tests, or does this indicate numerical saturation (e.g., SE(D) too small)?

Should I increase the block size (e.g., blgsize: 0.05 or larger) to avoid unrealistically small SE values?

Would it be more appropriate to limit the tests to quartets consistent with the species tree, instead of testing all possible combinations?

Could the high Z-scores result from using too distant an outgroup or from excessive divergence among species? In that case, would you recommend restricting the D-statistic tests within clades and choosing the nearest outgroup for each trio?

Any guidance on how to interpret these large Z-scores and how to adjust parameters or filtering strategies would be greatly appreciated.

Thank you very much for your time and for maintaining this excellent tool.

Best regards,
Na Wan

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the qpDstat entry point and its handling of the blgsize parameter, then examine how Z-scores and block-based standard errors are calculated. Done would require determining whether the reported Z = 100 values reflect expected input or a numerical problem and documenting evidence-based guidance on block size, quartet selection, and outgroup choice.

Written by the indexing model from the issue text.

Assessment

Tech stack
c
Domain
bioinformatics
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.