DReichLab / DReichLab/AdmixTools

High standard errors in qpDstat results

Open
#105 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C
Stars
235
Forks
76
PR merge metrics
No merged PRs in 30d

Description

Hi, I noticed that sometimes when I ran qpDstat(version: 662) with "printsd: YES", the output standard error (in column seven) can reach up to one, and the value of D/Z did not equal to the output standard_error. As the following results showed, when the output standard error is 1, the absolute of Z score still reach above 3.

result: Mbuti Afanasievo A B -0.0050 1.000000 -4.005 53480 54018 1099497
result: Mbuti AfontovaGora3 A B -0.0007 1.000000 -0.277 12366 12384 264327
result: Mbuti Aigyrzhal_BA A B -0.0061 1.000000 -3.680 39145 39627 811035
result: Mbuti Altaian.DG A B -0.0035 1.000000 -2.018 54744 55131 1090835
result: Mbuti Ami.DG A B -0.0058 1.000000 -3.473 55438 56082 1092691
result: Mbuti Anatolia_EBA A B -0.0025 1.000000 -1.588 46813 47045 989892
result: Mbuti Anatolia_N A B -0.0046 1.000000 -3.622 53485 53981 1109335
result: Mbuti Andronovo A B -0.0038 1.000000 -1.746 24760 24949 525918

So I am wondering, is this the real calculated standard error, or is this a sign of some calculating problem? Also, are these D results trust-worthy to infer population relationships? In which kind of scenario will this (high standard error) be more likely to happen?
Looking forward to your reply. If there is any thing that I misundertand, I will be really grateful to know. Thanks a lot!!

Han

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Begin with qpDstat version 662 using printsd: YES and the sample result rows in the issue; inspect how the seventh output column and D/Z are produced. Reproduce the high-standard-error cases and determine whether the calculation is faulty, which scenarios trigger it, and whether a code or test change is required.

Written by the indexing model from the issue text.

Assessment

Tech stack
c
Domain
bioinformatics
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.