scverse / scverse/PyDESeq2

dispersion algorithm

Open
#252 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
761
Forks
90
Avg merge
25m
Merged PRs (30d)
10

Description

Hi,
It is reported "Dispersion parameters are first estimated independently for each gene by fitting a negative binomial generalized linear model (GLM)" in pydeseq2 bioinformatics paper. Since I am not a statistician, I cannot understand the complicated statistics principle under pydeseq2 and DEA. However, I wanna know which group of samples are used to get dispersion, for example, if the control group contains three samples and the treatment group contains another three samples. Which samples will be used to calculate the dispersion ?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names no files, tests, or entry points. Read the PyDESeq2 dispersion-estimation documentation and the referenced bioinformatics paper first, then determine whether the answer should be documented; done means clearly explaining which samples contribute to dispersion estimation for the stated two-group example.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.