scverse / scverse/scanpy

sc.tl.rank_genes_groups: reference argument is ignored

Open
#1,485 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
2.6k
Forks
779
Avg merge
1d 4h
Merged PRs (30d)
27

Description

  • [ x] I have checked that this issue has not already been reported.
  • I have confirmed this bug exists on the latest version of scanpy.
  • (optional) I have confirmed this bug exists on the master branch of scanpy.

Hello,

I am having problems with the sc.tl.rank_genes_groups function. Specifically, I specify a reference level with reference = argument but it is ignored. The table that this function produces in the .uns object indicates the reference as rest (the default) when I have indicated otherwise.

print(set(noncycling_adult.obs.class_1))
#{'krt', 'dendritic', 'eccrine', 'T-cell', 'mel'}

sc.tl.rank_genes_groups(noncycling_adult, groupby = 'class_1', groups = ['eccrine', 'krt', 'T-cell', 'dendritic'], reference = 'mel', method = 'wilcoxon')

print(full_adata.uns['rank_genes_groups'])

"""{'params': {'groupby': 'class_1', 'reference': 'rest', 'method': 'wilcoxon', 'use_raw': True, 'corr_method': 'benjamini-hochberg'}, 'scores': rec.array([(8.494621 ,), (8.326364 ,), (8.24139  ,), (7.382108 ,),
           (7.340947 ,), (7.25889  ,), (7.2148457,), (7.0626616,),
           (6.991276 ,), (6.952865 ,)],
          dtype=[('T-cell', '<f4')]), 'names': rec.array([('IL32',), ('CD52',), ('CORO1A',), ('CD3D',), ('IL2RG',),
           ('PTPRCAP',), ('RAC2',), ('CD2',), ('LTB',), ('S100A4',)],
          dtype=[('T-cell', '<U50')]), 'logfoldchanges': rec.array([(10.175177 ,), (12.354224 ,), (11.05518  ,), (14.337216 ,),
           (11.3317585,), ( 9.758805 ,), ( 8.825092 ,), (14.170704 ,),
           (10.144425 ,), ( 5.6517367,)],
          dtype=[('T-cell', '<f4')]), 'pvals': rec.array([(1.98579427e-17,), (8.33632215e-17,), (1.70221006e-16,),
           (1.55802204e-13,), (2.12087430e-13,), (3.90279912e-13,),
           (5.39952731e-13,), (1.63343167e-12,), (2.72397796e-12,),
           (3.57940624e-12,)],
          dtype=[('T-cell', '<f8')]), 'pvals_adj': rec.array([(4.86400449e-13,), (1.02094937e-12,), (1.38979777e-12,),
           (9.54054799e-10,), (1.03897390e-09,), (1.59325269e-09,),
           (1.88937174e-09,), (5.00115941e-09,), (7.41345735e-09,),
           (8.76739764e-09,)],
          dtype=[('T-cell', '<f8')])}
"""

Thanks for your help

Versions

scanpy==1.4.4.post1 anndata==0.6.22.post1 umap==0.3.10 numpy==1.17.4 scipy==1.3.2 pandas==1.1.3 scikit-learn==0.22 statsmodels==0.12.0 python-igraph==0.7.1 louvain==0.6.1

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the example with sc.tl.rank_genes_groups and inspect the implementation and stored .uns['rank_genes_groups'] output. Confirm whether reference='mel' is retained and used instead of the default 'rest'; done means the parameters and ranking results reflect the requested reference.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
bioinformatics
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.