scverse / scverse/scanpy

scanpy.api.tl.pca compensate for few genes < 50 but not few cells.

Open
#432 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
2.6k
Forks
779
Avg merge
1d 4h
Merged PRs (30d)
27

Description

Minor bug I assume.

Using fewer than 50 cells raises the following error when trying to run sc.tl.pca. The code handles this when the n_vars < 50 but not when n_obs is.

ValueErrorTraceback (most recent call last)
<ipython-input-823-4c11b9b62e6d> in <module>
----> 1 sc.tl.pca(bla)

~/.virtualenvs/default/lib/python3.6/site-packages/scanpy/preprocessing/simple.py in pca(data, n_comps, zero_center, svd_solver, random_state, return_info, use_highly_variable, dtype, copy, chunked, chunk_size)
    504             pca_ = TruncatedSVD(n_components=n_comps, random_state=random_state)
    505             X = adata_comp.X
--> 506         X_pca = pca_.fit_transform(X)
    507 
    508     if X_pca.dtype.descr != np.dtype(dtype).descr: X_pca = X_pca.astype(dtype)

~/.virtualenvs/default/lib/python3.6/site-packages/sklearn/decomposition/pca.py in fit_transform(self, X, y)
    357 
    358         """
--> 359         U, S, V = self._fit(X)
    360         U = U[:, :self.n_components_]
    361 

~/.virtualenvs/default/lib/python3.6/site-packages/sklearn/decomposition/pca.py in _fit(self, X)
    404         # Call different fits for either full or truncated SVD
    405         if self._fit_svd_solver == 'full':
--> 406             return self._fit_full(X, n_components)
    407         elif self._fit_svd_solver in ['arpack', 'randomized']:
    408             return self._fit_truncated(X, n_components, self._fit_svd_solver)

~/.virtualenvs/default/lib/python3.6/site-packages/sklearn/decomposition/pca.py in _fit_full(self, X, n_components)
    423                              "min(n_samples, n_features)=%r with "
    424                              "svd_solver='full'"
--> 425                              % (n_components, min(n_samples, n_features)))
    426         elif n_components >= 1:
    427             if not isinstance(n_components, (numbers.Integral, np.integer)):

ValueError: n_components=50 must be between 0 and min(n_samples, n_features)=38 with svd_solver='full'

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in the PCA implementation referenced by the traceback, preprocessing/simple.py, and reproduce sc.tl.pca with fewer than 50 cells. Compare the existing handling for n_vars below 50 with the n_obs case. Done means PCA no longer raises the shown n_components error for small cell counts and regression coverage verifies the behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, scikit-learn
Domain
bioinformatics, machine-learning
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.