numpy / numpy/numpy

determining the datatype of recursive list is slow

Open
#9,311 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

component: numpy._core
Dominant language
Python
Stars
32.8k
Forks
12.8k
Avg merge
1d 7h
Merged PRs (30d)
197

Description

See https://github.com/pandas-dev/pandas/issues/16778

import numpy as np
c = []
c.append(c)
c.append(c)
np.array(c)

The code above hangs calling recursively the C level function PyArray_DTypeFromObjectHelper. Interestingly if I only append c once I get something rather strangely looking:

In [13]: c = []

In [14]: c.append(c)

In [15]: np.aray(c)
In [16]: np.array(c)
Out[16]: array([[[[[[[[[[[[[[[[[[[[[[[[[[[[[[[[list([[...]])]]]]]]]]]]]]]]]]]]]]]]]]]]]]]]]], dtype=object)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the recursive-list examples first, then inspect the C-level entry point PyArray_DTypeFromObjectHelper and the discussion in issue #16778. The issue names no source file or test and does not define the intended result; clarify the expected handling of self-referential lists before identifying a regression test and considering a fix.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, python
Domain
data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.