pymc-devs / pymc-devs/pytensor
User facing split does not match numpy
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 644
- Forks
- 208
- Avg merge
- 2d 14h
- Merged PRs (30d)
- 16
Description
Description
Our implementation of Split expects the sizes of the subarrays (instead of the truncation points like numpy) and these must match to the full size of the input.
This is arguably a more sensible implementation as this is valid in numpy:
import numpy as np
np.split(np.zeros(10), [5, 2])
However users wanting to use split may be tripped by the fact it is different than the numpy one. We should offer a helper for the numpy one and keep ours internally. We can canonicalize to ours when we know all split sizes are non-decreasing.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating the Split implementation and its current size-based behavior. Compare it with the numpy split behavior shown in the issue, then define the helper and canonicalization boundaries; done means users can request numpy-style split points while the existing internal representation remains supported.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- numpy, python
- Domain
- data
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100