Reductions return NumPy objects eagerly but blosc2 arrays lazily
- 主要语言
- Python
- 星标
- 211
- 派生
- 58
- 平均合并
- 1 天 17 小时
- 30 天内合并 PR
- 6
描述
Reductions return NumPy objects through the eager path but blosc2 arrays through
the lazy one, so the same reduction has two different return types depending on
how it is spelled.
```python
import numpy as np, blosc2
na = np.arange(100, dtype="f8").reshape(10, 10)
a, b = blosc2.asarray(na), blosc2.asarray(na * 2)
a.sum() # numpy.float64
blosc2.sum(a) # numpy.float64
a.sum(axis=0) # numpy.ndarray
(a + b).sum(axis=0) # numpy.ndarray
a.mean(), a.std() # numpy.float64
blosc2.any(a > 5) # numpy.bool
blosc2.lazyexpr("sum(a + b)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape ()
blosc2.lazyexpr("sum(a + b, axis=0)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape (10,)
(a + b).compute() # blosc2.NDArray (no reduction: stays blosc2)
```
The last group is the point: `sum(a + b)` written as a string expression gives a
`blosc2.NDArray`, while the same reduction written as `(a + b).sum()` gives a
`numpy.float64`. A caller cannot tell from the operation what they will get back.
Beyond the inconsistency, returning NumPy makes reductions a hole in the lazy
model — the result of a reduction over a 100 GB array is materialised eagerly
and cannot be fed back into a blosc2 pipeline without a round trip through
`asarray`.
### Array API
The array API specification has reduction functions return arrays of the calling
namespace rather than language scalars, so `blosc2.sum(a)` yielding
`numpy.float64` is non-compliant regardless of which way the inconsistency above
is resolved. Related to #466, but kept separate: that issue is a checklist
organised by failing test file, whereas this cuts across every reduction and
needs a decision before those boxes can be ticked consistently.
### Open questions
1. **Which way to converge?** Making everything return blosc2 arrays is the
array-API answer and the consistent one. Making the lazy path return NumPy
would also be consistent, but gives up the ability to keep a reduction lazy.
2. **What about full (0-d) reductions?** A full reduction is one number, and
wrapping it in a compressed container with a schunk, chunks and blocks costs
more than it returns — every subsequent `float()` pays a decompression. An
`axis=`-reduction is a different matter. A split rule ("arrays for partial,
scalars for full") would be pragmatic but is exactly what the array API
forbids.
3. **Deprecation.** This is a breaking change: `if a.sum() > 5`, `float(a.mean())`,
passing a result to matplotlib or to a C API all change behaviour. It needs a
transition plan — a keyword, a namespace flag, or a release where both are
documented — rather than a flag day.
Raised as a follow-up in #457 (whose indexing bug is now fixed), and split out
because it is an API semantics decision rather than a defect.
贡献指南
调研方向
首先跟踪示例中展示的 reduction 入口:blosc2.sum、array .sum()、array .mean()、array .std()、blosc2.any() 和 lazyexpr(...).compute()。将它们的返回类型与 Array API 的 reduction 要求进行比较,然后解决 full-reduction、partial-reduction 和 deprecation 相关的问题,再定义什么样的一致行为和测试才算完成。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- backend-api-design, data
- Issue 类型
- 功能
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 冷清
- 描述清晰度
- 需要澄清
- 新手友好度
- 35/100