Blosc / Blosc/python-blosc2

Reductions return NumPy objects eagerly but blosc2 arrays lazily

未关闭
#689 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
211
派生
58
平均合并
1 天 17 小时
30 天内合并 PR
6

描述

Reductions return NumPy objects through the eager path but blosc2 arrays through
the lazy one, so the same reduction has two different return types depending on
how it is spelled.

```python
import numpy as np, blosc2

na = np.arange(100, dtype="f8").reshape(10, 10)
a, b = blosc2.asarray(na), blosc2.asarray(na * 2)

a.sum() # numpy.float64
blosc2.sum(a) # numpy.float64
a.sum(axis=0) # numpy.ndarray
(a + b).sum(axis=0) # numpy.ndarray
a.mean(), a.std() # numpy.float64
blosc2.any(a > 5) # numpy.bool

blosc2.lazyexpr("sum(a + b)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape ()
blosc2.lazyexpr("sum(a + b, axis=0)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape (10,)

(a + b).compute() # blosc2.NDArray (no reduction: stays blosc2)
```

The last group is the point: `sum(a + b)` written as a string expression gives a
`blosc2.NDArray`, while the same reduction written as `(a + b).sum()` gives a
`numpy.float64`. A caller cannot tell from the operation what they will get back.

Beyond the inconsistency, returning NumPy makes reductions a hole in the lazy
model — the result of a reduction over a 100 GB array is materialised eagerly
and cannot be fed back into a blosc2 pipeline without a round trip through
`asarray`.

### Array API

The array API specification has reduction functions return arrays of the calling
namespace rather than language scalars, so `blosc2.sum(a)` yielding
`numpy.float64` is non-compliant regardless of which way the inconsistency above
is resolved. Related to #466, but kept separate: that issue is a checklist
organised by failing test file, whereas this cuts across every reduction and
needs a decision before those boxes can be ticked consistently.

### Open questions

1. **Which way to converge?** Making everything return blosc2 arrays is the
array-API answer and the consistent one. Making the lazy path return NumPy
would also be consistent, but gives up the ability to keep a reduction lazy.
2. **What about full (0-d) reductions?** A full reduction is one number, and
wrapping it in a compressed container with a schunk, chunks and blocks costs
more than it returns — every subsequent `float()` pays a decompression. An
`axis=`-reduction is a different matter. A split rule ("arrays for partial,
scalars for full") would be pragmatic but is exactly what the array API
forbids.
3. **Deprecation.** This is a breaking change: `if a.sum() > 5`, `float(a.mean())`,
passing a result to matplotlib or to a C API all change behaviour. It needs a
transition plan — a keyword, a namespace flag, or a release where both are
documented — rather than a flag day.

Raised as a follow-up in #457 (whose indexing bug is now fixed), and split out
because it is an API semantics decision rather than a defect.

贡献指南

打开贡献指南

调研方向

首先跟踪示例中展示的 reduction 入口:blosc2.sum、array .sum()、array .mean()、array .std()、blosc2.any() 和 lazyexpr(...).compute()。将它们的返回类型与 Array API 的 reduction 要求进行比较,然后解决 full-reduction、partial-reduction 和 deprecation 相关的问题,再定义什么样的一致行为和测试才算完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
backend-api-design, data
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
需要澄清
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。