Reductions return NumPy objects eagerly but blosc2 arrays lazily
- 主要言語
- Python
- スター
- 211
- フォーク
- 58
- 平均マージ
- 1日 17時間
- マージ済み PR(30日)
- 6
説明
Reductions return NumPy objects through the eager path but blosc2 arrays through
the lazy one, so the same reduction has two different return types depending on
how it is spelled.
```python
import numpy as np, blosc2
na = np.arange(100, dtype="f8").reshape(10, 10)
a, b = blosc2.asarray(na), blosc2.asarray(na * 2)
a.sum() # numpy.float64
blosc2.sum(a) # numpy.float64
a.sum(axis=0) # numpy.ndarray
(a + b).sum(axis=0) # numpy.ndarray
a.mean(), a.std() # numpy.float64
blosc2.any(a > 5) # numpy.bool
blosc2.lazyexpr("sum(a + b)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape ()
blosc2.lazyexpr("sum(a + b, axis=0)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape (10,)
(a + b).compute() # blosc2.NDArray (no reduction: stays blosc2)
```
The last group is the point: `sum(a + b)` written as a string expression gives a
`blosc2.NDArray`, while the same reduction written as `(a + b).sum()` gives a
`numpy.float64`. A caller cannot tell from the operation what they will get back.
Beyond the inconsistency, returning NumPy makes reductions a hole in the lazy
model — the result of a reduction over a 100 GB array is materialised eagerly
and cannot be fed back into a blosc2 pipeline without a round trip through
`asarray`.
### Array API
The array API specification has reduction functions return arrays of the calling
namespace rather than language scalars, so `blosc2.sum(a)` yielding
`numpy.float64` is non-compliant regardless of which way the inconsistency above
is resolved. Related to #466, but kept separate: that issue is a checklist
organised by failing test file, whereas this cuts across every reduction and
needs a decision before those boxes can be ticked consistently.
### Open questions
1. **Which way to converge?** Making everything return blosc2 arrays is the
array-API answer and the consistent one. Making the lazy path return NumPy
would also be consistent, but gives up the ability to keep a reduction lazy.
2. **What about full (0-d) reductions?** A full reduction is one number, and
wrapping it in a compressed container with a schunk, chunks and blocks costs
more than it returns — every subsequent `float()` pays a decompression. An
`axis=`-reduction is a different matter. A split rule ("arrays for partial,
scalars for full") would be pragmatic but is exactly what the array API
forbids.
3. **Deprecation.** This is a breaking change: `if a.sum() > 5`, `float(a.mean())`,
passing a result to matplotlib or to a C API all change behaviour. It needs a
transition plan — a keyword, a namespace flag, or a release where both are
documented — rather than a flag day.
Raised as a follow-up in #457 (whose indexing bug is now fixed), and split out
because it is an API semantics decision rather than a defect.
コントリビューションガイド
調査の方向性
まず、例に示されている reduction のエントリポイントを追跡します: blosc2.sum、array .sum()、array .mean()、array .std()、blosc2.any()、lazyexpr(...).compute()。戻り値の型を Array API の reduction 要件と比較し、そのうえで、完全 reduction、部分 reduction、deprecation に関する問題を解決してから、一貫した動作とどのテストが完了の条件となるかを定義します。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- backend-api-design, data
- issue の種類
- 機能追加
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 静か
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 35/100