Blosc / Blosc/python-blosc2

Reductions return NumPy objects eagerly but blosc2 arrays lazily

Đang mở
#689 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
211
Fork
58
Merge trung bình
1 ngày 17 giờ
Pull request đã merge (30 ngày)
6

Mô tả

Reductions return NumPy objects through the eager path but blosc2 arrays through
the lazy one, so the same reduction has two different return types depending on
how it is spelled.

```python
import numpy as np, blosc2

na = np.arange(100, dtype="f8").reshape(10, 10)
a, b = blosc2.asarray(na), blosc2.asarray(na * 2)

a.sum() # numpy.float64
blosc2.sum(a) # numpy.float64
a.sum(axis=0) # numpy.ndarray
(a + b).sum(axis=0) # numpy.ndarray
a.mean(), a.std() # numpy.float64
blosc2.any(a > 5) # numpy.bool

blosc2.lazyexpr("sum(a + b)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape ()
blosc2.lazyexpr("sum(a + b, axis=0)", {"a": a, "b": b}).compute() # blosc2.NDArray, shape (10,)

(a + b).compute() # blosc2.NDArray (no reduction: stays blosc2)
```

The last group is the point: `sum(a + b)` written as a string expression gives a
`blosc2.NDArray`, while the same reduction written as `(a + b).sum()` gives a
`numpy.float64`. A caller cannot tell from the operation what they will get back.

Beyond the inconsistency, returning NumPy makes reductions a hole in the lazy
model — the result of a reduction over a 100 GB array is materialised eagerly
and cannot be fed back into a blosc2 pipeline without a round trip through
`asarray`.

### Array API

The array API specification has reduction functions return arrays of the calling
namespace rather than language scalars, so `blosc2.sum(a)` yielding
`numpy.float64` is non-compliant regardless of which way the inconsistency above
is resolved. Related to #466, but kept separate: that issue is a checklist
organised by failing test file, whereas this cuts across every reduction and
needs a decision before those boxes can be ticked consistently.

### Open questions

1. **Which way to converge?** Making everything return blosc2 arrays is the
array-API answer and the consistent one. Making the lazy path return NumPy
would also be consistent, but gives up the ability to keep a reduction lazy.
2. **What about full (0-d) reductions?** A full reduction is one number, and
wrapping it in a compressed container with a schunk, chunks and blocks costs
more than it returns — every subsequent `float()` pays a decompression. An
`axis=`-reduction is a different matter. A split rule ("arrays for partial,
scalars for full") would be pragmatic but is exactly what the array API
forbids.
3. **Deprecation.** This is a breaking change: `if a.sum() > 5`, `float(a.mean())`,
passing a result to matplotlib or to a C API all change behaviour. It needs a
transition plan — a keyword, a namespace flag, or a release where both are
documented — rather than a flag day.

Raised as a follow-up in #457 (whose indexing bug is now fixed), and split out
because it is an API semantics decision rather than a defect.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Bắt đầu bằng cách lần theo các điểm truy cập reduction được minh họa trong các ví dụ: blosc2.sum, array .sum(), array .mean(), array .std(), blosc2.any() và lazyexpr(...).compute(). So sánh các kiểu trả về của chúng với các yêu cầu reduction của Array API, sau đó giải quyết các vấn đề về full-reduction, partial-reduction và deprecation trước khi xác định hành vi nhất quán và các bài kiểm thử nào sẽ được xem là hoàn tất.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python
Lĩnh vực
backend-api-design, data
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Ít trao đổi
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
35/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.