[RFC]: a simple flat array format for ndarrays
还没有人认领这个 Issue。
- 主要语言
- JavaScript
- 星标
- 6k
- 派生
- 1.3k
- 平均合并
- 1 天 3 小时
- 30 天内合并 PR
- 611
描述
Description
This RFC proposes introducing a simple flat array format for ndarrays and is inspired by work involving the integration of stdlib in Google Sheets. The motivation for this RFC is to provide a human-readable, non-binary format for serializing and deserializing ndarrays, which is JSON compatible.
At a high level, the format is comprised of a version, header, and list of data buffer elements.
<version> | <header> | <data>
The version component would be comprised of two elements:
[ 'version', '<semver>', ... ]
The first element is the string literal 'version' and is followed by a version string in semver format. It is not anticipated that the patch field of the version string will be used. Only major (breaking changes) and minor (new features/header fields) version fields should update over time.
The header component would be comprised as follows:
'ndarray' | shape | strides | offset | order | dtype | length | capacity | 'data'
and as part of the serialized array
[ ..., 'ndarray', 'shape', ...shape, 'strides', ...strides, 'offset', offset, 'order', order, 'dtype', dtype, 'length', length, 'capacity', capacity, 'data', ... ]
where
'ndarray'is the string literal'ndarray'.'shape'is the string literal'shape'....shapeis 0 or more dimension sizes. For a zero-dimensional array, no dimension sizes should be present.'strides'is the string literal'strides'....stridesis 1 or more dimension strides. For a zero-dimensional array, one stride should be present, which should be equal to0.'offset'is the string literal'offset'.offsetis a nonnegative integer indicating the index offset in the data buffer marking the first indexed element. The offset of the first indexed element in the serialized format would beversion_length + header_length + offset, where one must take into account the version and header lengths.'order'is the string literal'order'.orderis either'row-major'or'column-major'.'dtype'is the string literal'dtype'.dtypeis the ndarray data type string (e.g.,'float64','complex128','int32', etc).'length'is the string literal'length'.lengthis a nonnegative integer indicating how many elements are indexed by the ndarray. For a zero-dimensional array, this should equal1. For non-zero-dimensional arrays, this should be equal to the product of dimension sizes, as listed inshape.'capacity'is the string literal'capacity'.capacityis a nonnegative integer indicating how many elements are in the data buffer. This value should be compatible with the specified ndarray meta data (i.e., shape, strides, offset). For zero-dimensional arrays, this should be greater than or equal to1.'data'is the string literal'data'and should be followed by data buffer elements.
The 'ndarray' string literal is required to be the first header element. The 'data' string literal is required to be the last header element. For the other header elements, each string literal and associated value pair can be arranged in any order. E.g.,
[ ..., 'ndarray', 'capacity', capacity, 'length', length, 'dtype', dtype, 'order', order, 'offset', offset, 'strides', ...strides, 'shape', ...shape, 'data', ... ]
would be valid. Parsers should not assume any particular string literal and value pair order and should instead identify a sub-header element by the string literal indicating its beginning.
The data component is the linear data buffer atop which the serialized ndarray is a view. This data buffer is allowed to contain elements which are outside the view bounds and are not indexed by the view.
Example
The following is an example of a 2x2 ndarray serialized to the proposed linear exchange format:
[
'version',
'1.0.0',
'ndarray',
'shape',
2,
2,
'strides',
2,
1,
'offset',
0,
'order',
'row-major',
'dtype',
'float64',
'length',
4,
'capacity',
4,
'data',
1,
2,
3,
4
]
Note that this particular linear format is easily extendable to CSV/DSV serialization, where each column could represent a different ndarray.
Proposal
As part of this RFC, the following packages are proposed
@stdlib/ndarray/[base/]to-linear-exchange-format: serializes an ndarray to the proposed format.@stdlib/ndarray/[base/]from-linear-exchange-format: converts a serialized ndarray to an ndarray instance.
where [base/] indicates both base and non-base package versions.
The format name and associated package names are not set in stone. Any naming suggestions are welcome.
Prior Art
ndarrayobjects can already be serialized to JSON, using thendarray#toJSONmethod; however, the format is not a linear data structure (nor should it necessarily be) and does not serialize the data buffer outside of the array view. This prevents creating subsequent views of different sizes atop the same data buffer.- NumPy has an
*.npyformat; however, this does not include some of the meta data proposed in this RFC and is not human-readable. - NumPy also has an API,
savetxtfor saving an ndarray to text; however, this is primarily oriented to formatting, similar to@stdlib/string/format.
Related Issues
None.
Questions
No.
Other
No.
Checklist
- I have read and understood the Code of Conduct.
- Searched for existing issues and pull requests.
- The issue name begins with
RFC:.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
首先审查 RFC 中提出的线性格式,以及 Prior Art 下引用的现有 ndarray#toJSON 方法。将 to-linear-exchange-format 和 from-linear-exchange-format 这两个 package 入口点与当前的 ndarray API 进行比较。只有在开始编码之前就格式、命名和实现范围达成一致,才算完成。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- javascript
- 领域
- api, data
- Issue 类型
- 功能
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 25/100