anitnilay20 / anitnilay20/thoth
[Future] Arrow IPC: typed columnar format for the plugin dataset bus
- 主要言語
- Rust
- スター
- 70
- フォーク
- 5
- 平均マージ
- 10時間 33分
- マージ済み PR(30日)
- 2
説明
## Idea
Replace the current `list>` wire format in `dataset-bus.publish` with Arrow IPC — a flat typed columnar buffer that crosses the WASM boundary in a single copy instead of thousands of individual string copies.
## Why this matters
The plugin→host databus is the one place in the stack where data crosses a WASM boundary. Every cell in the current format is an individual string allocation and copy through the component model ABI. For plugins that push large datasets (100k+ rows), this becomes measurable overhead.
Arrow IPC replaces the WIT signature from:
```wit
publish: func(name: string, columns: list, rows: list>) -> string;
```
to something like:
```wit
publish-arrow: func(name: string, ipc-bytes: list) -> string;
```
The plugin serialises its data as an Arrow IPC stream; the host deserialises into a `RecordBatch`. DuckDB (already in the host) can scan Arrow record batches natively via `query_arrow()` — no second conversion needed.
## Not urgent
The current string format is fine for typical plugin datasets. The WIT file already notes this as a future direction. Revisit when profiling shows the databus crossing is a real bottleneck — likely only relevant for plugins that stream large, high-frequency datasets (monitoring, Kafka consumers, etc.).
## Prerequisites
- DuckDB native integration (#148) should be shipped first — Arrow's value compounds when DuckDB can consume the batch directly
- Requires Arrow encoding support in the plugin SDK (adds a dependency for plugin authors)
- WIT is a breaking change — needs a minor version bump and migration path for existing plugins
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
評価
この issue はまだ評価されていません。