kestra-io / kestra-io/plugin-transform

New sub-module: plugin-transform-arrow (jq queries over Arrow/Parquet/CSV/NDJSON)

Open
#70 0 comments 0 reactions 1 assignee Claimed by @fdelbrayelle View on GitHub
area/plugin good first issue
Dominant language
Java
Stars
1
Forks
8
Avg merge
1d 11h
Merged PRs (30d)
7

Description

## Summary

Add a new `plugin-transform-arrow` sub-module implementing [`aq`](https://github.com/Anaethelion/aq) — a tool that applies **jq-style filter expressions to columnar data files** (Parquet, Arrow IPC, CSV, NDJSON).

---

## What `aq` does

`aq` is \"jq for Apache Arrow\". Its pipeline:

1. **Read** — auto-detect and read Parquet, Arrow IPC, CSV, or NDJSON into Arrow `RecordBatch` objects
2. **Serialize** — convert each row to a JSON object (via NDJSON intermediate)
3. **Filter/transform** — apply a jq expression row-by-row (or in slurp mode over all rows)
4. **Output** — write results as NDJSON, JSON, CSV, TSV, or Arrow IPC

---

## Why `plugin-transform`, not `plugin-serdes`

`plugin-serdes` is purely a format converter (file ↔ Ion records) with no query or filter logic. `aq`'s core value is the **jq query engine applied to columnar data**, which belongs in `plugin-transform` alongside the existing `plugin-transform-json` (JSONata) and `plugin-transform-records` (SQL-like filter/map/aggregate).

---

## Why `plugin-transform-arrow`, not `plugin-transform-jq`

The name should reflect the **data format**, not the query language — consistent with how `plugin-transform-json` is named after JSON, not after JSONata. The differentiator here is Arrow/Parquet input support; jq is just the query mechanism. It also leaves room to add non-jq tasks later (schema inspection, format conversion within the Arrow ecosystem, etc.).

---

## Proposed tasks

| Task | Description |
|---|---|
| `Query` | Apply a jq expression to an Arrow/Parquet/CSV/NDJSON file, output as Ion records or file |

Mirrors the `Transform` / `TransformItems` pattern from `plugin-transform-json`.

---

## Java implementation

The Rust-to-Java mapping is straightforward:

| `aq` dependency (Rust) | Java equivalent |
|---|---|
| `parquet` crate | `org.apache.parquet:parquet-arrow` |
| `arrow` (IPC, CSV, JSON reader/writer) | `org.apache.arrow:arrow-vector`, `arrow-dataset`, `arrow-ipc` |
| Row → NDJSON serialization | Arrow Java JSON writer (`ArrowToJson`) |
| `jaq-core` (pure-Rust jq engine) | [`jackson-jq`](https://github.com/eiiches/jackson-jq) — best pure-Java jq implementation |
| `serde_json` | Jackson Databind |

**Note on `jackson-jq`:** it doesn't implement the full jq spec (missing `$ENV`, `input`/`inputs`, some path builtins). This is the same tradeoff `aq` itself makes (it uses `jaq`, not the reference `jq`), so it's acceptable for the target use cases.

---

## References

- [`aq` repository](https://github.com/Anaethelion/aq)
- [`plugin-transform-json`](https://github.com/kestra-io/plugin-transform/tree/main/plugin-transform-json) — existing sibling sub-module using JSONata
- [`jackson-jq`](https://github.com/eiiches/jackson-jq) — Java jq engine
- [Apache Arrow Java](https://arrow.apache.org/docs/java/)
- [Apache Parquet MR](https://github.com/apache/parquet-mr)

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.