apache / apache/echarts

[Feature] Native pivot / group-by-column dataset transform for long-format records

Open
#21,742 1 comment 0 reactions 0 assignees View on GitHub
en pending
Dominant language
TypeScript
Stars
67.3k
Forks
19.8k
Avg merge
11d 14h
Merged PRs (30d)
8

Description

### What problem does this feature solve?

ECharts `dataset` is strongest when the source is already a **pivot table** (one row per category, one column per series). A large share of real data — SQL result sets, CSV exports, analytics event logs — arrives in **long / tidy / record** form instead:

```js
dataset: {
dimensions: ['Country', 'Year', 'Population'],
source: [
['Brazil', 2011, 18203],
['Brazil', 2012, 19325],
['Indonesia', 2011, 23489],
['USA', 2011, 29034]
]
}
```

There is still no first-class way to turn that into N series (one per `Year`, or one per `Country`) without either:

- pre-pivoting outside ECharts, or
- one `filter` transform **per distinct group** (`O(groups × rows)`), which does not scale.

This was originally described in #16083 (closed stale; still requested in 2025–2026). @100pah agreed a **pivot transformer** belongs in the transform pipeline, possibly piped after aggregate.

Built-in `filter` / `sort` cannot express this. Third-party `echarts-simple-transform` only aggregates **one** `groupBy` dimension and does not reshape long → wide. Draft PR #16903 migrates aggregate/id into core but does not add pivot / “split into series by column”.

We hit this in production SQL → option pipelines: one query returns tidy rows, several cartesian/polar series need to share that table, and `encode` cannot invent series from a grouping column.

Related: #15306 (built-in aggregate; still open), #16903 (draft migration, stale).

This is not an alternative to an existing API — it is the missing reshape that makes `dataset` + `encode` usable for the format databases actually emit.

### What does the proposed API look like?

Keep it a **single-upstream** `dataset.transform`, consistent with filter/sort and with @100pah’s note on #16083 (standalone, or piped after aggregate).

**1. Split long data into multiple series (the #16083 bar case)**

```js
dataset: [
{ id: 'raw', source: /* Country, Year, Population */ },
{
fromDatasetId: 'raw',
transform: {
type: 'pivot', // or 'group'
config: {
row: 'Country',
column: 'Year',
value: 'Population',
fill: 0
}
}
}
]
// result dimensions: Country, 2011, 2012, ...
series: [
{ type: 'bar', encode: { x: 'Population', y: 'Country' } } // plus one series per year via encode, or
]
```

A companion option for “one series per distinct value of `column`” (so callers do not have to list years) would close the remaining gap:

```js
series: {
type: 'bar',
datasetId: 'raw',
encode: { x: 'Population', y: 'Country' },
// proposed — not currently valid
groupBy: 'Year'
}
```

`series.groupBy` is the better UX for bar/line/scatter; `dataset.transform: pivot` is the better fit for the existing transform plugin model and for heatmap/treemap that want a wide table. Both can share one implementation.

**2. Pipe after aggregate (#15306)**

Long data with duplicates first:

```js
dataset: [
{ source: rows },
{ transform: { type: 'aggregate', config: { groupBy: ['Country', 'Year'], output: [{ from: 'Population', method: 'sum' }] } } },
{ transform: { type: 'pivot', config: { row: 'Country', column: 'Year', value: 'Population' } } }
]
```

**3. What this should not try to do**

Cross-dataset joins (sankey nodes+edges, heatmap cells+axis catalogs) stay outside this transform — a pivot only sees one upstream `dataset`, which is the right scope for core. Multi-source composition belongs in application code.

**Prior art**

- Vega `pivot` / `fold`
- Production-validated row/col/value config: `{ from, rowKey, colKey, value, fill }`
- Community aggregate: https://github.com/100pah/echarts-simple-transform (SUM currently wrong on npm until https://github.com/100pah/echarts-simple-transform/pull/9)

Happy to iterate on names (`pivot` vs `spread` vs `group`) and on whether series-level `groupBy` is in scope for a first RFC. I am **not** opening a core PR until there is maintainer agreement on the shape — per the contributing wiki this is RFC-sized, not a drive-by patch.

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the existing dataset.transform filter/sort pipeline and the related aggregate work referenced in #15306 and #16903. First get maintainer agreement on the pivot versus series.groupBy API and its single-upstream scope; done means an accepted RFC and agreed implementation direction, since no files or tests are named.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
data-visualization
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.