graphprotocol / graphprotocol/graph-node

current: include emits an all-null bucket for dimensionless aggregations, nulling the whole response

Offen Anfängerfreundlich
#6,719 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
Rust
Sterne
3.2k
Forks
1.1k
Ø Merge
4 T. 1 Std.
Gemergte PRs (30 T.)
1

Beschreibung

When current: include is used on an aggregation that has no dimensions, and no unrolled source rows match the query's filters, graph-node emits a current bucket whose id, timestamp and @aggregate fields are all null. Those fields are non-null in the generated schema, so the query fails with Null value resolved for non-null field, and because a non-null violation propagates to the parent, the entire data becomes null — one empty bucket discards an otherwise valid response.

Introduced by #6293.

Reproduction

Any dimensionless aggregation. A far-future timestamp_gte deterministically guarantees zero matching rows:

{ myCandles(interval: "hour", current: include, where: { timestamp_gte: "9999999999000000" }) { id } }
{ "data": null, "errors": [{ "message": "Null value resolved for non-null field `id`" }] }

Selecting sum or timestamp instead fails on those fields. But the same query asking only for count succeeds, which shows the row really is emitted:

{ myCandles(interval: "hour", current: include, where: { timestamp_gte: "9999999999000000" }) { count } }
{ "data": { "myCandles": [{ "count": "0" }] } }
Root cause

select_current_bucket in store/postgres/src/relational/rollup.rs emits group by only when the aggregation has dimensions:

if !self.dimensions.is_empty() {
    write!(w, " group by ")?;
    write_dims(self.dimensions, w, false)?;
}

Without dimensions the generated query is a bare aggregate — see the repo's own fixture COUNT_ONLY_CURRENT_SQL, which ends at /*FILTERS*/) c with no GROUP BY. A bare aggregate always returns exactly one row, and over an empty input max()/sum() return NULL while count(*) returns 0. That is precisely the observed split. Aggregations with dimensions (STATS_HOUR_CURRENT_SQL) end in group by "token", and a GROUP BY over empty input yields zero rows, so they are unaffected.

Query filters are spliced in at the /*FILTERS*/ marker, which sits inside the inner subquery over the source timeseries table, so they run before aggregation and cannot suppress the synthesized row. I confirmed id_gte, id_gt and timestamp_gt all still fail.

This also contradicts docs/aggregations.md, which states the current bucket "covers the time period from the end of the last completed bucket up to the most recent data point" — with no data point there is no such period, and no bucket to return.

Expected

The current bucket should be omitted when no unrolled rows match, leaving only the completed buckets (or []).

Suggested fix

Emit having count(*) > 0 when self.dimensions.is_empty(), matching the zero-row behaviour GROUP BY already gives the dimensioned case.

Impact

This does not need a contrived filter. Any dimensionless aggregation queried with current: include over a wide range fails whenever the source timeseries has had no writes since the last rollup — e.g. an oracle that posts a few times an hour breaks every consumer for the first minutes of each hour. Clients see the whole response nulled, so a chart renders nothing rather than degrading to complete buckets.

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne in store/postgres/src/relational/rollup.rs bei select_current_bucket und vergleiche anschließend die Fixtures COUNT_ONLY_CURRENT_SQL und STATS_HOUR_CURRENT_SQL. Führe die Reproduktion der dimensionslosen aktuellen Aggregation mit einem weit in der Zukunft liegenden timestamp_gte aus; abgeschlossen ist die Aufgabe, wenn kein aktueller Bucket mit null-Wert zurückgegeben wird, sofern keine passenden nicht entrollten Zeilen vorhanden sind, sodass abgeschlossene Buckets oder eine leere Liste übrig bleiben.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
graphql, postgresql, rust
Bereich
api, databases
Issue-Typ
Bug
Schwierigkeit
2/5
Geschätzter Aufwand
1-3 Stunden
Aktivitätsstatus
Aktiv
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
78/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.