ClickHouse / ClickHouse/clickhouse-go

Scanning large arrays is very expensive

Open
#1,747 0 comments 0 reactions 0 assignees View on GitHub
performance Q1-FY-2026
Dominant language
Go
Stars
3.3k
Forks
680
Avg merge
2d 3h
Merged PRs (30d)
14

Description

Here's an example query that takes ~450ms with curl to execute and drain the results:

```
$ time curl -s -G http://localhost:8123/ --data-urlencode "query=SELECT stackMap.stack as stack, sum(stackMap.value) as value FROM profiles_v3 ARRAY JOIN stackMap WHERE profileType = 'process_cpu:cpu:nanoseconds:cpu:nanoseconds' AND timestamp >= now() - 3600 * 3 AND serviceName = 'grafana.alloy.ebpf' AND label_app_name = '/clickhouse' GROUP BY stackMap.stack" | wc -c
124736917

real 0m0.456s
user 0m0.041s
sys 0m0.105s
```

Note almost 125MiB of data here, it's not a tiny result.

We can do the same query with Go:

```go
package main

import (
"context"
"log"
"time"

"github.com/ClickHouse/clickhouse-go/v2"
)

func main() {
db, err := clickhouse.Open(&clickhouse.Options{
Protocol: clickhouse.HTTP,
Addr: []string{"127.0.0.1:8123"},
})
if err != nil {
log.Fatal(err)
}

started := time.Now()

rows, err := db.Query(context.Background(), "SELECT stackMap.stack as stack, sum(stackMap.value) as value FROM profiles_v3 ARRAY JOIN stackMap WHERE profileType = 'process_cpu:cpu:nanoseconds:cpu:nanoseconds' AND timestamp >= now() - 3600 * 3 AND serviceName = 'grafana.alloy.ebpf' AND label_app_name = '/clickhouse' GROUP BY stackMap.stack")
if err != nil {
log.Fatal(err)
}

log.Printf("query returned in %.2fs", time.Since(started).Seconds())

started = time.Now()

// for rows.Next() {}

stacks := []stack{}
var frames []string
var value uint64

for rows.Next() {
err := rows.Scan(&frames, &value)
if err != nil {
log.Fatal(err)
}

stacks = append(stacks, stack{frames: frames, value: value})
}

log.Printf("row scanning finished in %.2fs", time.Since(started).Seconds())
}

type stack struct {
frames []string
value uint64
}
```

If we just iterate rows without any scanning (uncomment `for rows.Next() {}` and comment the actual loop):

```
$ go build -o /tmp/whoa ./cmd/huh && time GOMAXPROCS=1 /tmp/whoa
2026/01/04 06:53:50 query returned in 0.25s
2026/01/04 06:53:50 row scanning finished in 0.22s

real 0m0.469s
user 0m0.160s
sys 0m0.054s
```

That's pretty close to curl. We're also actually parsing columns, so that's not too bad:

Image

If we add scanning (using the unchanged code above), wall time to execute pretty much doubles, while on-CPU user time goes up 3.625x:

```
$ go build -o /tmp/whoa ./cmd/huh && time GOMAXPROCS=1 /tmp/whoa
2026/01/04 06:54:37 query returned in 0.24s
2026/01/04 06:54:38 row scanning finished in 0.72s

real 0m0.968s
user 0m0.578s
sys 0m0.149s
```

Here's the flamegraph:

Image

Zoomed into scanning specifically:

Image

It would be nice for this to be a bit faster.

## Details

### Environment
* [x] `clickhouse-go` version: v2.41.0
* [x] Interface: ClickHouse API
* [x] Go version: go1.25.4
* [x] Operating system: Debian running Linux v6.19.0-rc1
* [x] ClickHouse version: v25.10.3.100
* [x] Is it a ClickHouse Cloud? It is not
* [x] ClickHouse Server non-default settings, if any: none
* [ ] `CREATE TABLE` statements for tables involved: not necessarily relevant
* [ ] Sample data for all these tables, use [clickhouse-obfuscator](https://github.com/ClickHouse/ClickHouse/blob/master/programs/obfuscator/Obfuscator.cpp#L42-L80) if necessary

Contributor guide

Open the contributing guide

Research direction

Start with the provided Go reproduction against clickhouse-go v2.41.0, comparing rows.Next() alone with rows.Scan(&frames, &value) for the large ARRAY JOIN result. Use the flamegraphs and timing measurements as the baseline; done means scanning the same result is measurably faster without changing returned values or scan behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
clickhouse, go
Domain
database
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.