Altinity / Altinity/clickhouse-operator

Bug: `SYSTEM DROP DATABASE REPLICA` uses wrong shard identifier during scale-down of Replicated database with custom shard/replica names

Open
#1,936 2 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Go
Stars
2.6k
Forks
574
Avg merge
8d 6h
Merged PRs (30d)
6

Description

Summary

When scaling down a ClickHouseInstallation that uses Replicated database engine, the operator generates an incorrect SYSTEM DROP DATABASE REPLICA command. The stale replica metadata is not cleaned from ClickHouse Keeper, causing REPLICA_ALREADY_EXISTS errors on subsequent scale-up.

Environment

  • ClickHouse Operator version: latest (tested on current main)
  • ClickHouse version: 25.3.2.1
  • Database engine: Replicated

Steps to Reproduce

  1. Create a ClickHouseInstallation with 1 shard, 1 replica using Replicated database engine:
    CREATE DATABASE myDB ON CLUSTER 'all-replicated' ENGINE = Replicated('/clickhouse/databases/myDB', '{all-sharded-shard}-{shard}', '{replica}')
    
  2. Scale up to 2 replicas
  3. Scale back down to 1 replica
  4. Scale up again to 2 replicas

At step 4, the new replica fails with:

Code: 253. DB::Exception: There was an error on [chi-clickhouse-scw-0-1:9000]:
Code: 253. DB::Exception: Replica /clickhouse/databases/default/replicas/1-0|chi-clickhouse-scw-0-1
already exists. (REPLICA_ALREADY_EXISTS)

Root Cause

In pkg/model/chi/schemer/sql.go:244-248, the sqlDropReplica function generates:

func (s *ClusterSchemer) sqlDropReplica(shard int, replica string) []string {
    return []string{
        fmt.Sprintf("SYSTEM DROP REPLICA '%s'", replica),
        fmt.Sprintf("SYSTEM DROP DATABASE REPLICA '%d|%s'", shard, replica),
    }
}

The shard parameter is an int (0-based ShardIndex), so the generated command is:

SYSTEM DROP DATABASE REPLICA '0|chi-clickhouse-scw-0-1'

However, the Replicated database engine is created with macros for the shard identifier (e.g. {all-sharded-shard}-{shard}), which resolve to a string like 1-0. ClickHouse Keeper stores replica entries using this resolved database_shard_name (visible in system.clusters), not the integer shard index. The actual znode path is:

/clickhouse/databases/default/replicas/1-0|chi-clickhouse-scw-0-1

So the correct command should be:

SYSTEM DROP DATABASE REPLICA '1-0|chi-clickhouse-scw-0-1'

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in pkg/model/chi/schemer/sql.go at sqlDropReplica and inspect how the shard identifier is represented when the Replicated database uses custom shard and replica names. Verify the generated SYSTEM DROP DATABASE REPLICA command against the resolved identifier shown in the issue, then reproduce the scale-down and scale-up sequence to confirm stale replica metadata is removed.

Written by the indexing model from the issue text.

Assessment

Tech stack
go, kubernetes
Domain
databases
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.