Altinity / Altinity/clickhouse-operator
Maintenance & Failover
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 2.6k
- Forks
- 574
- Avg merge
- 8d 6h
- Merged PRs (30d)
- 6
Description
We're running a Kubernetes cluster with three worker nodes and are planning to deploy ClickHouse and ClickHouse Keeper version 25.8 using the Altinity ClickHouse Operator 0.25.3. However, I couldn’t find detailed documentation regarding maintenance tasks.
Specifically, I’d like to know:
- Is there any official guidance on temporarily or permanently removing a ClickHouse node from the cluster for maintenance purposes in Kubernetes?
- If a node fails or its disk becomes corrupted, will adding a new node with a new disk to the ClickHouse cluster automatically trigger replication to that node for data recovery?
- Are there any official documents or example manifests for enabling in-transit and at-rest encryption for both ClickHouse and ClickHouse Keeper to ensure secure communication?
Any references or best practices would be greatly appreciated.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the three requested areas in the issue: Kubernetes maintenance and node removal, recovery when replacing a failed or corrupted node, and encryption for ClickHouse and ClickHouse Keeper. Identify the relevant official documentation and example manifests; done means the issue's questions are answered with authoritative references or concrete examples.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes
- Domain
- databases, documentation, infrastructure
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100