[Bug] cloudberry cluster becoming unavailable
- Dominant language
- C
- Stars
- 1.4k
- Forks
- 247
- Avg merge
- 4d 3h
- Merged PRs (30d)
- 39
Description
### Apache Cloudberry version
apache cloudberry 2.0.0
### What happened
Hi Team,
First there was segment crash, where the mirror became active and primary is in failed state, then on top of that there is another crash on the same segment where primary became active and mirror went into failed state.
As soon as the primary became active and mirror went to failed state for the second time crash, the cloudberry cluster became unavailable because the **Standby. Signal** file was not removed from the primary segment instance data directory and we need to stop and start the cloudberry cluster to make it available.
Can you please kindly provide your input\solution on this.
### What you think should happen instead
_No response_
### How to reproduce
first crash a segment and then immediately conduct one more crash on the same segment,
### Operating System
OEL 9.7
### Anything else
_No response_
### Are you willing to submit PR?
- [ ] Yes, I am willing to submit a PR!
### Code of Conduct
- [x] I agree to follow this project's [Code of Conduct](https://github.com/apache/cloudberry/blob/main/CODE_OF_CONDUCT.md).
Contributor guide
Research direction
No files or tests are named. First reproduce the two successive crashes on one segment in Apache Cloudberry 2.0.0, then trace segment failover and handling of the Standby.Signal file when the primary and mirror swap states. Done means the cluster remains available without requiring a stop and start after the second crash.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- c, postgresql
- Domain
- databases, distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100