[BYDB-Replica] Support Liaison Groups with Global-Level Failover in Liaison
- Dominant language
- Java
- Stars
- 25k
- Forks
- 6.6k
- Avg merge
- 9h 32m
- Merged PRs (30d)
- 19
Description
### Search before asking
- [X] I had searched in the [issues](https://github.com/apache/skywalking/issues?q=is%3Aissue) and found no similar feature requirement.
### Description
Based upon the existing Load Balancer(#12874) feature in BanyanDB's Liaison component, we would like to propose the addition of **Liaison Groups**. This enhancement will enable the creation of logically grouped liaison nodes, facilitating more granular load balancing and robust failover strategies. By organizing liaison nodes into groups, the BanyanDB can ensure higher availability, better resource management, and improved resilience against group-level failures.
**Key Requirements:**
- **Group Definition:** Allow administrators to define multiple liaison groups, each containing several liaison nodes.
- **Group Availability Status:** Continuously assess the availability of each liaison group. A group is **unavailable** if the number of healthy data nodes falls below the specified replica number.
- **Automatic Failover:** If a group becomes unavailable, traffic will be automatically redirected to other healthy groups to maintain service continuity.
- **Health Reporting:** Upstream groups must communicate their availability status to the global liaison to inform traffic distribution decisions.
### Use case
_No response_
### Related issues
_No response_
### Are you willing to submit a pull request to implement this on your own?
- [ ] Yes I am willing to submit a pull request on my own!
### Code of Conduct
- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)
Contributor guide
Research direction
Start by reading the existing Load Balancer feature referenced in issue #12874 and the Liaison component's current behavior. Trace how liaison nodes report health and how traffic is distributed, then define how groups, replica-based availability, global failover, and upstream health reporting should fit together. Done means unavailable groups redirect traffic to healthy groups while preserving service continuity.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- databases, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100