apache / apache/pinot

Pinot Server Graceful Shutdown Improvements

Open
#10,876 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
6.1k
Forks
1.5k
Avg merge
2d 55m
Merged PRs (30d)
182

Description

The current graceful shutdown steps are listed below in order (along with the issues):

1. `PinotFSFactory` is shutdown first. **Issue**: This will fail segment upload for the segments that are committing when a server shutdown is happening.
2. HelixManager disconnects which shuts down the ZkClient. **Issue**: This happens before segment data managers are destroyed, so if there is a segment that tries to do a commit after the disconnect is called, it will throw because the ZkClient is already shutdown.
3. Server instance shutdown issues destroy for all `SegmentDataManager`. For realtime tables, this attempts to stop the ingestion by joining with the consumer thread for each consuming segment. (no issue with this)

All this means that we could run into scenarios where a segment goes into error-state, because the deep-store link for the segment is missing and peer download wouldn't work because server restarts usually take at least 2 minutes or more and current retry logic only waits 10-15s for an Online peer (by default the replica will take 31 seconds to catch-up to the final offset. if it can't, then we try to download the segment instead).

If instead of a server restart it was a host failure, then we could also have data loss.

I am working offline with some stakeholders for the fixes.

image

Exception seen due to closed ZkClient:

```
2023-06-09 03:00:29.552 [some_table__242__1314__20230609T0216Z] ERROR o.a.p.c.d.m.r.LLRealtimeSegmentDataManager_some_table__242__1314__20230609T0216Z - Exception while in work
java.lang.IllegalStateException: ZkClient already closed!
at org.apache.helix.zookeeper.zkclient.ZkClient.retryUntilConnected(ZkClient.java:1977)
at org.apache.helix.zookeeper.zkclient.ZkClient.readData(ZkClient.java:2139)
at org.apache.helix.zookeeper.zkclient.ZkClient.readData(ZkClient.java:2131)
at org.apache.helix.manager.zk.ZkBaseDataAccessor.get(ZkBaseDataAccessor.java:495)
at org.apache.helix.manager.zk.ZkCacheBaseDataAccessor.get(ZkCacheBaseDataAccessor.java:397)
at org.apache.helix.store.zk.AutoFallbackPropertyStore.get(AutoFallbackPropertyStore.java:101)
at org.apache.pinot.common.metadata.ZKMetadataProvider.getTableConfig(ZKMetadataProvider.java:308)
at org.apache.pinot.core.data.manager.realtime.RealtimeTableDataManager.replaceLLSegment(RealtimeTableDataManager.java:665)
at org.apache.pinot.core.data.manager.realtime.LLRealtimeSegmentDataManager.commitSegment(LLRealtimeSegmentDataManager.java:1012)
at org.apache.pinot.core.data.manager.realtime.LLRealtimeSegmentDataManager$PartitionConsumer.run(LLRealtimeSegmentDataManager.java:728)
at java.base/java.lang.Thread.run(Thread.java:829)
```

cc: @Jackie-Jiang

Contributor guide

Open the contributing guide

Research direction

Start by tracing the Pinot server shutdown sequence around PinotFSFactory, HelixManager, and SegmentDataManager, then inspect the realtime commit path shown in the LLRealtimeSegmentDataManager stack trace. The work is done when shutdown ordering prevents segment upload failures and closed-ZkClient errors during in-flight commits without leaving segments in error state.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
databases, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.