apache / apache/beam

[Bug]: HBase read could stuck at HBaseReader.advance indefinitely

Open
#28,025 2 comments 0 reactions 0 assignees View on GitHub
awaiting triage bug hbase io java P2 pinned
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

### What happened?

There is report that HBaseIO read get stuck intermittently, possible due to failed connection never recovered

At here:

https://github.com/apache/beam/blob/16e7af2ea4148c65e3e1045e6c05f0d40f498db2/sdks/java/io/hbase/src/main/java/org/apache/beam/sdk/io/hbase/HBaseIO.java#L535C21-L535C21

```
Operation ongoing in bundle process_bundle-8636262876883182297-1613 for PTransform{id=Read-from-HBase-Read-HBaseSource--ParDo-BoundedSourceAsSDFWrapper--ParMultiDo-BoundedSourceAsSDFWrap/ProcessElementAndRestrictionWithSizing-ptransform-66, name=Read-from-HBase-Read-HBaseSource--ParDo-BoundedSourceAsSDFWrapper--ParMultiDo-BoundedSourceAsSDFWrap/ProcessElementAndRestrictionWithSizing-ptransform-66, state=process} for at
least 01h15m05s without outputting or completing: at
java.lang.Object.wait(Native Method) at
java.lang.Object.wait(Object.java:460) at
java.util.concurrent.TimeUnit.timedWait(TimeUnit.java:348) at
org.apache.hadoop.hbase.client.ResultBoundedCompletionService.poll(ResultBoundedCompletionService.java:159) at
org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:198) at
org.apache.hadoop.hbase.client.ScannerCallableWithReplicas.call(ScannerCallableWithReplicas.java:60) at
org.apache.hadoop.hbase.client.RpcRetryingCaller.callWithoutRetries(RpcRetryingCaller.java:200) at
org.apache.hadoop.hbase.client.ClientScanner.call(ClientScanner.java:320) at
org.apache.hadoop.hbase.client.ClientScanner.loadCache(ClientScanner.java:403) at
org.apache.hadoop.hbase.client.ClientScanner.next(ClientScanner.java:364) at
org.apache.hadoop.hbase.client.AbstractClientScanner$1.hasNext(AbstractClientScanner.java:94) at
org.apache.beam.sdk.io.hbase.HBaseIO$HBaseReader.advance(HBaseIO.java:514) at
org.apache.beam.sdk.io.Read$BoundedSourceAsSDFWrapperFn$BoundedSourceAsSDFRestrictionTracker.tryClaim(Read.java:366) at
...
java.lang.Thread.run(Thread.java:750)
```

Shall we add a timeout to advance?

### Issue Priority

Priority: 2 (default / most bugs should be filed as P2)

### Issue Components

- [ ] Component: Python SDK
- [X] Component: Java SDK
- [ ] Component: Go SDK
- [ ] Component: Typescript SDK
- [X] Component: IO connector
- [ ] Component: Beam examples
- [ ] Component: Beam playground
- [ ] Component: Beam katas
- [ ] Component: Website
- [ ] Component: Spark Runner
- [ ] Component: Flink Runner
- [ ] Component: Samza Runner
- [ ] Component: Twister2 Runner
- [ ] Component: Hazelcast Jet Runner
- [ ] Component: Google Cloud Dataflow Runner

Contributor guide

Open the contributing guide

Research direction

Start with HBaseIO.java at the referenced HBaseReader.advance location and compare it with the HBase client stack trace, especially the blocking scanner calls. Investigate how a failed connection can recover or time out, then determine whether advance can be bounded safely. Done means a failed or stalled read no longer waits indefinitely, with regression coverage for the observed behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.