temporalio / temporalio/sdk-java
Make activity heartbeats more robust to network outages by retrying them
Nobody has claimed this yet.
- Dominant language
- Java
- Stars
- 433
- Forks
- 249
- Avg merge
- 5d 6h
- Merged PRs (30d)
- 26
Description
Right now we don't retry failed heartbeats and just ignore the error.
This may create a problem if heartbeat timeout is, for example, 20 seconds and activity heartbeats every 15 seconds. If we get a network blip, such an activity will be terminated with heartbeat timeout.
We should use heartbeatExecutor to schedule an asynchronous retry for a failed heartbeat. We can utilize an exponential backoff throttling strategy for it. ActivityContext.heartbeat() shouldn't be kept blocked because we retry.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with ActivityContext.heartbeat() and the heartbeatExecutor path described in the issue. Trace how failed heartbeats are currently handled, then verify that retries use exponential backoff, run asynchronously, and do not block the heartbeat call. Done means a transient network outage does not cause heartbeat timeout while retries remain bounded by the existing execution behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- distributed-systems
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100