temporalio / temporalio/sdk-java

Make activity heartbeats more robust to network outages by retrying them

Open
#1,258 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Java
Stars
433
Forks
249
Avg merge
5d 6h
Merged PRs (30d)
26

Description

Right now we don't retry failed heartbeats and just ignore the error.
This may create a problem if heartbeat timeout is, for example, 20 seconds and activity heartbeats every 15 seconds. If we get a network blip, such an activity will be terminated with heartbeat timeout.
We should use heartbeatExecutor to schedule an asynchronous retry for a failed heartbeat. We can utilize an exponential backoff throttling strategy for it. ActivityContext.heartbeat() shouldn't be kept blocked because we retry.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with ActivityContext.heartbeat() and the heartbeatExecutor path described in the issue. Trace how failed heartbeats are currently handled, then verify that retries use exponential backoff, run asynchronously, and do not block the heartbeat call. Done means a transient network outage does not cause heartbeat timeout while retries remain bounded by the existing execution behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
distributed-systems
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.