GoogleCloudPlatform / GoogleCloudPlatform/spring-cloud-gcp

[spring-cloud-gcp-autoconfigure] PubSubHealthIndicator Fails Under Large GCP Pub/Sub Backlogs, Triggering Negative Feedback Loop

Open
#3,438 4 comments 0 reactions 0 assignees View on GitHub
priority: p3 type: bug
Dominant language
Java
Stars
551
Forks
349
Avg merge
1d 13h
Merged PRs (30d)
14

Description

**Describe the bug**

We’re experiencing periodic spikes in latency from the `PubSubHealthIndicator` in **Spring Cloud GCP** whenever there’s a large backlog in **Google Cloud Pub/Sub**. Although the backlog itself isn’t being pulled by the health check, the overall **Pub/Sub system (or network)** slows down enough that the “quick pull” call hangs or times out. This marks our service as **DOWN** in `/actuator/health`, which can **trigger restarts in Kubernetes**, creating a negative feedback loop.

**Logs:**

```
Health contributor org.springframework.cloud.gcp.autoconfigure.pubsub.health.PubSubHealthIndicator (pubSub) took 10936ms to respond
Health contributor org.springframework.cloud.gcp.autoconfigure.pubsub.health.PubSubHealthIndicator (pubSub) took 89844ms to respond
```

Im not sure if updating the health-check settings (e.g., timeouts) would resolve this issue, or if we should exclude the `PubSubHealthIndicator` from the group of core `/actuator/health`.
Since Pub/Sub is designed to handle backlogs to protect our service from being overwhelmed during high traffic periods, relying on a “quick pull” to measure health may not be best practice in production.
btw, I’m not entirely sure why the “quick pull” call either hangs or times out when other topics start having backlogs.

Any guidance or recommended patterns on handling these scenarios while still monitoring Pub/Sub health would be greatly appreciated.

Contributor guide

Open the contributing guide

Research direction

Start with the PubSubHealthIndicator entry point and the /actuator/health behavior described in the issue. Trace the “quick pull” health check and review how its timeout or failure marks Pub/Sub health as DOWN; done should define a safe production behavior for large Pub/Sub backlogs without causing unnecessary service restarts.

Written by the indexing model from the issue text.

Assessment

Tech stack
gcp, java, kubernetes, spring
Domain
backend, cloud, devops, observability
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.