GoogleContainerTools / GoogleContainerTools/skaffold

Support tolerateFailuresUntilDeadline for Helm deployments

Open
#9,809 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
15.9k
Forks
1.7k
Avg merge
3d 9h
Merged PRs (30d)
10

Description

Could you consider adding support for the `tolerateFailuresUntilDeadline` field for Helm deployments in Skaffold?

**Context**
In GKE Autopilot clusters, Helm deployments sometimes fail in Skaffold due to delays caused by node autoscaling. For example, if a node is deleted during a deployment, the associated pod needs to be recreated on a new node. This process can take some time.

Even though Kubernetes eventually recreates the pod and the deployment completes successfully, Skaffold may already report the deployment as failed.

**Why this is needed**
Currently, there’s no mechanism for Helm deployments in Skaffold to tolerate temporary scheduling issues. Supporting `tolerateFailuresUntilDeadline` for Helm—similar to what was introduced for Cloud Run in [v2.16.0](https://github.com/GoogleContainerTools/skaffold/blob/main/CHANGELOG.md#v2160-release---05022025) — would allow Skaffold to wait before marking the deployment as failed.

This would make deployments in Autopilot environments more resilient and improve reliability in GitHub Actions and other CI/CD setups.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.