NVIDIA / NVIDIA/nvcf

Keep unhealthy function pods available briefly for diagnostics

Open
#84 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

roadmap
Dominant language
Go
Stars
218
Forks
72
Avg merge
1d 12h
Merged PRs (30d)
427

Description

Description

When a function health check fails, rapid pod termination can prevent users from collecting logs, telemetry, and other failure evidence. Add a per-function diagnostic grace period that can keep an unhealthy function pod running for up to 10 minutes before termination.

Definition of Done

  • A function can be configured with a diagnostic grace period of up to 10 minutes.
  • After a health check failure, the affected pod remains running for the configured period before termination.
  • Logs and telemetry remain collectable during the grace period, and self-managed operators can inspect the pod with Kubernetes tools.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing per-function configuration, health-check failure handling, and the Kubernetes pod termination path. Determine where the diagnostic grace period can be configured and enforced, then verify that a failed pod remains inspectable for the configured period and is terminated afterward, with logs and telemetry still available.

Written by the indexing model from the issue text.

Assessment

Tech stack
go, kubernetes
Domain
backend, cloud, infrastructure
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.