Azure / Azure/azure-sdk-tools

[Stress] OOMKilled errors are happening sporadically

Open
#5,910 0 comments 0 reactions 0 assignees View on GitHub
Stress
Dominant language
C#
Stars
135
Forks
260
Avg merge
3d 1h
Merged PRs (30d)
143

Description

We're seeing sporadic OOMKilled errors in pods that don't appear to be pushing their `request` limits. There's been some flux in the baseline memory usage of our Kubenetes nodes and we've noticed other issues (like restarted system pods) that makes it seem like we might need to tweak or adjust our base SKU to handle it.

Contributor guide

Open the contributing guide

Research direction

The issue names no files, tests, or entry points. Begin by examining pod memory limits, Kubernetes node baseline usage, restarted system pods, and the current base SKU; confirm whether resource pressure explains the OOMKilled events. Done means identifying the cause and documenting or validating a specific remediation.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes
Domain
infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.