[Stress] OOMKilled errors are happening sporadically
Open
Stress
- Dominant language
- C#
- Stars
- 135
- Forks
- 260
- Avg merge
- 3d 1h
- Merged PRs (30d)
- 143
Description
We're seeing sporadic OOMKilled errors in pods that don't appear to be pushing their `request` limits. There's been some flux in the baseline memory usage of our Kubenetes nodes and we've noticed other issues (like restarted system pods) that makes it seem like we might need to tweak or adjust our base SKU to handle it.
Contributor guide
Research direction
The issue names no files, tests, or entry points. Begin by examining pod memory limits, Kubernetes node baseline usage, restarted system pods, and the current base SKU; confirm whether resource pressure explains the OOMKilled events. Done means identifying the cause and documenting or validating a specific remediation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes
- Domain
- infrastructure
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100