Retina agent pods crash looping after node upgrade (AKS)
- Dominant language
- Go
- Stars
- 3.2k
- Forks
- 304
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 78
Description
**Describe the bug**
After a node upgrade to AKSUbuntu-2004gen2fipscontainerd-202503.02.0 my retina agents are crash looping.
**To Reproduce**
Steps to reproduce the behavior:
1. Create a private AKS cluster (v1.30.4) with Azure monitor enabled
2. Ensure your system node pool uses the AKSUbuntu-2004gen2fipscontainerd-202503.02.0 node image version
3. Observe the retina agent pods crash looping.
**Expected behavior**
The retina service to run as expected.
**Screenshots**
**Platform (please complete the following information):**
- OS: AKSUbuntu-2004gen2fipscontainerd-202503.02.0
- Kubernetes Version: [e.g. 1.30.4]
- Host: AKS
- Retina Version: v0.0.25
**Additional context**
Add any other context about the problem here.
Contributor guide
Research direction
Reproduce the crash loop on a private AKS cluster using Kubernetes 1.30.4, Azure Monitor, and node image AKSUbuntu-2004gen2fipscontainerd-202503.02.0 with Retina v0.0.25. Start by inspecting the agent pod status and logs; done means the Retina agent pods remain running and the service operates as expected.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- azure, go, kubernetes
- Domain
- infrastructure, networking, observability
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 28/100