microsoft / microsoft/retina

Retina agent pods crash looping after node upgrade (AKS)

Open
#1,439 4 comments 2 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
3.2k
Forks
304
Avg merge
1d 19h
Merged PRs (30d)
78

Description

**Describe the bug**
After a node upgrade to AKSUbuntu-2004gen2fipscontainerd-202503.02.0 my retina agents are crash looping.

**To Reproduce**
Steps to reproduce the behavior:

1. Create a private AKS cluster (v1.30.4) with Azure monitor enabled
2. Ensure your system node pool uses the AKSUbuntu-2004gen2fipscontainerd-202503.02.0 node image version
3. Observe the retina agent pods crash looping.

**Expected behavior**
The retina service to run as expected.

**Screenshots**

Image

**Platform (please complete the following information):**

- OS: AKSUbuntu-2004gen2fipscontainerd-202503.02.0
- Kubernetes Version: [e.g. 1.30.4]
- Host: AKS
- Retina Version: v0.0.25

**Additional context**
Add any other context about the problem here.

Contributor guide

Open the contributing guide

Research direction

Reproduce the crash loop on a private AKS cluster using Kubernetes 1.30.4, Azure Monitor, and node image AKSUbuntu-2004gen2fipscontainerd-202503.02.0 with Retina v0.0.25. Start by inspecting the agent pod status and logs; done means the Retina agent pods remain running and the service operates as expected.

Written by the indexing model from the issue text.

Assessment

Tech stack
azure, go, kubernetes
Domain
infrastructure, networking, observability
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.