globus / globus/globus-compute

Properly kill manager when a ManagerLost problem happens on k8s

Open
#255 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
162
Forks
53
Avg merge
15h 29m
Merged PRs (30d)
26

Description

Currently interchange does not force a manager on k8s to kill when a ManagerLost problem happens on k8s, and the manager will keep crashloop and stay there. Relevant to #254

Contributor guide

Open the contributing guide

Research direction

Start by tracing the interchange manager lifecycle for the Kubernetes ManagerLost problem described in this issue. Confirm the expected behavior for a lost manager and verify that a manager in this state is terminated instead of remaining in a crash loop.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, python
Domain
infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.