[EPIC] Workspace Pod/Node Metrics
- Dominant language
- No language data
- Stars
- 84
- Forks
- 149
- Avg merge
- 5d 15h
- Merged PRs (30d)
- 29
Description
### Certification
- [x] I certify I am an Epic Owner for Kubeflow Notebooks 2.0 and expected to create planning-related issues.
### User Story
Bella needs to understand how her Kubernetes workloads are performing. She can easily observe current basic resource usage, such as CPU and memory utilization, to gain insights into her workload's efficiency. This allows her to optimize resource allocation for her models and experiments, ensuring she gets the most out of the available resources. In the event a Metrics Server is not deployed in her k8s cluster, it should be obvious why the metrics aren't available.
Joel needs to be able to observe basic Kubernetes workload resource usage, such as CPU and memory utilization, across the cluster when a Metrics Server is deployed. He can proactively identify potential issues and ensure that the platform continues to operate smoothly, even in capacity-constrained environments. This capability is crucial for him to guarantee a consistent user experience and prevent disruptions.
Contributor guide
Research direction
Start by decomposing the two user stories into the required workspace, pod, and node metrics, including the behavior when a Metrics Server is unavailable. Identify the relevant Kubeflow Notebooks entry points and tests before deciding how to implement the epic; done means the described CPU and memory visibility works across the cluster and the unavailable-server state is obvious.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes
- Domain
- infrastructure, observability
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100