kernelci / kernelci/kernelci-pipeline

Proposed Pipeline Improvements for Lab Operations

Open
#1,419 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
11
Forks
40
Avg merge
2d 13h
Merged PRs (30d)
14

Description

I am planning significant updates to the lab operations pipeline to address current visibility and reliability issues.

# Weekly Lab Statistics

We will generate public reports (viewable on Maestro) for lab owners, we can make them private by sending over email only, i need feedback from lab owners on this. These reports will include:

- Availability: Early failures for labs (online checks) and specific device submissions.
- Volume: Total submitted jobs and jobs per device type.
- Outcomes: Result breakdown per device type (OK, Fail, Infrastructure error, Early failure).

# Availability Alerts
We will implement notifications for lab owners when a lab is unavailable for more than N hours. Might be even device, but i am not sure it worth it.

# Optimization Plan (Based on Collected Data)

Prioritize High-Value Tests: Ensure sufficient device coverage for tests with active stakeholders.
Load Balancing: Reduce load on overloaded devices and improve tests distribution.
Reliability: Prioritize successful completion over broad coverage (i.e., fewer devices but higher reliability). The same as build numbers, many developers expect to see N tests, X fail, Y pass, so they see the trend. But when number of test results are unstable, it is hard to see the trend and it is hard to trust the results.
Job/Test Scoring: Estimate runtime based on history. Labs may block non-whitelisted tests that exceed duration thresholds (e.g., >3 hours).

# Extended observability
Maestro dashboard where lab can see current errors, latest submission details, "stale" answers and other possible problems. This will help lab owners to quickly identify and address issues.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are identified; start by locating the existing lab operations, reporting, alerting, and Maestro dashboard components. Before implementation, clarify the report visibility, alert thresholds, optimization scope, scoring rules, and dashboard requirements, then define completion criteria for each proposed improvement.

Written by the indexing model from the issue text.

Assessment

Domain
observability
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.