ContextLab / ContextLab/clustrix

Restore Kubernetes execution and provisioning: removed from v0.2.0 as never verified against a real cluster

Open
#142 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
10
Forks
4
Avg merge
6h 27m
Merged PRs (30d)
9

Description

Kubernetes execution and auto-provisioning are implemented and removed from v0.2.0 as unverified. No job has ever been run against a real cluster.

## What exists, and where

| Code | Lines | Introduced |
|-|-|-|
| `clustrix/executor_kubernetes.py` | 594 | `6061247` (2025-08-23) |
| `clustrix/kubernetes/cluster_provisioner.py` | 423 | `d6b4b7b` (2025-08-16) |
| `clustrix/kubernetes/local_provisioner.py` | 455 | `d6b4b7b` |
| `clustrix/kubernetes/{aws,azure,gcp,lambda,huggingface}_provisioner.py` | 3,611 | `d6b4b7b` |
| `clustrix/kubernetes/__init__.py` | 60 | `d6b4b7b` |

Plus the widget's Kubernetes section, the `k8s_*` configuration fields, and `docs/source/notebooks/kubernetes_tutorial.ipynb`.

## Defects found while auditing

Two were real and are worth carrying forward whenever this is restored:

1. **`check_k8s_job_status` reported failed jobs as successful.** Its exception paths returned `"completed"`, so a job that died came back as a normal completion.
2. **Results were decoded with `ast.literal_eval` on the pod log**, falling back to returning the raw log text as if it were the function's return value.

Both were fixed during the v0.2.0 sweep, and neither fix has been run against a real cluster.

## Note on scope

The five cloud provisioners under `clustrix/kubernetes/` are removed with this issue rather than with the per-provider issues, because they are Kubernetes cluster provisioners, not the VM-based cloud execution backends. See the AWS/GCP/Azure/Lambda issues for those.

## Why it is being removed rather than fixed

Not because the code is known to be wrong. Because it has **never been run against the real thing**, and shipping it in the cluster-type dropdown states otherwise. A user who selects it gets a code path no one has ever seen succeed.

v0.2.0 keeps exactly the four backends that have been demonstrated end to end -- `local`, `ssh`, `slurm`, `huggingface` -- and the documentation now says the rest are planned for a future release rather than currently supported.

## Restoring it

Nothing is lost: every line cited above stays reachable in git history at the commits named. Reinstating it means reverting the removal commit and then doing the part that was never done -- running it against real hardware and recording the evidence in this issue.

## Definition of done

- [ ] Backend restored from history
- [ ] A real job submitted, executed and its result returned, with the transcript pasted into this issue
- [ ] Failure paths exercised (job rejected, job killed, node lost)
- [ ] Re-added to `SUPPORTED_CLUSTER_TYPES`, the widget dropdown and the CLI
- [ ] Documentation moved from "planned" to "supported"

Contributor guide

Open the contributing guide

Research direction

Start by reading the historical implementations in clustrix/executor_kubernetes.py and clustrix/kubernetes/, using commits 6061247 and d6b4b7b, then inspect the widget, k8s_* fields, CLI, and Kubernetes tutorial. Restore the backend and registrations, run a real job, exercise rejected, killed, and node-lost failures, and paste the execution transcript before moving the documentation to supported.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, python
Domain
devops, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.