ContextLab / ContextLab/clustrix
Restore Kubernetes execution and provisioning: removed from v0.2.0 as never verified against a real cluster
- Dominant language
- Python
- Stars
- 10
- Forks
- 4
- Avg merge
- 6h 27m
- Merged PRs (30d)
- 9
Description
Kubernetes execution and auto-provisioning are implemented and removed from v0.2.0 as unverified. No job has ever been run against a real cluster.
## What exists, and where
| Code | Lines | Introduced |
|-|-|-|
| `clustrix/executor_kubernetes.py` | 594 | `6061247` (2025-08-23) |
| `clustrix/kubernetes/cluster_provisioner.py` | 423 | `d6b4b7b` (2025-08-16) |
| `clustrix/kubernetes/local_provisioner.py` | 455 | `d6b4b7b` |
| `clustrix/kubernetes/{aws,azure,gcp,lambda,huggingface}_provisioner.py` | 3,611 | `d6b4b7b` |
| `clustrix/kubernetes/__init__.py` | 60 | `d6b4b7b` |
Plus the widget's Kubernetes section, the `k8s_*` configuration fields, and `docs/source/notebooks/kubernetes_tutorial.ipynb`.
## Defects found while auditing
Two were real and are worth carrying forward whenever this is restored:
1. **`check_k8s_job_status` reported failed jobs as successful.** Its exception paths returned `"completed"`, so a job that died came back as a normal completion.
2. **Results were decoded with `ast.literal_eval` on the pod log**, falling back to returning the raw log text as if it were the function's return value.
Both were fixed during the v0.2.0 sweep, and neither fix has been run against a real cluster.
## Note on scope
The five cloud provisioners under `clustrix/kubernetes/` are removed with this issue rather than with the per-provider issues, because they are Kubernetes cluster provisioners, not the VM-based cloud execution backends. See the AWS/GCP/Azure/Lambda issues for those.
## Why it is being removed rather than fixed
Not because the code is known to be wrong. Because it has **never been run against the real thing**, and shipping it in the cluster-type dropdown states otherwise. A user who selects it gets a code path no one has ever seen succeed.
v0.2.0 keeps exactly the four backends that have been demonstrated end to end -- `local`, `ssh`, `slurm`, `huggingface` -- and the documentation now says the rest are planned for a future release rather than currently supported.
## Restoring it
Nothing is lost: every line cited above stays reachable in git history at the commits named. Reinstating it means reverting the removal commit and then doing the part that was never done -- running it against real hardware and recording the evidence in this issue.
## Definition of done
- [ ] Backend restored from history
- [ ] A real job submitted, executed and its result returned, with the transcript pasted into this issue
- [ ] Failure paths exercised (job rejected, job killed, node lost)
- [ ] Re-added to `SUPPORTED_CLUSTER_TYPES`, the widget dropdown and the CLI
- [ ] Documentation moved from "planned" to "supported"
Contributor guide
Research direction
Start by reading the historical implementations in clustrix/executor_kubernetes.py and clustrix/kubernetes/, using commits 6061247 and d6b4b7b, then inspect the widget, k8s_* fields, CLI, and Kubernetes tutorial. Restore the backend and registrations, run a real job, exercise rejected, killed, and node-lost failures, and paste the execution transcript before moving the documentation to supported.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes, python
- Domain
- devops, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100