oxidecomputer / oxidecomputer/propolis
talos linux showing as not ready after using `kexec` during decommissioning
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 270
- Forks
- 42
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 6
Description
@rothgar was testing Talos Linux on Oxide and noticed that, when decommissioning Talos Linux nodes, some nodes would remain in the Not Ready state in Omni, a Talos Linux management program, after attempting to use kexec to reboot after wiping itself.
During a troubleshooting session we collected propolis logs for an instance that failed to reboot with kexec (talos-10) and an instance that successfully rebooted with kexec disabled (talos-11).
- talos-10-oxz_propolis-server_7bcbb1b1-9f72-40e6-aecd-ca49995dc0ac.log
- talos-11-oxz_propolis-server_9468a2bd-ec9a-46da-bf1c-7f4c15706d5f.log
However, it's not yet clear what's going on here. The minimum reproduction case is still not guaranteed since maybe 3 nodes out of 25 experience this behavior when decommissioning. The following items are outstanding and would help narrow down what's happening here.
- Narrow down and document the reproduction case to reliably reproduce this issue at a higher yield. This would help in gathering guest OS logs and Oxide logs.
- Gather more guest OS information during a reproduction.
- Output of
dmesg. Can this be reliably retrieved viatalosctlwhen the instance is decommissioning? - Serial console output.
- Output of
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by comparing the attached propolis logs for talos-10 and talos-11, then investigate the decommissioning path that uses kexec. Attempt to make the reproduction reliable at a higher yield and collect guest dmesg through talosctl and serial console output. Done means the failure is reproducible and the relevant guest and Oxide logs are documented.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- linux, rust
- Domain
- operating-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100