bottlerocket-os / bottlerocket-os/bottlerocket

aws-k8s-nvidia variant fails to boot on g7 instance family: ghostdog selects tesla driver but GPU (10de:2c3a) requires open kernel modules

Open
#4,889 2 comments 5 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
9.7k
Forks
586
Avg merge
1d 11h
Merged PRs (30d)
11

Description

## Image I'm using

`bottlerocket-aws-k8s-1.36-nvidia-x86_64-v1.62.1` (ami resolved via Karpenter alias `bottlerocket@latest`, us-east-1)

## What I expected to happen

Nodes from the **g7 instance family** (Blackwell-generation GPU, PCI ID `10de:2c3a`) boot, load a working NVIDIA driver, and join the EKS cluster.

## What actually happened

Boot fails before kubelet ever starts, so the node never joins the cluster (Karpenter NodeClaim stuck at `Registered: Unknown — Node not registered with cluster`, then relaunch loop). From the EC2 console log:

```
ghostdog[2336]: Error: grid is not preferred driver: tesla
ghostdog[2337]: Error: open-gpu is not preferred driver: tesla
...
nvidia-migmanager[2755]: '/usr/bin/nvidia-smi' failed - stderr:
[FAILED] Failed to start NVIDIA MIG manager service.
[DEPEND] Dependency failed for Generate CDI specifications.
[DEPEND] Dependency failed for Bottlerocket initial configuration complete.
[DEPEND] Dependency failed for Activate configured.target.
```

`ghostdog` selects the proprietary **tesla** driver for this GPU. The g7 family's GPU requires the NVIDIA **open** kernel modules — the proprietary driver rejects it. (Verified on AL2023 with the same 580-series driver: proprietary logs `NVRM: The NVIDIA GPU ... (PCI ID: 10de:2c3a) requires use of the NVIDIA open kernel modules` and fails `RmInitAdapter`, while the open kernel module at 580.159.03 drives the GPU fine — `nvidia-smi` shows "NVIDIA RTX PRO 4500 Blackwell Server Edition", CUDA 13.0.)

Because `nvidia-smi` fails, `nvidia-migmanager` fails, and the dependency chain through `configured.target` collapses — the host never reaches the point of starting kubelet.

## How to reproduce

Launch any **g7-family** instance with the `aws-k8s-*-nvidia` variant (reproduced on multiple g7 sizes/launches; the failure is deterministic).

## Suggested fix

Teach the driver-selection logic (ghostdog / driver preference) that device ID `10de:2c3a` (and other g7-family GPU IDs) must use the **open-gpu** driver, not tesla.

## Impact

Bottlerocket nvidia variants are currently unusable on the g7 instance family, with no settings/userdata escape hatch (admin/control host-containers can't intervene before `configured.target` fails). We had to move our GPU node group to AL2023 as a workaround.

Contributor guide

Open the contributing guide

Research direction

Start by locating ghostdog's driver-selection logic and the aws-k8s-*-nvidia variant configuration; trace how PCI ID 10de:2c3a is classified. Reproduce on a g7 instance and verify that open-gpu is selected, nvidia-smi and nvidia-migmanager succeed, configured.target completes, and kubelet starts.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, kubernetes, linux
Domain
devops, operating-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
56/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.