aws / aws/containers-roadmap

[EKS] [request]: Windows node - VolumeAttachment stuck

Open
#2,663 1 comment 1 reaction 0 assignees View on GitHub
EKS Proposed Windows
Dominant language
Shell
Stars
5.4k
Forks
334
PR merge metrics
No merged PRs in 30d

Description

### Community Note

* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or "me too" comments, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment

**Tell us about your request**
What do you want us to build?

In a mixed cluster, on a Windows node VolumeAttachment stuck for hours with detach error:

Detach Error:
Message: rpc error: code = DeadlineExceeded desc = context deadline exceeded
Time: 2024-07-29T06:49:42Z

This issue does not happen on Windows, but on Windows nodes after a certain run time we face the above issue. Usually it happens after like 2-3 days of uptime despite the daily workload (and the operation volume of attach - detach) is identical every day.

The volume state is "In-use" on AWS UI and this can be found in the logs:

> ebs-plugin I0729 07:26:42.701157 1 controller.go:463] "ControllerUnpublishVolume: detaching" volumeID="vol-05adad3e3cf7854ab" nodeID="i-0ec4abe6f62602eaa"
> ebs-plugin I0729 07:26:44.476134 1 cloud.go:1057] "Waiting for volume state" volumeID="vol-05adad3e3cf7854ab" actual="busy" desired="detached"
> ebs-plugin I0729 07:26:46.544284 1 cloud.go:1057] "Waiting for volume state" volumeID="vol-05adad3e3cf7854ab" actual="busy" desired="detached"
> ebs-plugin I0729 07:26:49.419443 1 cloud.go:1057] "Waiting for volume state" volumeID="vol-05adad3e3cf7854ab" actual="busy" desired="detached"
> ebs-plugin I0729 07:26:53.726079 1 cloud.go:1057] "Waiting for volume state" volumeID="vol-05adad3e3cf7854ab" actual="busy" desired="detached"
> ebs-plugin I0729 07:27:00.630341 1 cloud.go:1057] "Waiting for volume state" volumeID="vol-05adad3e3cf7854ab" actual="busy" desired="detached"
> ebs-plugin I0729 07:27:12.233899 1 cloud.go:1057] "Waiting for volume state" volumeID="vol-05adad3e3cf7854ab" actual="busy" desired="detached"
> ebs-plugin I0729 07:27:32.238823 1 cloud.go:1057] "Waiting for volume state" volumeID="vol-05adad3e3cf7854ab" actual="busy" desired="detached"
> ebs-plugin E0729 07:27:42.701238 1 driver.go:107] "GRPC error" err="rpc error: code = Internal desc = Could not detach volume "vol-05adad3e3cf7854ab" from node "i-0ec4abe6f62602eaa": context canceled"
> csi-attacher I0729 07:27:42.709052 1 csi_handler.go:234] Error processing "csi-37d0378ea70f8f66ea0687a4b86524b8e89182f80d29903bfcb9056f83393306": failed to detach: rpc error: code = DeadlineExceeded desc = context deadline exceeded

> ebs-plugin E0728 19:59:48.121284 7608 driver.go:107] "GRPC error" err="rpc error: code = Internal desc = Failed to find device path /dev/xvdaa. disk number for device path "/dev/xvdaa" volume id "vol-05adad3e3cf7854ab" not found"
> ebs-plugin E0728 20:01:52.276644 7608 driver.go:107] "GRPC error" err="rpc error: code = Internal desc = Failed to find device path /dev/xvdaa. disk number for device path "/dev/xvdaa" volume id "vol-05adad3e3cf7854ab" not found"
> ebs-plugin E0728 20:03:56.517699 7608 driver.go:107] "GRPC error" err="rpc error: code = Internal desc = Failed to find device path /dev/xvdaa. disk number for device path "/dev/xvdaa" volume id "vol-05adad3e3cf7854ab" not found"
> ebs-plugin E0728 20:06:00.523852 7608 driver.go:107] "GRPC error" err="rpc error: code = Internal desc = Failed to find device path /dev/xvdaa. disk number for device path "/dev/xvdaa" volume id "vol-05adad3e3cf7854ab" not found"

**Which service(s) is this request for?**
EKS

**Tell us about the problem you're trying to solve. What are you trying to do, and why is it hard?**
What outcome are you trying to achieve, ultimately, and why is it hard/impossible to do right now? What is the impact of not having this problem solved? The more details you can provide, the better we'll be able to understand and solve the problem.

**Are you currently working around this issue?**
Using m7a instances makes the workload stable. Initially, we were using t3 instances. AWS support suggested to use m7a instances but those did not work either with the same issue.

**Additional context**
Anything else we should know?
To reproduce the issue:

Configure and startup the driver with option enableWindows: true on mixed cluster on a windows node. Define the Storage class and PVC for your workload. We use Argo Workflows to schedule jobs which utilize the defined PVCs and attach volumes to pods.
The pods execute activities and upon completion it detaches. After a few day of correct execution the communication between the
controller which runs on a linux machine and the ebs-csi-node-windows pod which runs on a windows machine will experience communication problem, which later result in volumes gets stuck.

- Kubernetes version: v1.29.4-eks-036c24b
- Driver version: v1.30.0

---

Another [GH Issue was opened](https://github.com/kubernetes-sigs/aws-ebs-csi-driver/issues/2100) on the `aws-ebs-csi-driver` repo as well documenting this issue

**Attachments**
If you think you might have additional information that you'd like to include via an attachment, please do - we'll take a look. (Remember to remove any personally-identifiable information.)

- VolumeAttachment
![](https://private-user-images.githubusercontent.com/66022687/353072920-d9d0d0e9-541e-4e59-86f7-376ece81412b.jpg?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NTUwMTQzMjYsIm5iZiI6MTc1NTAxNDAyNiwicGF0aCI6Ii82NjAyMjY4Ny8zNTMwNzI5MjAtZDlkMGQwZTktNTQxZS00ZTU5LTg2ZjctMzc2ZWNlODE0MTJiLmpwZz9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNTA4MTIlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjUwODEyVDE1NTM0NlomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPWUzZjQyYTIwMjNlZWM3OTRhYjJhZWM5NjYxNGE4ODJmN2U3ZDU0MmEzMWVmNzM1YmIyYmYxYzViNTBmYjkyZGMmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0In0.kjvpYt_8ul6sgbjjtdobXDscZWIE_2DBOz4ioKOCQQc)

- PVC
![](https://private-user-images.githubusercontent.com/66022687/353075250-23ef907f-a70b-4e7b-9060-d46a4660ffb5.jpg?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NTUwMTQzMjYsIm5iZiI6MTc1NTAxNDAyNiwicGF0aCI6Ii82NjAyMjY4Ny8zNTMwNzUyNTAtMjNlZjkwN2YtYTcwYi00ZTdiLTkwNjAtZDQ2YTQ2NjBmZmI1LmpwZz9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNTA4MTIlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjUwODEyVDE1NTM0NlomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPWMzOWQ1YjgxNWIxZDA3MjUzMzNiZDA0NGEzYjM1MjA5OWJmYjI2MmIyNWIzZDNiMzZjMzZiODBmMmZlYjU5NGEmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0In0.M55ZH-fHqtLVHPoMB-pAGDYBRL_B5QxXBU8hii3ciKY)

This is how the volume looks exactly on UI:

![](https://private-user-images.githubusercontent.com/66022687/353350999-9b8a2b57-4b7d-4152-a884-4eb1966785d7.jpg?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NTUwMTQ1MjksIm5iZiI6MTc1NTAxNDIyOSwicGF0aCI6Ii82NjAyMjY4Ny8zNTMzNTA5OTktOWI4YTJiNTctNGI3ZC00MTUyLWE4ODQtNGViMTk2Njc4NWQ3LmpwZz9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNTA4MTIlMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjUwODEyVDE1NTcwOVomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPTVhNmVkMGEzOTAzYTQ1N2YwNjU2YzVlODBiODQzMjI5MTgzOTYxNzhlNGM2OTYzNzhlMTQ4Y2UxNjYwZjU2MzAmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0In0.m3QESbqT4Ly39O_WwcKFigBcGF4fmJZFk5cDzHMVgZE)

Contributor guide

Open the contributing guide

Research direction

Start with the mixed-cluster reproduction described in the issue: enableWindows, create the StorageClass and PVC, and observe the controller and ebs-csi-node-windows pod during attach and detach. Review the VolumeAttachment and logs around the DeadlineExceeded and missing device-path errors, then compare the related aws-ebs-csi-driver issue 2100. Done should include a confirmed cause and evidence that volumes no longer remain stuck after repeated workload cycles.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, kubernetes
Domain
cloud, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.