iscsi path goes down
- Dominant language
- Python
- Stars
- 68
- Forks
- 58
- PR merge metrics
- No merged PRs in 30d
Description
Hello to all!
I have some problems with ceph iscsi and esxi. I have a 2 HP c7000 with 20 blade servers, on each of him installed esxi 6.7U3 and placed on sd-cards, on each I have a 2 HDD usable for ceph:)
My ceph cluster: 20 VMs on each blade with 2 disk for ceph. One admin VM with ceph-ansible.
Environment of VMs:
- CentOS 8
- Linux 5.6.12-1.el8.elrepo.x86_64 #1 SMP Fri May 8 19:48:47 EDT 2020 x86_64 GNU/Linux
- Branch - stable 5.0
- ceph version 15.2.1 (9fd2f65f91d9246fae2c841a6222d34d121680ee) octopus (stable)
**Bug Report**
All components installed successfully. I create rbd image and present it by iscsi(target created via gui) to 4 esxi hosts. All goes fine(from 1 to 5 days), but from one mystical moment all path going down by one. I find moment at one node and collect logs in attachment. Probably, I missed some major setting, I hope u can help me:)
After hit a problem I can not restart the rbd-target-gw service and I can not restart OS, the kernel is blocked. tcmu-runner log stops at the time of the crash. No other issues on ceph cluster I can't find, my proxmox cluster with rbd images from that ceph working fine.
Attachments:
[ansible_hosts.txt](https://github.com/ceph/ceph-iscsi/files/4660703/ansible_hosts.txt)
[messages.txt](https://github.com/ceph/ceph-iscsi/files/4660704/messages.txt)
[tcmu-runner.log](https://github.com/ceph/ceph-iscsi/files/4660705/tcmu-runner.log)
[rmp.txt](https://github.com/ceph/ceph-iscsi/files/4660706/rmp.txt)
[group_vars.txt](https://github.com/ceph/ceph-iscsi/files/4660707/group_vars.txt)
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing messages.txt and tcmu-runner.log around the reported path failure, then compare the environment in ansible_hosts.txt and group_vars.txt with the rmp.txt details. Establish what causes the rbd-target-gw service and kernel to become unresponsive; done would require a confirmed cause and a reproducible remediation, but the issue does not define a specific code or test target.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- centos, linux
- Domain
- infrastructure, networking
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100