ceph / ceph/ceph-iscsi

iscsi path goes down

Open
#190 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
68
Forks
58
PR merge metrics
No merged PRs in 30d

Description

Hello to all!

I have some problems with ceph iscsi and esxi. I have a 2 HP c7000 with 20 blade servers, on each of him installed esxi 6.7U3 and placed on sd-cards, on each I have a 2 HDD usable for ceph:)
My ceph cluster: 20 VMs on each blade with 2 disk for ceph. One admin VM with ceph-ansible.
Environment of VMs:

- CentOS 8
- Linux 5.6.12-1.el8.elrepo.x86_64 #1 SMP Fri May 8 19:48:47 EDT 2020 x86_64 GNU/Linux
- Branch - stable 5.0
- ceph version 15.2.1 (9fd2f65f91d9246fae2c841a6222d34d121680ee) octopus (stable)

**Bug Report**

All components installed successfully. I create rbd image and present it by iscsi(target created via gui) to 4 esxi hosts. All goes fine(from 1 to 5 days), but from one mystical moment all path going down by one. I find moment at one node and collect logs in attachment. Probably, I missed some major setting, I hope u can help me:)

After hit a problem I can not restart the rbd-target-gw service and I can not restart OS, the kernel is blocked. tcmu-runner log stops at the time of the crash. No other issues on ceph cluster I can't find, my proxmox cluster with rbd images from that ceph working fine.

Attachments:
[ansible_hosts.txt](https://github.com/ceph/ceph-iscsi/files/4660703/ansible_hosts.txt)
[messages.txt](https://github.com/ceph/ceph-iscsi/files/4660704/messages.txt)
[tcmu-runner.log](https://github.com/ceph/ceph-iscsi/files/4660705/tcmu-runner.log)
[rmp.txt](https://github.com/ceph/ceph-iscsi/files/4660706/rmp.txt)
[group_vars.txt](https://github.com/ceph/ceph-iscsi/files/4660707/group_vars.txt)

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing messages.txt and tcmu-runner.log around the reported path failure, then compare the environment in ansible_hosts.txt and group_vars.txt with the rmp.txt details. Establish what causes the rbd-target-gw service and kernel to become unresponsive; done would require a confirmed cause and a reproducible remediation, but the issue does not define a specific code or test target.

Written by the indexing model from the issue text.

Assessment

Tech stack
centos, linux
Domain
infrastructure, networking
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.