4paradigm / 4paradigm/OpenMLDB
Zk too many connections when too many online loaddata jobs or too many work threads
- Ngôn ngữ chính
- C++
- Star
- 1.7k
- Fork
- 331
- Merge trung bình
- 12 ngày 12 giờ
- Pull request đã merge (30 ngày)
- 1
Mô tả
Too many **load online jobs** create too many connections with zk. And zk default setting `maxClientCnxns=60`.
(Online loaddata will generate too many clients, one zk client for each one.)
If `too many connections` occur, offine job will fail at `fail to init zk handler with hosts`, zk client error is `110(Connection timed out)`, we can't know it's too many connections of zk from offline jobs.
So if you can use diag tool or CLI to connect cluster, but the new offline job will get failed, it may be the zk connections error. To check it, `grep "Too many" `.
If you just run a onebox cluster(taskmanager is local) and just submit one job(no other running jobs), still get the `Too many` error, you might have got it wrong. You should check `show jobs`, no other running jobs in list, and `jps -v` to ensure no orphan offline jobs.
Diag tool can do the checks, it should support:
1. inspect offline --running to get all running jobs
2. execute jps and grep to get all running jobs in system
3. zk log parse(but maybe it’s not allowed to ssh the zk host)
Hướng dẫn đóng góp
Hướng nghiên cứu
The issue describes a ZooKeeper connection limit problem under high load from online jobs. Look at the ZooKeeper client initialization code and connection pooling. Check the error handling for 'Too many connections' in the logs. The diag tool or CLI needs enhancements to detect orphan jobs and parse ZK logs. Understanding the job submission and ZK client lifecycle in the codebase is required to propose a fix.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Lĩnh vực
- databases, devops, distributed-systems
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 35/100