apache / apache/iotdb

[Bug] Version 1.3.3 After 1 or 2 weeks of stable operation, the database error Failed to connect to config node.

未关闭
#14,702 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Java
星标
6.4k
派生
1.2k
平均合并
1 天 23 小时
30 天内合并 PR
115

描述

### Search before asking

- [x] I searched in the [issues](https://github.com/apache/iotdb/issues) and found nothing similar.

### Version

Operating system: CentOS 7
iotdb version 1.3.3 (Build: ad95a7e)
Docker version 27.4.0, build bde2b89
docker-compose version 1.24.1, build 4667896b

### Describe the bug and provide the minimal reproduce step

I use docker-compose script to start a stand-alone iotdb service, and when my service runs stably for 1-2 weeks, the system will report the following error, causing my Java SessionPool to fail to connect to the iotdb database,Writes and reads per second are about 100-300
`
2025-01-15 00:51:06,966 [pool-30-IoTDB-ClientRPC-Processor-141] WARN o.a.i.d.p.c.ConfigNodeClient:299 - The current node may have been down TEndPoint(ip:127.0.0.1, port:10710),try next node
2025-01-15 00:51:07,967 [pool-30-IoTDB-ClientRPC-Processor-141] INFO o.a.i.c.c.ThriftClient:91 - Broken pipe error happened in sending RPC, we need to clear all previous cached connection, error msg is java.net.ConnectException: Connection refused
2025-01-15 00:51:07,968 [pool-30-IoTDB-ClientRPC-Processor-141] WARN o.a.i.d.p.c.ConfigNodeClient:299 - The current node may have been down TEndPoint(ip:127.0.0.1, port:10710),try next node
2025-01-15 00:51:07,970 [pool-30-IoTDB-ClientRPC-Processor-141] ERROR o.a.i.d.a.ClusterAuthorityFetcher:410 - Failed to connect to config node.
2025-01-15 00:51:07,989 [pool-30-IoTDB-ClientRPC-Processor-141] INFO o.a.i.c.c.ThriftClient:91 - Broken pipe error happened in sending RPC, we need to clear all previous cached connection, error msg is java.net.ConnectException: Connection refused
`
The content of my docker-compose file is:
`
version: "3.3"

services:
iotdb-service:
image: apache/iotdb:latest
hostname: iotdb
container_name: iotdb
restart: always
network_mode: "host"
environment:
- cn_internal_address=127.0.0.1
- cn_internal_port=10710
- cn_consensus_port=10720
- cn_seed_config_node=127.0.0.1:10710
- dn_rpc_address=0.0.0.0
- dn_internal_address=127.0.0.1
- dn_rpc_port=6667
- dn_mpp_data_exchange_port=10740
- dn_schema_region_consensus_port=10750
- dn_data_region_consensus_port=10760
- dn_seed_config_node=127.0.0.1:10710
privileged: true
volumes:
- ./iotdb/data:/iotdb/data
- ./iotdb/logs:/iotdb/logs
`
When I restart the container, the service is restored

### What did you expect to see?

Is there a problem with my configuration? Or whether the docker container runs this problem

### What did you see instead?

![Image](https://github.com/user-attachments/assets/01713aef-a8fd-4dd6-b1b3-b75482ae2105)

### Anything else?

_No response_

### Are you willing to submit a PR?

- [ ] I'm willing to submit a PR!

贡献指南

打开贡献指南

调研方向

使用提供的 docker-compose 配置以及 ConfigNodeClient、ClusterAuthorityFetcher 和 ThriftClient 的日志条目作为切入点。调查在所述的 CentOS 7、Docker 和 IoTDB 1.3.3 设置下,独立服务为何会在 1–2 周后失去与 config-node 的连接。完成标准是确定一个已确认的原因,并记录一个可复现的修复方案或配置变更。

由索引模型根据 Issue 内容生成。

评估

技术栈
docker, docker-compose, java
领域
databases, devops
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。