apache / apache/rocketmq-flink

Run on k8s, it's throw RemotingTimeoutException

Open
#100 1 comment 1 reaction 0 assignees View on GitHub
Dominant language
Java
Stars
174
Forks
104
PR merge metrics
No merged PRs in 30d

Description

my Flink Job run on k8s with flink-kubernetes-operator
[https://github.com/apache/flink-kubernetes-operator]([flink-kubernetes-operator])

Tasks always report errors from time to time, but it doesn’t always appear
![image](https://github.com/apache/rocketmq-flink/assets/47728686/231c500c-fd64-4679-9b0f-ff4a025dd766)

but when I execute it on the yarn cluster, this error will not appear again.

rocketmq cluster status is normal.

does anyone have the same problem?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the intermittent RemotingTimeoutException with the Flink Kubernetes Operator and compare the behavior with the same job on a YARN cluster. Inspect the task error details alongside RocketMQ cluster status and determine a reliable reproduction or root cause; done means documenting the cause and a verified resolution.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, kubernetes
Domain
cloud, distributed-systems, stream-processing
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.