apache / apache/hudi

facing connection issue in SPARK-HADOOP standalone mode while connecting on LAN between machines [SUPPORT]

Open
#9,968 4 comments 0 reactions 0 assignees View on GitHub
engine:spark priority:medium
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

I am using SPARK-Hadoop 3.3.2 bundle On IP 192.168.1.4x: I have started spark master (port 7077) and worker

On IP 192.168.1.22y: I have my webapp.py which: a. creates spark session (see below config):
```
spark = SparkSession.builder \
.appName("dataHudi") \
.master('spark://192.168.1.40:7077') \
.config("spark.submit.deployMode","client") \
.config('spark.driver.bindAddress', '192.168.1.40') \
.config('spark.driver.host', '192.168.1.40') \
.config('spark.driver.port', '33037') \
.config('spark.jars.packages', 'org.apache.hudi:hudi-spark3.3-bundle_2.12:0.13.1') \
.config('spark.serializer', 'org.apache.spark.serializer.KryoSerializer') \
.config('spark.sql.catalog.spark_catalog', 'org.apache.spark.sql.hudi.catalog.HoodieCatalog') \
.config('spark.sql.extensions', 'org.apache.spark.sql.hudi.HoodieSparkSessionExtension') \
.getOrCreate()
```
b. submits a job:
```
spark_df = spark.createDataFrame(ingested_df)

spark_df.write \
.format("org.apache.hudi") \
.options(**hudi_options) \
.mode("append") \
.save(basePath_ID +"/"+f"{unique_filename}")
```
when I check logs on 192.168.1.4x:8080 and 192.168.1.4x:8081 then I see that application is running and executors exiting and starting. but then when I check executor stderr and stdout logs then I see that the spark is trying to connect to a random port say 33037 and connection is failing on that port.

well, I tried to run both spark and application on the same machine and on ip 192.168.1.22y This worked.

But on LAN it fails. we tried with different configurations. changing the bind address, driver.host, driver.port etc..

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with webapp.py and the SparkSession configuration, then inspect the executor stderr and stdout logs for the failed connection to port 33037. Compare the working same-machine setup with the LAN setup and verify the driver and executor connection addresses and ports. Done means the application runs successfully across the two LAN machines without executor connection failures.

Written by the indexing model from the issue text.

Assessment

Tech stack
hadoop, python, spark
Domain
distributed-systems, networking
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.