Optimization of input file split
未关闭
enhancement
- 主要语言
- Scala
- 星标
- 170
- 派生
- 96
- 平均合并
- 57 分钟
- 30 天内合并 PR
- 2
描述
## Background
I have fixed length of 200 bytes file with 100 multi segments present in copybook with 2k columns. It is taking almost 12 hours for just 1gb of file.
## Feature
Can you add feature for input file split in mb for all file formats, currently it is working for record format =VB, where input_split_size_mb
I tried to adjust block size in spark code, but cobrix taking default cluster block size.
## Proposed Solution [Optional]
Solution Ideas
1. Allow this input_split_size_mb for all file formats
2. How to take custom block size specified in spark configuration.
Ex:spark.conf.set(“dfs.blocksize”,”32m”)
贡献指南
这个仓库没有索引到贡献指南
评估
这个 Issue 还没有评估数据。