apache / apache/hudi

Bootstrap reads still hand Hive a path outside the table root when data columns are projected

Open
#19,643 0 comments 0 reactions 1 assignee Claimed by @rangareddy View on GitHub
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

Follow-up to #15676 / PR #19510, which fixed only the no-projection shape. Raised from [voonhous's review](https://github.com/apache/hudi/pull/19510#discussion_r3774251680).

## What is still broken

`HoodieParquetInputFormat.createBootstrappingRecordReader` opens one of two files for a `BootstrapBaseFileSplit`. The split's own path is the skeleton, inside the table root; `getBootstrapFileSplit()` is the external source file, which is not. Two branches still pass the external path to Hive:

- **external-only** — when the query projects data columns but no Hudi meta columns, the reader is opened on the external split;
- **stitch** — when both are projected, `rightSplit` (the external split) is passed to `getRecordReaderInternal`.

Hive's `VectorizedParquetRecordReader.initPartitionValues` resolves partition values by looking the split path up in `pathToPartitionInfo`, which only contains the table's own partition directories, so on a **partitioned** bootstrap table both branches fail with:

```
java.io.IOException: cannot find dir = s3://.../parquet-source-tables///part-....parquet
in pathToPartitionInfo: [s3://...//, ...]
```

So `SELECT FROM bootstrap_tbl` still throws after #19510, which only reordered the two single-file branches so that a no-column projection (`SELECT COUNT(*)`) resolves to the skeleton.

## Not Hive-3 specific

This is gated by a config, not a Hive version. `hive-exec` 2.3.10 ships the same `VectorizedParquetRecordReader.initPartitionValues`; what differs is the default of `hive.vectorized.execution.enabled` — `false` in `hive-common` 2.3.10, `true` in 3.1.3 (verified with `javap` on `HiveConf$ConfVars`: `iconst_0` vs `iconst_1`). A Hive 2 deployment with vectorization enabled hits it too.

## Workaround

```sql
set hive.vectorized.execution.enabled=false;
```

## Why it is not a simple reorder

Unlike the no-projection case, these two branches genuinely need the external file's data, so they cannot just be pointed at the skeleton. Options worth weighing:

1. present the skeleton path to Hive while reading external data underneath (the split path and the read path stop being the same thing);
2. register the external location in `pathToPartitionInfo`;
3. route bootstrap reads through the file-group reader — `shouldUseFilegroupReader` currently excludes `BootstrapBaseFileSplit` (`HoodieInputFormatUtils.java:571`), which is why this legacy path is still live in 1.x.

## Testing note

`TestBootstrap` and `TestOrcBootstrap` already drive these branches end to end ("RO Input Format Read - Project only non-hoodie column", `TestBootstrap.java:633`), but both classes are `@Disabled("HUDI-7353")` since #10551, so nothing exercises them today.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.