AbsaOSS / AbsaOSS/ABRiS

Preserve nested column resolution in AvroSchemaUtils.toAvroSchema

オープン
#375 コメント 0 件 リアクション 0 件 担当者 1 名 @kevinwallimann が担当を希望しています GitHub で見る
主要言語
Scala
スター
242
フォーク
84
平均マージ
10時間 12分
マージ済み PR(30日)
2

説明

## Summary

`AvroSchemaUtils.toAvroSchema` accepts a sequence of column names. Its current implementation builds a `StructType` from `dataFrame.schema(columnName)`. This lookup only resolves top-level schema fields.

Investigate and define the intended behavior for nested column paths such as `Seq("parent.child")`. If nested paths must be supported, preserve Catalyst column resolution when building the schema.

## Rationale

This behavior is outside the Spark 4 upgrade scope. A change can alter behavior for existing callers, so it requires separate investigation and regression coverage.

## Affected area

- `src/main/scala/za/co/absa/abris/avro/parsing/utils/AvroSchemaUtils.scala`
- Spark schema conversion tests for `AvroSchemaUtils`

## Required work

1. Identify existing callers that pass nested column paths to the `columnNames: Seq[String]` overload.
2. Define the compatibility contract for nested paths.
3. If nested paths are supported, construct the selected schema through analyzed DataFrame column expressions rather than direct top-level `StructType` field lookup.
4. Add regression tests for a nested path such as `Seq("parent.child")`.
5. Document any intentional behavior change or compatibility limitation.

## Acceptance criteria

- The nested-column behavior is explicitly defined.
- Tests cover the defined behavior on the supported Spark 4 version.
- The implementation does not unintentionally change top-level column behavior.

## Backlinks

- Source pull request: https://github.com/AbsaOSS/ABRiS/pull/370
- Source review comment: https://github.com/AbsaOSS/ABRiS/pull/370#discussion_r3898865601

Requested by: @kevinwallimann

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。