Need an example of creating DDL for a Hive Parquet table with EEL
- 主要言語
- Scala
- スター
- 147
- フォーク
- 32
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
- CSVSource to HiveSink
```scala
val schema = AvroSchemaFns.fromAvroSchema(new Schema.Parser().parse(new File("user.avsc")))
CsvSource(path)
.withSchema(schema)
.to(HiveSink("mydatabase", "myTable"))
```
- Table field: **fname**, **lname**, **age**, **salary**
- 2 partition keys of **country** and **city**
```scala
object EelCreateTableExample extends App {
val crateTableCommand = HiveDDL.showDDL(
tableName = "mydatabase.mytable",
partitions = Seq(
PartitionColumn("country", StringType),
PartitionColumn("city", StringType)
),
fields = Seq(
Field("fname", StringType),
Field("lname", StringType),
Field("age", IntType.Signed),
Field("salary", DecimalType(38, 5))
),
tableType = TableType.EXTERNAL_TABLE,
location = Some("hdfs://nameservice1/blah/mytable_location"),
serde = "org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe",
inputFormat = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat",
outputFormat = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat",
props = Map.empty,
tableComment = Some("my lovely table"),
ifNotExists = true
)
println(crateTableCommand)
}
```
- Ouput:
```
CREATE EXTERNAL TABLE IF NOT EXISTS `mydatabase.mytable` (
`fname` string,
`lname` string,
`age` int,
`salary` decimal(38,5))
PARTITIONED BY (
`country` string,
`city` string)
ROW FORMAT SERDE
'org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe'
STORED AS INPUTFORMAT
'org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat'
OUTPUTFORMAT
'org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat'
LOCATION 'hdfs://nameservice1/blah/mytable_location'
```
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
調査の方向性
The example is already provided in the issue body. Look for the HiveDDL class in the codebase to understand how showDDL is implemented. Verify the output matches the expected CREATE EXTERNAL TABLE statement. Add this example to the project's documentation, likely in a examples or docs directory. Run the example to ensure it compiles and prints the correct DDL.
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- hadoop, scala
- 領域
- data-engineering, databases
- issue の種類
- ドキュメント
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 活発さ
- 停滞
- 明瞭さ
- 明確に書かれている
- 初心者へのやさしさ
- 75/100