Need an example of creating DDL for a Hive Parquet table with EEL
- 主要语言
- Scala
- 星标
- 147
- 派生
- 32
- PR 合并指标
- 30 天内没有已合并 PR
描述
- CSVSource to HiveSink
```scala
val schema = AvroSchemaFns.fromAvroSchema(new Schema.Parser().parse(new File("user.avsc")))
CsvSource(path)
.withSchema(schema)
.to(HiveSink("mydatabase", "myTable"))
```
- Table field: **fname**, **lname**, **age**, **salary**
- 2 partition keys of **country** and **city**
```scala
object EelCreateTableExample extends App {
val crateTableCommand = HiveDDL.showDDL(
tableName = "mydatabase.mytable",
partitions = Seq(
PartitionColumn("country", StringType),
PartitionColumn("city", StringType)
),
fields = Seq(
Field("fname", StringType),
Field("lname", StringType),
Field("age", IntType.Signed),
Field("salary", DecimalType(38, 5))
),
tableType = TableType.EXTERNAL_TABLE,
location = Some("hdfs://nameservice1/blah/mytable_location"),
serde = "org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe",
inputFormat = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat",
outputFormat = "org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat",
props = Map.empty,
tableComment = Some("my lovely table"),
ifNotExists = true
)
println(crateTableCommand)
}
```
- Ouput:
```
CREATE EXTERNAL TABLE IF NOT EXISTS `mydatabase.mytable` (
`fname` string,
`lname` string,
`age` int,
`salary` decimal(38,5))
PARTITIONED BY (
`country` string,
`city` string)
ROW FORMAT SERDE
'org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe'
STORED AS INPUTFORMAT
'org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat'
OUTPUTFORMAT
'org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat'
LOCATION 'hdfs://nameservice1/blah/mytable_location'
```
贡献指南
这个仓库没有索引到贡献指南
调研方向
The example is already provided in the issue body. Look for the HiveDDL class in the codebase to understand how showDDL is implemented. Verify the output matches the expected CREATE EXTERNAL TABLE statement. Add this example to the project's documentation, likely in a examples or docs directory. Run the example to ensure it compiles and prints the correct DDL.
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- hadoop, scala
- 领域
- data-engineering, databases
- Issue 类型
- 文档
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 75/100