4paradigm / 4paradigm/OpenMLDB

Support UnsafeRowOpt for groupby agg physical node

Open
#1,346 1 comment 0 reactions 1 assignee Claimed by @tobegit3hub View on GitHub
enhancement
Dominant language
C++
Stars
1.7k
Forks
331
Avg merge
12d 12h
Merged PRs (30d)
1

Description

Now we can not enable UnsafeRowOpt for SQL with `group by` which may output incorrect result or crash because of C++ core.

Here is the simple case to reproduce.

```
test("Test unsafe groupby") {
val spark = getSparkSession
val sess = new OpenmldbSession(spark)

val data = Seq(
Row(1, "tom", 100, 1),
Row(2, "amy", 200, 2),
Row(3, "tom", 300, 3),
Row(4, "amy", 400, 4),
Row(5, "tom", 500, 5),
Row(6, "amy", 600, 6),
Row(7, "tom", 700, 7),
Row(8, "amy", 800, 8),
Row(9, "tom", 900, 9),
Row(10, "amy", 1000, 10))
val schema = StructType(List(
StructField("id", IntegerType),
StructField("user", StringType),
StructField("trans_amount", IntegerType),
StructField("trans_time", IntegerType)))
val df = spark.createDataFrame(spark.sparkContext.makeRDD(data), schema)

sess.registerTable("t1", df)
df.createOrReplaceTempView("t1")

val sqlText = "SELECT max(id) AS max_id, sum(trans_amount) AS sum_amount FROM t1 GROUP BY user"
// core for unaligned memory
//val sqlText = "SELECT user, max(id) AS max_id, sum(trans_amount) AS sum_amount FROM t1 GROUP BY user"

val outputDf = sess.sql(sqlText)
outputDf.show()
}
```

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.