4paradigm / 4paradigm/OpenMLDB

Support UnsafeRowOpt for groupby agg physical node

Aperta
#1,346 1 commento 0 reazioni 1 assegnatario Rivendicata da @tobegit3hub Vedi su GitHub
enhancement
Lingua principale
C++
Stelle
1.7k
Fork
331
Merge medio
12g 12h
PR unite (30g)
1

Descrizione

Now we can not enable UnsafeRowOpt for SQL with `group by` which may output incorrect result or crash because of C++ core.

Here is the simple case to reproduce.

```
test("Test unsafe groupby") {
val spark = getSparkSession
val sess = new OpenmldbSession(spark)

val data = Seq(
Row(1, "tom", 100, 1),
Row(2, "amy", 200, 2),
Row(3, "tom", 300, 3),
Row(4, "amy", 400, 4),
Row(5, "tom", 500, 5),
Row(6, "amy", 600, 6),
Row(7, "tom", 700, 7),
Row(8, "amy", 800, 8),
Row(9, "tom", 900, 9),
Row(10, "amy", 1000, 10))
val schema = StructType(List(
StructField("id", IntegerType),
StructField("user", StringType),
StructField("trans_amount", IntegerType),
StructField("trans_time", IntegerType)))
val df = spark.createDataFrame(spark.sparkContext.makeRDD(data), schema)

sess.registerTable("t1", df)
df.createOrReplaceTempView("t1")

val sqlText = "SELECT max(id) AS max_id, sum(trans_amount) AS sum_amount FROM t1 GROUP BY user"
// core for unaligned memory
//val sqlText = "SELECT user, max(id) AS max_id, sum(trans_amount) AS sum_amount FROM t1 GROUP BY user"

val outputDf = sess.sql(sqlText)
outputDf.show()
}
```

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

The issue involves the groupby aggregation physical node in OpenMLDB's execution engine, specifically the UnsafeRowOpt optimization. Start by examining the code around groupby physical node implementation and UnsafeRowOpt handling. Look for memory alignment issues or incorrect result logic when UnsafeRowOpt is enabled. The provided test case reproduces the crash; run it to see the failure, then trace through the relevant C++ code to understand the root cause.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
spark, sql
Ambito
backend, databases, machine-learning
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.