4paradigm / 4paradigm/OpenMLDB

Support UnsafeRowOpt for groupby agg physical node

Ouverte
#1,346 1 commentaire 0 réactions 1 personne assignée Réclamée par @tobegit3hub Voir sur GitHub
enhancement
Langage dominant
C++
Étoiles
1.7k
Forks
331
Merge moyen
12 j 12 h
PR mergées (30 j)
1

Description

Now we can not enable UnsafeRowOpt for SQL with `group by` which may output incorrect result or crash because of C++ core.

Here is the simple case to reproduce.

```
test("Test unsafe groupby") {
val spark = getSparkSession
val sess = new OpenmldbSession(spark)

val data = Seq(
Row(1, "tom", 100, 1),
Row(2, "amy", 200, 2),
Row(3, "tom", 300, 3),
Row(4, "amy", 400, 4),
Row(5, "tom", 500, 5),
Row(6, "amy", 600, 6),
Row(7, "tom", 700, 7),
Row(8, "amy", 800, 8),
Row(9, "tom", 900, 9),
Row(10, "amy", 1000, 10))
val schema = StructType(List(
StructField("id", IntegerType),
StructField("user", StringType),
StructField("trans_amount", IntegerType),
StructField("trans_time", IntegerType)))
val df = spark.createDataFrame(spark.sparkContext.makeRDD(data), schema)

sess.registerTable("t1", df)
df.createOrReplaceTempView("t1")

val sqlText = "SELECT max(id) AS max_id, sum(trans_amount) AS sum_amount FROM t1 GROUP BY user"
// core for unaligned memory
//val sqlText = "SELECT user, max(id) AS max_id, sum(trans_amount) AS sum_amount FROM t1 GROUP BY user"

val outputDf = sess.sql(sqlText)
outputDf.show()
}
```

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

The issue involves the groupby aggregation physical node in OpenMLDB's execution engine, specifically the UnsafeRowOpt optimization. Start by examining the code around groupby physical node implementation and UnsafeRowOpt handling. Look for memory alignment issues or incorrect result logic when UnsafeRowOpt is enabled. The provided test case reproduces the crash; run it to see the failure, then trace through the relevant C++ code to understand the root cause.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
spark, sql
Domaine
backend, databases, machine-learning
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.