4paradigm / 4paradigm/OpenMLDB

The discrete feature values in the gcformat sample data generated by the OpenMLDB SQL feature extraction script are inconsistent with those calculated by the PICO script

オープン
#3,923 コメント 1 件 リアクション 0 件 担当者 1 名 @wyl4pd に割り当て済み GitHub で見る
bug
主要言語
C++
スター
1.7k
フォーク
331
平均マージ
12日 12時間
マージ済み PR(30日)
1

説明

**Bug Description**
Service Version: 0.9.0
The discrete feature values in the gcformat sample data generated by the OpenMLDB SQL feature extraction script are inconsistent with those calculated by the PICO script.

**Expected Behavior**
Current incorrect format: label| slot:sign:origin-value
Correct format: label index| slot:sign:origin-value

**Relation Case**
OpenMLDB SQL Feature Extraction Example:
```
0| 1:0:1 2:4599670039981440374 3:6365000770384461703 4:0:93.200000
1| 1:0:2 2:5613161932270271752 3:-1384602352766124944 4:0:93.075000
0| 1:0:3 2:4599670039981440374 3:-6239076729344379818 4:0:92.893000
```
PICO Feature Extraction Example:
```
0 0| 2:-8773247204422130117:1 3:4042412524814531440 4:6048373541161169225 5:4681710344575317709:0x1.74ccccccccccdp6
1 1| 2:-8773247204422130117:2 3:6142047291687075953 4:1461111459061395210 5:4681710344575317709:0x1.744cccccccccdp6
0 2| 2:-8773247204422130117:3 3:4042412524814531440 4:3353218529862650678 5:4681710344575317709:0x1.73926e978d4fep6
```

**Steps to Reproduce**
1. data schema:
```
id[Int],age[Int],job[String],cons_price_idx[Double],y[Int]
```
2. PICO Feature Extraction Script:
```
target_y = binary_label(y)
f_id = continuous(id)
f_age = discrete(age)
f_job = discrete(job)
f_cons_price_idx = continuous(cons_price_idx)
```
4. OpenMLDB SQL Feature Extraction Script:
```
select gcformat(
binary_label(bool(y)),
continuous(id),
discrete(age),
discrete(job),
continuous(cons_price_idx)
) as instance from main_table
```

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

The issue is about the gcformat output in OpenMLDB SQL. Examine the SQL feature extraction script and the underlying C++ implementation of the gcformat function. Compare the output format with the PICO script's expected format. The fix likely involves modifying the string formatting logic for discrete features to include the label index. Look for tests related to gcformat or feature extraction to validate the change.

索引モデルが issue の本文から書いたものです。

評価

技術スタック
sql
領域
databases, machine-learning
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
停滞
明瞭さ
明確に書かれている
初心者へのやさしさ
40/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。