4paradigm / 4paradigm/OpenMLDB

spark load data(online&offline), the blank value in the data will become null

Offen
#2,459 3 Kommentare 0 Reaktionen 1 zugewiesene Person Beansprucht von @vagetablechicken Auf GitHub ansehen
bug high-priority storage-engine
Vorherrschende Sprache
C++
Sterne
1.7k
Forks
331
Ø Merge
12 T. 12 Std.
Gemergte PRs (30 T.)
1

Beschreibung

**Bug Description**
When using load data to import data in CSV format, the empty string in the data will become null
![bbc42822e5496834964e8270e7cdcbb7](https://user-images.githubusercontent.com/18275570/189319963-ea3fe480-2661-46a9-a7a1-14cf81a2d055.png)

select into result:
9667b640bdc199847a35f6e0654691bf

**Expected Behavior**

**Relation Case**
test_select_into_load_data.yaml id:0-1. 0-2、18-2、62

**Steps to Reproduce**
```sql
create table auto_YuMSwEvO(
id int,
c1 string,
c2 smallint,
c3 int,
c4 bigint,
c5 float,
c6 double,
c7 timestamp,
c8 date,
c9 bool,
index(key=(c1),ts=c7))options(partitionnum=1,replicanum=1);
insert into auto_YuMSwEvO values
(3,'',3,22,32,1.3,2.3,1590738991000,'2020-05-03',true);
create table auto_FwREKcjN(
id int,
c1 string,
c2 smallint,
c3 int,
c4 bigint,
c5 float,
c6 double,
c7 timestamp,
c8 date,
c9 bool,
index(key=(c1),ts=c7))options(partitionnum=1,replicanum=1);
set @@SESSION.execute_mode = "online";
select * from auto_YuMSwEvO into outfile '/Users/zhaowei/code/4paradigm/OpenMLDB/auto_YuMSwEvO.csv' ;
LOAD DATA INFILE 'file:///Users/zhaowei/code/4paradigm/OpenMLDB/auto_YuMSwEvO.csv' into table auto_FwREKcjN options(mode='append');
```

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

The issue involves CSV data loading where empty strings become NULL. Look at the Spark load data path in OpenMLDB, likely in the Spark connector or data import modules. Check the test case referenced (test_select_into_load_data.yaml) to understand the expected behavior. Reproduce the bug using the provided SQL steps, then examine how CSV parsing handles empty values versus nulls.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
spark, sql
Bereich
data-engineering, databases
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
45/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.