4paradigm / 4paradigm/OpenMLDB

spark load data(online&offline), the blank value in the data will become null

Abierto
#2,459 3 comentarios 0 reacciones 1 asignado Reclamado por @vagetablechicken Ver en GitHub
bug high-priority storage-engine
Lenguaje dominante
C++
Estrellas
1.7k
Forks
331
Merge medio
12 d 12 h
PR fusionados (30 d)
1

Descripción

**Bug Description**
When using load data to import data in CSV format, the empty string in the data will become null
![bbc42822e5496834964e8270e7cdcbb7](https://user-images.githubusercontent.com/18275570/189319963-ea3fe480-2661-46a9-a7a1-14cf81a2d055.png)

select into result:
9667b640bdc199847a35f6e0654691bf

**Expected Behavior**

**Relation Case**
test_select_into_load_data.yaml id:0-1. 0-2、18-2、62

**Steps to Reproduce**
```sql
create table auto_YuMSwEvO(
id int,
c1 string,
c2 smallint,
c3 int,
c4 bigint,
c5 float,
c6 double,
c7 timestamp,
c8 date,
c9 bool,
index(key=(c1),ts=c7))options(partitionnum=1,replicanum=1);
insert into auto_YuMSwEvO values
(3,'',3,22,32,1.3,2.3,1590738991000,'2020-05-03',true);
create table auto_FwREKcjN(
id int,
c1 string,
c2 smallint,
c3 int,
c4 bigint,
c5 float,
c6 double,
c7 timestamp,
c8 date,
c9 bool,
index(key=(c1),ts=c7))options(partitionnum=1,replicanum=1);
set @@SESSION.execute_mode = "online";
select * from auto_YuMSwEvO into outfile '/Users/zhaowei/code/4paradigm/OpenMLDB/auto_YuMSwEvO.csv' ;
LOAD DATA INFILE 'file:///Users/zhaowei/code/4paradigm/OpenMLDB/auto_YuMSwEvO.csv' into table auto_FwREKcjN options(mode='append');
```

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

The issue involves CSV data loading where empty strings become NULL. Look at the Spark load data path in OpenMLDB, likely in the Spark connector or data import modules. Check the test case referenced (test_select_into_load_data.yaml) to understand the expected behavior. Reproduce the bug using the provided SQL steps, then examine how CSV parsing handles empty values versus nulls.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
spark, sql
Área
data-engineering, databases
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
45/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.