apache / apache/datafusion-python

CsvReadOptions.null_regex has no effect: matching values are read as literal strings, and fail the read in a numeric column

Aberta
#1,735 1 comentário 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
604
Forks
174
Merge médio
2d 22h
PRs com merge (30d)
5

Descrição

**Describe the bug**

`CsvReadOptions.null_regex` is accepted everywhere it is offered, but the CSV
reader never applies it. Values matching the regex are read as literal strings,
and when a matching value sits in a column typed as an integer the read fails
outright instead of producing NULL.

This is not a bug in this crate — `python/datafusion/options.py` stores the
value and `crates/core/src/options.rs` copies it into DataFusion's
`CsvReadOptions` correctly. The root cause is in DataFusion itself; see below.

**To Reproduce**

```python
from pathlib import Path
from datafusion import CsvReadOptions, SessionContext

p = Path("probe.csv")
p.write_text("id,name\n1,alice\n2,N/A\n3,carol\n")

ctx = SessionContext()
options = CsvReadOptions().with_has_header(True).with_null_regex(r"^(null|NULL|N/A)$")
ctx.read_csv(p, options=options).show()
```

```
+----+-------+
| id | name |
+----+-------+
| 1 | alice |
| 2 | N/A | <- expected NULL
| 3 | carol |
+----+-------+
```

Every entry point behaves the same way — the `CsvReadOptions(null_regex=...)`
constructor, `with_null_regex()`, `read_csv()`, `register_csv()`, and SQL:

```python
ctx.sql("""
CREATE EXTERNAL TABLE t (id INT, name VARCHAR)
STORED AS CSV LOCATION 'probe.csv'
OPTIONS ('format.has_header' 'true', 'format.null_regex' '^(null|NULL|N/A)$')
""")
ctx.sql("select * from t").show() # same output, N/A not nulled
```

The SQL form does not reject the option, and giving an explicit schema does not
change anything, so this is not schema inference choosing `Utf8`.

When the matching value is in a numeric column, the read fails rather than
returning the wrong value:

```python
p.write_text("id,value\n1,10\n2,N/A\n3,30\n")
ctx.sql("""CREATE EXTERNAL TABLE t2 (id INT, value BIGINT) STORED AS CSV
LOCATION 'probe.csv'
OPTIONS ('format.has_header' 'true', 'format.null_regex' '^(null|NULL|N/A)$')""")
ctx.sql("select * from t2").show()
```

```
DataFusion error: Arrow error: Parser error: Error while parsing value 'N/A' as
type 'Int64' for column 1 at line 2. Row data: '[2,N/A]'
```

This is the case that matters in practice: `N/A`, `NULL` and `-` placeholders
in otherwise numeric columns are the reason to reach for `null_regex` at all.

**Expected behavior**

A field matching `null_regex` is read as NULL, whatever the column's type.

**Additional context**

Why the existing coverage does not catch it: `test_read_csv_with_options` in
`python/tests/test_context.py` does set `null_regex="[pP]+aris"`, but the only
`Paris` in its fixture is on the `#Charlie;35;Paris` line, which `comment="#"`
removes before the reader sees it. The `None` in that test's expected output
comes from `truncated_rows=True` on the `Bob;25` row, not from `null_regex`. The
test pins that the option parses, which is what its comment says it is for — it
does not pin the behavior.

`docs/source/user-guide/io/csv.md` documents `with_null_regex` as "Treat these
as NULL", so the documented behavior and the actual behavior disagree.

**Root cause (upstream).** In `datafusion-datasource-csv` 55.0.0, the version
this repository pins, `null_regex` is applied during schema inference but never
to the reader that actually parses the rows:

- `src/file_format.rs:547` sets it on the `arrow::csv::reader::Format` used for
`infer_schema`.
- `src/source.rs:187` builds the `csv::ReaderBuilder` that reads the data, and
calls `with_delimiter`, `with_header`, `with_quote`, `with_truncated_rows`,
`with_terminator`, `with_escape` and `with_comment` — but not
`with_null_regex`. `arrow_csv::reader::ReaderBuilder::with_null_regex` exists
in arrow 59.2.0 and is simply never called.

`CsvSource` already holds the full `CsvOptions`, so `self.options.null_regex` is
in scope at that point — it looks like a few lines in `builder()`, mirroring the
`escape` and `comment` blocks just below it.

I could not find this reported in either this repository or apache/datafusion.
If you would rather track it upstream, I am happy to open it against
apache/datafusion and link it back here.

Found while working on #1728 / #1732, which is why the advanced-options example
there keeps `with_null_regex` set but puts its `N/A` in a string column, so the
example runs. Happy to take the fix if it turns out to be on this side.

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Direção de pesquisa

Comece em crates/core/src/options.rs e em src/source.rs:187 de datafusion-datasource-csv e, em seguida, compare o reader builder com o tratamento de null_regex em src/file_format.rs:547. Adicione cobertura a python/tests/test_context.py usando um valor correspondente em uma coluna numérica e execute os testes das opções de CSV. O trabalho estará concluído quando os campos correspondentes se tornarem NULL tanto para colunas string quanto para colunas numéricas, e o comportamento documentado em docs/source/user-guide/io/csv.md continuar correto.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
python, rust
Domínio
backend, data-engineering
Tipo de issue
Bug
Dificuldade
3/5
Tempo estimado
1-2 dias
Status de atividade
Ativa
Clareza
Claramente especificada
Facilidade para iniciantes
76/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.