apache / apache/arrow

[R] read_csv_arrow fails when a string contains a backslash-escaped quote mark followed by a comma

Open
#33,405 7 comments 0 reactions 0 assignees View on GitHub
Component: R good-second-issue Priority: Critical Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 18h
Merged PRs (30d)
91

Description

`read_csv_arrow()` incorrectly parses CSV files when a string value contains a comma that appears after a backslash-escaped quote mark. Originally noted by Thomas Klebel

This is an example that throws the error:

```r

x <- tempfile()
readr::write_lines(
'
id,text
1,"some text on \\"BLAH
" and X, and Y also"
', x)

cat(system(paste('cat', x), intern = TRUE), sep = "\n")
#>
#> id,text
#> 1,"some text on \"BLAH\" and X, and Y also"
arrow::read_csv_arrow(x, escape_backslash = TRUE)
#> Error:
#> ! Invalid: CSV parse error: Expected 2 columns, got 3: 1,"some text on \"BLAH\" and X, and Y also"

#> Backtrace:
#> ▆
#> 1. └─arrow (local) ``(file = x, escape_backslash = TRUE, delim = ",")
#> 2. └─base::tryCatch(...) at r/R/csv.R:217:2
#> 3. └─base (local) tryCatchList(expr, classes, parentenv, handlers)
#> 4. └─base (local) tryCatchOne(expr, names, parentenv, handlers[[1L]])
#> 5. └─value[[3L]](cond)
#> 6. └─arrow:::augment_io_error_msg(e, call, schema = schema) at r/R/csv.R:222:6
#> 7. └─rlang::abort(msg, call = call) at r/R/util.R:251:2
```

Created on 2022-11-02 with [reprex v2.0.2]([https://reprex.tidyverse.org](https://reprex.tidyverse.org/))

This version includes four lines that might potentially error but do not:

```r

x <- tempfile()
readr::write_lines(
'
id,text
2,"some text on X and Y"
3,"some text on X, and Y"
4,"some text on \\"BLAH
"
5,"some text on X and Y, and \\"BLAH
" also"
', x)

cat(system(paste('cat', x), intern = TRUE), sep = "\n")
#>
#> id,text
#> 2,"some text on X and Y"
#> 3,"some text on X, and Y"
#> 4,"some text on \"BLAH\"
#> 5,"some text on X and Y, and \"BLAH\" also"
arrow::read_csv_arrow(x, escape_backslash = TRUE)
#> # A tibble: 4 × 2
#> id text
#>
#> 1 2 "some text on X and Y"
#> 2 3 "some text on X, and Y"
#> 3 4 "some text on \\BLAH\\\""
#> 4 5 "some text on X and Y, and \\BLAH\\\" also\""
```

Created on 2022-11-02 with [reprex v2.0.2]([https://reprex.tidyverse.org](https://reprex.tidyverse.org/))

I'm not sure if the problem is R specific. I've partially reproduced the error using reticulate and pyarrow as follows, but notice that this errors at a different point: the pyarrow version appears to fail with the comma preceding the backslash-escaped quote mark:

```r

x <- tempfile()
readr::write_lines(
'
id,text
1,"some text on X and Y"
2,"some text on X, and Y"
3,"some text on \\"BLAH
"
4,"some text on X and Y, and \\"BLAH
" also"
5,"some text on \\"BLAH
" and X, and Y also"
', x)

cat(system(paste('cat', x), intern = TRUE), sep = "\n")
#>
#> id,text
#> 1,"some text on X and Y"
#> 2,"some text on X, and Y"
#> 3,"some text on \"BLAH\"
#> 4,"some text on X and Y, and \"BLAH\" also"
#> 5,"some text on \"BLAH\" and X, and Y also"

csv <- reticulate::import("pyarrow.csv")
opt <- csv$ParseOptions(escape_char='
')
csv$read_csv(x, parse_options = opt)
#> Error in py_call_impl(callable, dots$args, dots$keywords): pyarrow.lib.ArrowInvalid: CSV parse error: Expected 2 columns, got 3: 3,"some text on \"BLAH\"
#> 4,"some text on X and Y, and \"BLAH\" also"
```

Created on 2022-11-02 with [reprex v2.0.2]([https://reprex.tidyverse.org](https://reprex.tidyverse.org/))

**Reporter**: [Danielle Navarro](https://issues.apache.org/jira/browse/ARROW-18219) / @djnavarro

**Note**: *This issue was originally created as [ARROW-18219](https://issues.apache.org/jira/browse/ARROW-18219). Please see the [migration documentation](https://github.com/apache/arrow/issues/14542) for further details.*

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the failing R example with arrow::read_csv_arrow(..., escape_backslash = TRUE), then inspect the wrapper and error path in r/R/csv.R, especially around lines 217-222. Compare the failing input with the four successful cases and verify that the reported CSV parses as two columns without regressing those cases.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python, r
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.