apache / apache/arrow-cookbook

[R] Add recipe for opening datasets with multiple file types

Open
#301 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
108
Forks
49
Avg merge
2h 14m
Merged PRs (30d)
1

Description

``` r
library(arrow)
#> Some features are not enabled in this build of Arrow. Run `arrow_info()` for more information.
#>
#> Attaching package: 'arrow'
#> The following object is masked from 'package:utils':
#>
#> timestamp
library(dplyr)
#>
#> Attaching package: 'dplyr'
#> The following objects are masked from 'package:stats':
#>
#> filter, lag
#> The following objects are masked from 'package:base':
#>
#> intersect, setdiff, setequal, union

tf <- tempfile()
dir.create(tf)

arrow::write_csv_arrow(mtcars, file.path(tf, "mtcars.csv"))
arrow::write_parquet(mtcars, file.path(tf, "mtcars.parquet"))

schema <- schema(mpg = float64(), cyl = int64(), disp = float64(), hp = int64(),
drat = float64(), wt = float64(), qsec = float64(), vs = int64(),
am = int64(), gear = int64(), carb = int64())

csv_dataset <-
open_dataset(tf,
format = "csv",
factory_options = list(exclude_invalid_files = TRUE),
schema = schema, skip = 1)

parquet_dataset <-
open_dataset(tf,
format = "parquet",
factory_options = list(exclude_invalid_files = TRUE), schema = schema)

open_dataset(list(csv_dataset, parquet_dataset)) %>% collect()
#> # A tibble: 64 × 11
#> mpg cyl disp hp drat wt qsec vs am gear carb
#>
#> 1 21 6 160 110 3.9 2.62 16.5 0 1 4 4
#> 2 21 6 160 110 3.9 2.88 17.0 0 1 4 4
#> 3 22.8 4 108 93 3.85 2.32 18.6 1 1 4 1
#> 4 21.4 6 258 110 3.08 3.22 19.4 1 0 3 1
#> 5 18.7 8 360 175 3.15 3.44 17.0 0 0 3 2
#> 6 18.1 6 225 105 2.76 3.46 20.2 1 0 3 1
#> 7 14.3 8 360 245 3.21 3.57 15.8 0 0 3 4
#> 8 24.4 4 147. 62 3.69 3.19 20 1 0 4 2
#> 9 22.8 4 141. 95 3.92 3.15 22.9 1 0 4 2
#> 10 19.2 6 168. 123 3.92 3.44 18.3 1 0 4 4
#> # … with 54 more rows
```

Created on 2023-03-17 with [reprex v2.0.2](https://reprex.tidyverse.org)

Contributor guide

Open the contributing guide

Research direction

Use the supplied R example as the starting point, and first locate the cookbook's existing recipe structure and the appropriate section for dataset-opening examples. Add a recipe showing CSV and Parquet datasets opened together with a shared schema, and verify that it documents the demonstrated collected result.

Written by the indexing model from the issue text.

Assessment

Tech stack
r
Domain
documentation
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.