apache / apache/arrow-cookbook
[R] Add recipe for opening datasets with multiple file types
- Dominant language
- C++
- Stars
- 108
- Forks
- 49
- Avg merge
- 2h 14m
- Merged PRs (30d)
- 1
Description
``` r
library(arrow)
#> Some features are not enabled in this build of Arrow. Run `arrow_info()` for more information.
#>
#> Attaching package: 'arrow'
#> The following object is masked from 'package:utils':
#>
#> timestamp
library(dplyr)
#>
#> Attaching package: 'dplyr'
#> The following objects are masked from 'package:stats':
#>
#> filter, lag
#> The following objects are masked from 'package:base':
#>
#> intersect, setdiff, setequal, union
tf <- tempfile()
dir.create(tf)
arrow::write_csv_arrow(mtcars, file.path(tf, "mtcars.csv"))
arrow::write_parquet(mtcars, file.path(tf, "mtcars.parquet"))
schema <- schema(mpg = float64(), cyl = int64(), disp = float64(), hp = int64(),
drat = float64(), wt = float64(), qsec = float64(), vs = int64(),
am = int64(), gear = int64(), carb = int64())
csv_dataset <-
open_dataset(tf,
format = "csv",
factory_options = list(exclude_invalid_files = TRUE),
schema = schema, skip = 1)
parquet_dataset <-
open_dataset(tf,
format = "parquet",
factory_options = list(exclude_invalid_files = TRUE), schema = schema)
open_dataset(list(csv_dataset, parquet_dataset)) %>% collect()
#> # A tibble: 64 × 11
#> mpg cyl disp hp drat wt qsec vs am gear carb
#>
#> 1 21 6 160 110 3.9 2.62 16.5 0 1 4 4
#> 2 21 6 160 110 3.9 2.88 17.0 0 1 4 4
#> 3 22.8 4 108 93 3.85 2.32 18.6 1 1 4 1
#> 4 21.4 6 258 110 3.08 3.22 19.4 1 0 3 1
#> 5 18.7 8 360 175 3.15 3.44 17.0 0 0 3 2
#> 6 18.1 6 225 105 2.76 3.46 20.2 1 0 3 1
#> 7 14.3 8 360 245 3.21 3.57 15.8 0 0 3 4
#> 8 24.4 4 147. 62 3.69 3.19 20 1 0 4 2
#> 9 22.8 4 141. 95 3.92 3.15 22.9 1 0 4 2
#> 10 19.2 6 168. 123 3.92 3.44 18.3 1 0 4 4
#> # … with 54 more rows
```
Created on 2023-03-17 with [reprex v2.0.2](https://reprex.tidyverse.org)
Contributor guide
Research direction
Use the supplied R example as the starting point, and first locate the cookbook's existing recipe structure and the appropriate section for dataset-opening examples. Add a recipe showing CSV and Parquet datasets opened together with a shared schema, and verify that it documents the demonstrated collected result.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100