embulk / embulk/embulk-output-bigquery

Split file each 4GB for BigQuery Quota Policy

Open
#6 3 comments 1 reaction 0 assignees View on GitHub
new feature
Dominant language
Ruby
Stars
126
Forks
59
PR merge metrics
No merged PRs in 30d

Description

BigQuery has following [Quota Policy](https://cloud.google.com/bigquery/quota-policy).

So, It's better to split output file each 4GB.

| File Type | Compressed | Uncompressed |
| --- | --- | --- |
| CSV | 4 GB | With new-lines in strings: 4 GB
Without new-lines in strings: 5 TB |
| JSON | 4 GB | 5TB |
## Problems
- Have to split newline(CRLF/LF/CR) at EOL, not only filesize.
- Split before output beforehand is better way than split output file, Because Embulk run multiple tasks with multiple CPU cores.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.