BigQueryIO reading stalls if no data is returned by query
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 4.7k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 196
Description
When running a BigQueryIO query that doesn't return any rows (e.g. nothing has changed in a delta job) the job seems to stall and nothing happens as no temp files are being written which I think might be what it is waiting for. Just adding one row to the source table will make the job run through successfully.
Code:
```
PCollection rows = p.apply("ReadFromBQ",
BigQueryIO.read()
.fromQuery("SELECT * FROM `myproject.dataset.table`")
.withoutResultFlattening().usingStandardSql());
```
Log:
```
Jun 02, 2017 9:00:36 AM org.apache.beam.sdk.io.gcp.bigquery.BigQueryServicesImpl$JobServiceImpl startJob
INFO:
Started BigQuery job: {jobId=beam_job_batch-query, projectId=my-project}.
bq show -j --format=prettyjson
--project_id=my-project beam_job_batch-query
Jun 02, 2017 9:03:11 AM org.apache.beam.sdk.io.gcp.bigquery.BigQuerySourceBase
executeExtract
INFO: Starting BigQuery extract job: beam_job_batch-extract
Jun 02, 2017 9:03:12 AM org.apache.beam.sdk.io.gcp.bigquery.BigQueryServicesImpl$JobServiceImpl
startJob
INFO: Started BigQuery job: {jobId=beam_job_batch-extract, projectId=my-project}.
bq show -j
--format=prettyjson --project_id=my-project beam_job_batch-extract
Jun 02, 2017 9:04:06 AM org.apache.beam.sdk.io.gcp.bigquery.BigQuerySourceBase
executeExtract
INFO: BigQuery extract job completed: beam_job_batch-extract
Jun 02, 2017 9:04:08 AM
org.apache.beam.sdk.io.FileBasedSource expandFilePattern
INFO: Matched 1 files for pattern gs://my-bucket/tmp/BigQueryExtractTemp/ff594d003c6440a1ad84b9e02858b5c6/000000000000.avro
Jun
02, 2017 9:04:09 AM org.apache.beam.sdk.io.FileBasedSource getEstimatedSizeBytes
INFO: Filepattern gs://my-bucket/tmp/BigQueryExtractTemp/ff594d003c6440a1ad84b9e02858b5c6/000000000000.avro
matched 1 files with total size 9750
```
Imported from Jira [BEAM-2404](https://issues.apache.org/jira/browse/BEAM-2404). Original Jira may contain additional context.
Reported by: jroxtheworld.
Contributor guide
Research direction
Start at BigQueryIO and BigQuerySourceBase, especially the executeExtract and FileBasedSource expansion points shown in the log. Reproduce the provided Standard SQL query when it returns zero rows and inspect the temporary-file path. Done means the pipeline completes without requiring a source row, with regression coverage for the empty-result path.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- google-cloud, java
- Domain
- data-engineering, databases
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 42/100