apache / apache/beam

BigQueryIO reading stalls if no data is returned by query

Open
#18,316 0 comments 0 reactions 0 assignees View on GitHub
bug gcp io java P3
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

When running a BigQueryIO query that doesn't return any rows (e.g. nothing has changed in a delta job) the job seems to stall and nothing happens as no temp files are being written which I think might be what it is waiting for. Just adding one row to the source table will make the job run through successfully.

Code:
```

PCollection rows = p.apply("ReadFromBQ",
BigQueryIO.read()
.fromQuery("SELECT * FROM `myproject.dataset.table`")

.withoutResultFlattening().usingStandardSql());

```


Log:
```

Jun 02, 2017 9:00:36 AM org.apache.beam.sdk.io.gcp.bigquery.BigQueryServicesImpl$JobServiceImpl startJob
INFO:
Started BigQuery job: {jobId=beam_job_batch-query, projectId=my-project}.
bq show -j --format=prettyjson
--project_id=my-project beam_job_batch-query
Jun 02, 2017 9:03:11 AM org.apache.beam.sdk.io.gcp.bigquery.BigQuerySourceBase
executeExtract
INFO: Starting BigQuery extract job: beam_job_batch-extract
Jun 02, 2017 9:03:12 AM org.apache.beam.sdk.io.gcp.bigquery.BigQueryServicesImpl$JobServiceImpl
startJob
INFO: Started BigQuery job: {jobId=beam_job_batch-extract, projectId=my-project}.
bq show -j
--format=prettyjson --project_id=my-project beam_job_batch-extract
Jun 02, 2017 9:04:06 AM org.apache.beam.sdk.io.gcp.bigquery.BigQuerySourceBase
executeExtract
INFO: BigQuery extract job completed: beam_job_batch-extract
Jun 02, 2017 9:04:08 AM
org.apache.beam.sdk.io.FileBasedSource expandFilePattern
INFO: Matched 1 files for pattern gs://my-bucket/tmp/BigQueryExtractTemp/ff594d003c6440a1ad84b9e02858b5c6/000000000000.avro
Jun
02, 2017 9:04:09 AM org.apache.beam.sdk.io.FileBasedSource getEstimatedSizeBytes
INFO: Filepattern gs://my-bucket/tmp/BigQueryExtractTemp/ff594d003c6440a1ad84b9e02858b5c6/000000000000.avro
matched 1 files with total size 9750

```

Imported from Jira [BEAM-2404](https://issues.apache.org/jira/browse/BEAM-2404). Original Jira may contain additional context.
Reported by: jroxtheworld.

Contributor guide

Open the contributing guide

Research direction

Start at BigQueryIO and BigQuerySourceBase, especially the executeExtract and FileBasedSource expansion points shown in the log. Reproduce the provided Standard SQL query when it returns zero rows and inspect the temporary-file path. Done means the pipeline completes without requiring a source row, with regression coverage for the empty-result path.

Written by the indexing model from the issue text.

Assessment

Tech stack
google-cloud, java
Domain
data-engineering, databases
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.