apache / apache/beam

BigQuery DIRECT_READ does not validate pipeline's project ID and instead tries to read from a null project

Open
#21,507 1 comment 0 reactions 0 assignees View on GitHub
bug gcp io java P2
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

When a pipeline is created without a GCP project ID and tries to read from BigQuery using Storage Read API, it runs into the following unhelpful error:
```

org.apache.beam.sdk.Pipeline$PipelineExecutionException: com.google.api.gax.rpc.PermissionDeniedException:
io.grpc.StatusRuntimeException: PERMISSION_DENIED: BigQuery Storage API has not been used in project
770406736630 before or it is disabled. Enable it by visiting https://console.developers.google.com/apis/api/bigquerystorage.googleapis.com/overview?project=770406736630
then retry. If you enabled this API recently, wait a few minutes for the action to propagate to our
systems and retry.
```

It looks like no validation for project ID is happening, and Beam tries to read without a project ID. Project 770406736630 mentioned in the error is a `null` project and throws off the user because it isn't their project.

 

Doing the same but using the EXPORT read method results in this more helpful error.
```

org.apache.beam.sdk.Pipeline$PipelineExecutionException: java.lang.NullPointerException: Required parameter
projectId must be specified.
```

 

Imported from Jira [BEAM-14119](https://issues.apache.org/jira/browse/BEAM-14119). Original Jira may contain additional context.
Reported by: ahmedabu.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.