tensorflow / tensorflow/datasets

BeamWriter hits "Record exceeds maximum record size" in Dataflow with autosharding

Open
#10,995 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
4.6k
Forks
1.6k
Avg merge
3h 54m
Merged PRs (30d)
1

Description

For TFDS 4.9.7 on Dataflow 2.60.0, I have a company-internal Dataflow job that fails. Given the input collection:

Elements added 332,090
Estimated size 1.74 TB

to train_write/GroupShards, where the output collection reports:

Elements added 2
Estimated size 1.8 GB

it then fails on the next element with

"E0123 207 recordwriter.cc:401] Record exceeds maximum record size (1096571470 > 1073741823)."

Workaround

By installing the TFDS prerelease after https://github.com/tensorflow/datasets/commit/37007453e963424b5cd5f81f19e0c4698dbed8e3 and controlling --num_shards=4096 (auto-detection choose 2048), the DatasetBuilder runs to completion on Dataflow. I'm curious why the auto-detection didn't choose more file shards however, as all training examples should be roughly the same size in this DatasetBuilder.

Suggested fix

Maybe this https://github.com/tensorflow/datasets/blob/9969ce542f4b0e1cbf0a085e8e0df11bccea5c17/tensorflow_datasets/core/utils/shard_utils.py#L79 is too little headroom for the training examples. The FeatureDict in this particular DatasetBuilder is large, and perhaps the key overhead is unusually large. Should that number be 0.8 instead? Or whether https://github.com/tensorflow/datasets/blob/9969ce542f4b0e1cbf0a085e8e0df11bccea5c17/tensorflow_datasets/core/utils/shard_utils.py#L54 should be larger when the FeatureDict contains many keys?

Side remark

Surprisingly Dataflow limits mention

Maximum size for a single element (except where stricter conditions apply, for example Streaming Engine). 2 GB

which doesn't seem to be true in practice since the GroupBy fails on ~1 GB as per the logged error.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with tensorflow_datasets/core/utils/shard_utils.py, especially the sizing logic at the referenced lines, and trace how it determines shards for train_write/GroupShards. Compare the automatic 2048-shard result with the successful explicit 4096-shard Dataflow run, including the recordwriter size error. Done means explaining the sizing mismatch and defining a supported correction or clarified limitation.

Written by the indexing model from the issue text.

Assessment

Tech stack
gcp, python, tensorflow
Domain
cloud, data, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.