tensorflow / tensorflow/datasets
Multi-threaded compression?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 4.6k
- Forks
- 1.6k
- Avg merge
- 3h 54m
- Merged PRs (30d)
- 1
Description
What I need help with / What I was wondering
I need to build a large dataset of imagery that has > 3 channels (multi-spectral satellite imagery), so I'm relying on the tfds.features.Tensor feature connector. As writing data uncompressed is highly inefficient, I'm using tfds.features.Encoding.ZLIB for compression.
However, this compression step actually becomes the bottleneck in my dataset building process as it is single-threaded, causing my dataset build to take longer than a month.
What I've tried so far
Read up on the docs, also checked the tf.io namespace for any possible workarounds.
It would be nice if...
- Is there any way of speeding up the encoding/compression of the examples by using multiple cores?
- Are there plans to support a faster compression method than
ZLIBfor generic Tensor features?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading the tfds.features.Tensor and tfds.features.Encoding.ZLIB documentation, then trace how Tensor examples are encoded and check the tf.io namespace for compression options. The issue is done only when a supported approach for multi-core compression or a faster generic Tensor compression method is defined and validated for large dataset builds.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, tensorflow
- Domain
- data, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100