facebook / facebook/zstd

Optimal pointcloud compression

Open
#4,449 3 comments 1 reaction 1 assignee Claimed by @Cyan4973 View on GitHub
question
Dominant language
C
Stars
27.9k
Forks
2.6k
Avg merge
1d 3h
Merged PRs (30d)
8

Description

**Is your feature request related to a problem? Please describe.**
I am looking for the best general-purpose solution (prioritizing a balance of compression ratio and decompression speed) for compressing pointcloud data, specifically long arrays of `int32` values where each chunk of 3 corresponds to the x, y, z coordinates for a given point. The industry standard is [laz](https://laszip.org/) but that suffers from slow decompression speed (more than 10x worse than zstd) due to the use of slow arithmetic coders. There is also [draco](https://github.com/google/draco)(which implements https://arxiv.org/abs/cs/9909018) which offers competitive compression ratios but is generally slower than zstd for both compression and decompression.

I have found pretty good results by:
1. Sorting the points in [morton order](https://en.wikipedia.org/wiki/Z-order_curve)
2. Shuffling/transposing the bytes from a `N*12` array to a `12*N` array (so that all the first bytes of each point are clustered together and so on) using the blosc2 "shuffle" filter
3. Storing the byte delta using the blosc2 [bytedelta](https://www.blosc.org/posts/bytedelta-enhance-compression-toolset/) filter
4. Compressing with zstd -16

This quite often gets a better compression ratio than laz, with an order of magnitude faster decompression.

For some datasets this outperforms (in terms of compression ratio) draco by more than 10%, but for other datasets draco outperforms the above scheme by more than 10%.

I could compress a subset of points with both my scheme and draco and use the best algorithm for the whole set of points, but I feel like there might be a better solution out there.

**Describe the solution you'd like**
Ideally I would like to configure and/or modify zstd to take more advantage of the inherent structure in the data and reliably get the best compression ratio on most/all datasets.

I would like any advice on where to look or what experiments I could try.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.