apache / apache/datafusion

Open Variant Type for semi-structured data

Open
#10,987 4 comments 8 reactions 1 assignee Claimed by @wjones127 View on GitHub
Dominant language
Rust
Stars
9.3k
Forks
2.4k
Avg merge
3d 7h
Merged PRs (30d)
344

Description

I've been starting to experiment with implementing the Open Variant Type [^1] in Rust / DataFusion. There is a specification and Java library for this, and Spark will release this type in 4.0. There are also plans to integrate this into table formats such as Delta Lake [^3] and Iceberg [^4]. This would be a high-performance data type for semi-structured data, designed for better OLAP performance than JSON or BSON (discussed in #7845). I've discussed a little bit in the Arrow repo about it's potential as an Arrow extension type [^2].

I'm working on creating an extension similar to [datafusion-functions-json](https://github.com/datafusion-contrib/datafusion-functions-json). If we could create a new repo `datafusion-functions-variant`, I'd be happy to develop that in the open.

[^1]: https://github.com/apache/spark/tree/master/common/variant
[^2]: https://github.com/apache/arrow/issues/42069
[^3]: https://www.databricks.com/blog/introducing-open-variant-data-type-delta-lake-and-apache-spark
[^4]: https://lists.apache.org/thread/xnyo1k66dxh0ffpg7j9f04xgos0kwc34

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.