apache / apache/datafusion

Decimal division compatibility mode with spark

Open
#7,301 1 comment 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Rust
Stars
9.3k
Forks
2.4k
Avg merge
3d 7h
Merged PRs (30d)
344

Description

### Is your feature request related to a problem or challenge?

As described in detail by @liukun4515 and @tustvold and @viirya on https://github.com/apache/arrow-datafusion/pull/6832, DataFusion's decimal devision semantics.

@liukun4515 notes https://github.com/apache/arrow-datafusion/pull/6832#issuecomment-1680098056 that spark has the config to control the precision loss : https://github.com/apache/spark/blob/2be20e54a2222f6cdf64e8486d1910133b43665f/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/arithmetic.scala#L246

And @tustvold notes For people looking to emulate spark which only supports precision up to 38, casting to Decimal256 and then truncating down to Decimal128 will be equivalent, and is what a precision loss arithmetic kernel would do

### Describe the solution you'd like

If anyone needs spark compatible decimal division rules, I suggest:

1. Add a new config option
2. Apply the rewrite suggested by @tustvold (cast to Decimal256, divide, and then cast to Decimal128) as an [AnalyzerRule](https://docs.rs/datafusion/latest/datafusion/optimizer/analyzer/trait.AnalyzerRule.html#)

### Describe alternatives you've considered

See ticket -- we discussed at length changing the semantics of division in arrow-rs and concluded there was no one agreed upon ideal behavior

### Additional context

_No response_

Contributor guide

Open the contributing guide

Research direction

Start with the discussion in pull request 6832 and the AnalyzerRule documentation linked in the issue. Trace how decimal division is analyzed, then determine where a new configuration option and the Decimal256-to-Decimal128 rewrite belong. Done means Spark-compatible decimal division behavior is covered by tests, though no test file is named in the issue.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust, spark, sql
Domain
backend, databases
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.