[VL][1.5] Not enough spark off-heap execution memory on window
- Dominant language
- Scala
- Stars
- 1.6k
- Forks
- 657
- Avg merge
- 2d 14h
- Merged PRs (30d)
- 80
Description
### Backend
VL (Velox)
### Bug description
```
26/01/31 00:58:43 ERROR Executor: Exception in task 198.2 in stage 84.0 (TID 38940)
org.apache.gluten.exception.GlutenException: org.apache.gluten.exception.GlutenException: Error during calling Java code from native code: org.apache.gluten.exception.GlutenException: org.apache.gluten.exception.GlutenException: Exception: VeloxRuntimeError
Error Source: RUNTIME
Error Code: INVALID_STATE
Reason: Operator::addInput failed for [operator: Window, plan node ID: 2]: Error during calling Java code from native code: org.apache.gluten.memory.memtarget.ThrowOnOomMemoryTarget$OutOfMemoryException: Not enough spark off-heap execution memory. Acquired: 8.0 MiB, granted: 0.0 B. Try tweaking config option spark.memory.offHeap.size to get larger space to run this application (if spark.gluten.memory.dynamic.offHeap.sizing.enabled is not enabled).
Current config settings:
spark.gluten.memory.offHeap.size.in.bytes=2.0 GiB
spark.gluten.memory.task.offHeap.size.in.bytes=2.0 GiB
spark.gluten.memory.conservative.task.offHeap.size.in.bytes=1024.0 MiB
spark.memory.offHeap.enabled=true
spark.gluten.memory.dynamic.offHeap.sizing.enabled=false
Memory consumer stats:
Task.38940: Current used bytes: 2.0 GiB, peak bytes: N/A
\- Gluten.Tree.112: Current used bytes: 2.0 GiB, peak bytes: 2.0 GiB
\- Capacity[8.0 EiB].112: Current used bytes: 2.0 GiB, peak bytes: 2.0 GiB
+- NativePlanEvaluator-118.0: Current used bytes: 2040.0 MiB, peak bytes: 2040.0 MiB
| \- single: Current used bytes: 2040.0 MiB, peak bytes: 2040.0 MiB
| +- root: Current used bytes: 2024.4 MiB, peak bytes: 2033.0 MiB
| | +- task.Gluten_Stage_84_TID_38940_VTID_118: Current used bytes: 2024.4 MiB, peak bytes: 2033.0 MiB
| | | +- node.2: Current used bytes: 1958.6 MiB, peak bytes: 1960.0 MiB
| | | | \- op.2.0.0.Window: Current used bytes: 1958.6 MiB, peak bytes: 1958.6 MiB
| | | +- node.1: Current used bytes: 65.8 MiB, peak bytes: 1552.0 MiB
| | | | \- op.1.0.0.OrderBy: Current used bytes: 65.8 MiB, peak bytes: 1478.5 MiB
| | | +- node.4: Current used bytes: 128.0 B, peak bytes: 1024.0 KiB
| | | | \- op.4.0.0.FilterProject: Current used bytes: 128.0 B, peak bytes: 128.0 B
| | | +- node.0: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | | | \- op.0.0.0.ValueStream: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | | +- node.5: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | | | \- op.5.0.0.Aggregation: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | | \- node.6: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | | \- op.6.0.0.FilterProject: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | \- default_leaf: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- gluten::MemoryAllocator: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- default: Current used bytes: 0.0 B, peak bytes: 0.0 B
+- ArrowContextInstance.35: Current used bytes: 8.0 MiB, peak bytes: 8.0 MiB
+- IteratorMetrics.112: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- single: Current used bytes: 0.0 B, peak bytes: 0.0 B
| +- gluten::MemoryAllocator: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | \- default: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- root: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- default_leaf: Current used bytes: 0.0 B, peak bytes: 0.0 B
+- VeloxBatchResizer.112.OverAcquire.0: Current used bytes: 0.0 B, peak bytes: 2.4 MiB
+- NativePlanEvaluator-118.0.OverAcquire.0: Current used bytes: 0.0 B, peak bytes: 559.2 MiB
+- VeloxBatchResizer.112: Current used bytes: 0.0 B, peak bytes: 8.0 MiB
| \- single: Current used bytes: 0.0 B, peak bytes: 8.0 MiB
| +- root: Current used bytes: 0.0 B, peak bytes: 1024.0 KiB
| | \- default_leaf: Current used bytes: 0.0 B, peak bytes: 956.5 KiB
| \- gluten::MemoryAllocator: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- default: Current used bytes: 0.0 B, peak bytes: 0.0 B
+- ShuffleReader.30: Current used bytes: 0.0 B, peak bytes: 16.0 MiB
| \- single: Current used bytes: 0.0 B, peak bytes: 16.0 MiB
| +- root: Current used bytes: 0.0 B, peak bytes: 1024.0 KiB
| | \- default_leaf: Current used bytes: 0.0 B, peak bytes: 448.0 KiB
| \- gluten::MemoryAllocator: Current used bytes: 0.0 B, peak bytes: 451.6 KiB
| \- default: Current used bytes: 0.0 B, peak bytes: 451.6 KiB
+- IteratorMetrics.112.OverAcquire.0: Current used bytes: 0.0 B, peak bytes: 0.0 B
+- ShuffleReader.30.OverAcquire.0: Current used bytes: 0.0 B, peak bytes: 4.8 MiB
+- UniffleShuffleWriter.112: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- single: Current used bytes: 0.0 B, peak bytes: 0.0 B
| +- gluten::MemoryAllocator: Current used bytes: 0.0 B, peak bytes: 0.0 B
| | \- default: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- root: Current used bytes: 0.0 B, peak bytes: 0.0 B
| \- default_leaf: Current used bytes: 0.0 B, peak bytes: 0.0 B
\- UniffleShuffleWriter.112.OverAcquire.0: Current used bytes: 0.0 B, peak bytes: 0.0 B
at org.apache.gluten.memory.memtarget.ThrowOnOomMemoryTarget.borrow(ThrowOnOomMemoryTarget.java:104)
at org.apache.gluten.memory.listener.ManagedReservationListener.reserve(ManagedReservationListener.java:49)
at org.apache.gluten.vectorized.ColumnarBatchOutIterator.nativeHasNext(Native Method)
at org.apache.gluten.vectorized.ColumnarBatchOutIterator.hasNext0(ColumnarBatchOutIterator.java:57)
at org.apache.gluten.iterator.ClosableIterator.hasNext(ClosableIterator.java:39)
at scala.collection.convert.Wrappers$JIteratorWrapper.hasNext(Wrappers.scala:45)
at org.apache.gluten.iterator.IteratorsV1$InvocationFlowProtection.hasNext(IteratorsV1.scala:154)
at org.apache.gluten.iterator.IteratorsV1$IteratorCompleter.hasNext(IteratorsV1.scala:66)
at org.apache.gluten.iterator.IteratorsV1$PayloadCloser.hasNext(IteratorsV1.scala:38)
at org.apache.gluten.iterator.IteratorsV1$LifeTimeAccumulator.hasNext(IteratorsV1.scala:95)
at org.apache.gluten.iterator.IteratorsV1$ReadTimeAccumulator.hasNext(IteratorsV1.scala:122)
at scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:460)
at scala.collection.convert.Wrappers$IteratorWrapper.hasNext(Wrappers.scala:32)
at org.apache.gluten.vectorized.ColumnarBatchInIterator.hasNext(ColumnarBatchInIterator.java:36)
at org.apache.gluten.vectorized.ColumnarBatchOutIterator.nativeHasNext(Native Method)
at org.apache.gluten.vectorized.ColumnarBatchOutIterator.hasNext0(ColumnarBatchOutIterator.java:57)
at org.apache.gluten.iterator.ClosableIterator.hasNext(ClosableIterator.java:39)
at scala.collection.convert.Wrappers$JIteratorWrapper.hasNext(Wrappers.scala:45)
at org.apache.gluten.iterator.IteratorsV1$InvocationFlowProtection.hasNext(IteratorsV1.scala:154)
at org.apache.gluten.iterator.IteratorsV1$ReadTimeAccumulator.hasNext(IteratorsV1.scala:122)
at org.apache.gluten.iterator.IteratorsV1$PayloadCloser.hasNext(IteratorsV1.scala:38)
at org.apache.gluten.iterator.IteratorsV1$IteratorCompleter.hasNext(IteratorsV1.scala:66)
at scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:460)
at scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:460)
at org.apache.spark.shuffle.writer.VeloxUniffleColumnarShuffleWriter.writeImpl(VeloxUniffleColumnarShuffleWriter.java:145)
at org.apache.spark.shuffle.writer.RssShuffleWriter.write(RssShuffleWriter.java:344)
at org.apache.spark.shuffle.ShuffleWriteProcessor.write(ShuffleWriteProcessor.scala:59)
at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:104)
at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:54)
at org.apache.spark.TaskContext.runTaskWithListeners(TaskContext.scala:161)
at org.apache.spark.scheduler.Task.run(Task.scala:141)
at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$4(Executor.scala:620)
at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally(SparkErrorUtils.scala:64)
at org.apache.spark.util.SparkErrorUtils.tryWithSafeFinally$(SparkErrorUtils.scala:61)
at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:94)
at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:623)
at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
at java.lang.Thread.run(Thread.java:750)
```
### Gluten version
Gluten-1.5
### Spark version
None
### Spark configurations
_No response_
### System information
_No response_
### Relevant logs
```bash
```
Contributor guide
Research direction
Start with the stack-trace entry points ThrowOnOomMemoryTarget.borrow and ManagedReservationListener.reserve, then review the reported Window memory statistics and the listed off-heap configuration values. Reproduction details and an expected fix are not provided, so completion criteria must be established from the active discussion before implementation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- scala
- Domain
- backend, performance
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100