apache / apache/druid

Druid index_parallel ingestion failing in merge phase

Open
#19,035 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
14.1k
Forks
3.8k
Avg merge
2d 58m
Merged PRs (30d)
233

Description

I am trying to ingest data in druid datasource using index_parallel but the ingestion is getting failed again and again in merge phase

Here is the task spec :

{
"type": "index_parallel",
"spec": {
"dataSchema": {
"dataSource": "X",
"timestampSpec": {
"column": "timestamp",
"format": "auto"
},
"dimensionsSpec": {
"dimensions": [
"segment",
"region",
"deviceIdType",
"scope"
]
},
"metricsSpec": [
{
"fieldName": "X",
"type": "thetaSketch",
"name": "X",
"isInputThetaSketch": true,
"size": 65536
}
],
"granularitySpec": {
"type": "uniform",
"segmentGranularity": "day",
"queryGranularity": "day",
"intervals": ["PLACEHOLDER"],
"rollup": false
}
},
"ioConfig": {
"type": "index_parallel",
"inputSource": {
"type": "combining",
"delegates": [
{
"type": "druid",
"dataSource": "X",
"interval": "PLACEHOLDER",
"filter": {
"type": "not",
"field": {
"type": "in",
"dimension": "X",
"values": ["X"]
}
}
},
}
}
]
},
"inputFormat": {
"type": "parquet",
"binaryAsString": false
},
"appendToExisting": false
},
"tuningConfig": {
"type": "index_parallel",
"forceGuaranteedRollup": true,
"maxRowsInMemory": 25000,
"maxBytesInMemory": 500000000,
"partitionsSpec": {
"type": "hashed",
"targetRowsPerSegment": 6000
},
"maxNumConcurrentSubTasks": 10,
"maxRetry": 3,
"taskStatusCheckPeriodMs": 1000,
"chatHandlerTimeout": "PT10S",
"chatHandlerNumRetries": 5,
"pushTimeout": 0,
"ignoreInvalidRows": true,
"buildV9Directly": true
}
}
}

I am using a single middle manager with r6gd.8xlarge instance type with 12 worker slots, with 4GB heap for middle manager and 20GB ( 16GB heap + 4GB direct ) for peon task

Contributor guide

Open the contributing guide

Research direction

No source file, test, stack trace, or failing entry point is named. Start by obtaining the merge-phase task logs and a reproducible task specification, then trace the index_parallel merge path; done means identifying a concrete Druid defect and adding a regression test.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.