dmlc / dmlc/xgboost

[jvm-package] how to make xgboost4j-spark use DenseVector to build model on version 0.90

Open
#5,068 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
28.8k
Forks
8.9k
Avg merge
1d 12h
Merged PRs (30d)
54

Description

First, thanks for your great job.
I have one question which need your help.
Our dataset has a lot of 0.0 as meaning value, so we do not want set 0.0 as missing value. What I can do is to make xgboost4j-spark use DenseVector. But I do not know how to do this.
Can anybody help? Some examples would be appriciated.
Thank you.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by checking the xgboost4j-spark 0.90 input-vector handling and any existing examples for DenseVector and missing values. Reproduce the dataset setup described in the issue; done means providing a documented, maintainer-confirmed example or clearly recording that the requested input is unsupported.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, scala, spark
Domain
data-engineering, machine-learning
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.