[doc] Enhance documentation for distributed training.
- Dominant language
- C++
- Stars
- 28.8k
- Forks
- 8.9k
- Avg merge
- 1d 12h
- Merged PRs (30d)
- 54
Description
Currently, we have some high-level documents for various interfaces, such as Dask and Spark. But the documentation around actual XGBoost components is scant. It would be great if we could have a more coherent document specifically for distributed training, to aid third-party integration and customization.
This is particularly important, as existing distributed frameworks are no longer sufficient for training large models when we want to max out hardware capacity.
related:
- https://github.com/kubeflow/trainer/pull/3118
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files or tests; start by reviewing the existing Dask and Spark interface documents and the XGBoost distributed-training components. Done means a coherent XGBoost-focused distributed-training document that explains third-party integration and customization.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp, spark
- Domain
- distributed-systems, documentation, machine-learning
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100