Distributed Training with TensorFlow Java
還沒有人認領這個 Issue。
- 主要語言
- Java
- 星號
- 928
- 分支
- 227
- PR 合併指標
- 30 天內沒有已合併 PR
描述
Please make sure that this is a feature request. As per our GitHub Policy, we only address code/doc bugs, performance issues, feature requests and build/installation issues on GitHub. tag:feature_template
System information
- TensorFlow version (you are using): 2.X
- Are you willing to contribute it (Yes/No): Yes, when able and available
Describe the feature and the current behavior/state.
Tensorflow on Python has tf.distribute.Strategy API to distribute training across multiple GPUs or multiple machines.
Will this change the current api? How?
Yes, it will add a new awesome feature
Who will benefit with this feature?
- Anyone that requires to speed up training a DL model
- Anyone that requires to train a DL model with big data
- Anyone who wants to create or add Java support for APIs that leverages tf.distribute.Strategy such as TensorflowOnSpark, Spark Tensorflow Distributor or Horovod
Any Other info.
https://www.tensorflow.org/guide/distributed_training
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
先閱讀 TensorFlow 的分散式訓練指南,並將其 tf.distribute.Strategy API 與此儲存庫中的 Java bindings 進行比較。確定哪些分散式訓練功能和 Java API 表面屬於範圍,然後在實作之前定義能展示多 GPU 或多機器訓練的測試或範例。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- java, tensorflow
- 領域
- distributed-systems, machine-learning
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 停滯
- 描述清晰度
- 需要釐清
- 新手友好度
- 20/100