aws / aws/amazon-sagemaker-examples
Training job for `rl_distributed_robotschool` failed
Open
- Dominant language
- Jupyter Notebook
- Stars
- 11k
- Forks
- 7k
- Avg merge
- 8h 29m
- Merged PRs (30d)
- 8
Description
Problem notebook:
```
https://github.com/aws/amazon-sagemaker-examples/blob/master/reinforcement_learning/rl_roboschool_ray/rl_roboschool_ray_distributed.ipynb
```
Training jobs for both primary cluster and secondary cluster failed
Contributor guide
Research direction
Start with reinforcement_learning/rl_roboschool_ray/rl_roboschool_ray_distributed.ipynb, the linked problem notebook, and inspect the training setup for the primary and secondary clusters. Reproduce the failure if possible and determine what prevents both training jobs from completing successfully.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, jupyter-notebook
- Domain
- cloud, distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100