aws / aws/amazon-sagemaker-examples

[Debugger] tensorflow_nlp_sentiment_analysis training job runs much longer than expected

Open
#1,957 7 comments 0 reactions 0 assignees View on GitHub
status: awaiting response
Dominant language
Jupyter Notebook
Stars
11k
Forks
7k
Avg merge
8h 29m
Merged PRs (30d)
8

Description

https://github.com/aws/amazon-sagemaker-examples/tree/master/sagemaker-debugger/tensorflow_nlp_sentiment_analysis

Since the notebook published with the public SDK version, the training job runs for over an hour with epoch=25.
It's expected to run for ~500 seconds.

Contributor guide

Open the contributing guide

Research direction

Open the sagemaker-debugger/tensorflow_nlp_sentiment_analysis notebook and run the training job with epoch=25 using the public SDK version. Compare its runtime with the expected roughly 500 seconds and inspect the notebook's training configuration for the source of the longer run. Done means the job completes near the expected duration.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, jupyter-notebook, tensorflow
Domain
machine-learning, performance
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.