tensorflow / tensorflow/text

GPU Memory Error

Open
#1,120 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
1.3k
Forks
379
Avg merge
3h 30m
Merged PRs (30d)
8

Description

Hello,

When attempting to run this notebook, I consistently received a "cpu to gpu memory limit exceeded" on Google colab with a GPU allocated (free account) - and my entire colab notebook session would crash. If I connected to an instance without a GPU (CPU only), then the notebook would run successfully but the training would take way too long. I ended up paying for a colab pro membership in order to get access to a "premium" GPU and allocated more RAM...then I was able to run the notebook with a GPU without failure.

We are facing a similar issue to the one described above when trying to follow this notebook in an internal environment leveraging Sagemaker. Can someone provide some helpful tips on getting around GPU memory errors in general with this notebook? We are using batch size of 32 with each item in the batch having roughly 115 tokens. Is lowering the batch size the only way to try to counter the memory error? We have not changed anything significant in this notebook other than the data we use it on, but because I faced a similar error just trying to run this notebook as-is in colab, I am wondering if something can be done to try to make this notebook more GPU efficient. Thank you!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue does not name a notebook file, test, or entry point. Locate the referenced training notebook and reproduce the GPU memory failure with the stated batch size and token length in Colab or SageMaker. Done means documenting reproducible memory limits and validated guidance or a confirmed efficiency improvement.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, jupyter-notebook
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.