NVIDIA-Merlin / NVIDIA-Merlin/Transformers4Rec
[QST] Help with end-to-end example of yoochoose
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1.3k
- Forks
- 165
- Avg merge
- 1m
- Merged PRs (30d)
- 2
Description
❓ Questions & Help
Details
The end-to-end session based with yoochoose example notebook is getting stuck at the last training iteration on setting gpu size greater than 2. The code works fine on 1 and 2 gpus but on 3 and 4 GPU the iteration stuck at training output od 180th day...
### Tasks
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the end-to-end session-based Yoochoose example notebook and reproduce its training run on one and two GPUs, then on three and four GPUs. Compare the final training iteration, where the report says execution stalls around day 180, and document or resolve the multi-GPU behavior so the notebook completes consistently.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100