alibaba / alibaba/ROLL

oom

Open
#219 7 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.4k
Forks
312
Avg merge
1h 2m
Merged PRs (30d)
2

Description

Reward Feedback Learning (Reward FL)
中wan2.2训练需要什么配置呢,两张H100 80G ,batchsize为2会OOM

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the Reward Feedback Learning training entry point and reproducing wan2.2 training with two H100 80G GPUs and batch size 2. Trace the configuration and memory usage to identify why it runs out of memory; done means the reported setup has a documented working configuration or no longer OOMs.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.