the model is underrated , are you trying to fly under the radar ?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1.1k
- Forks
- 93
- Avg merge
- 1m
- Merged PRs (30d)
- 1
Description
MOVA training 1024 GPUs for 42 days (~43 000 GPU‑days). and you don't bother to optimize and get your model out there on reddit and forums .
i'm reading the paper and this model should give more than even the demos you have don't reflect it's true power
"Computational Resources. All three phases run on 1024 GPUs (128 nodes, 8 GPUs per node). For 360p
training (Phases 1–2), we use CP=8, yielding effective batch size 128. For 720p fine-tuning (Phase 3), increased
sequence length requires CP=16, reducing effective batch size to 64. The complete training spans 42 days,
totaling approximately 43,000 GPU-days. "
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the paper, the current demos, and the reported training setup of 1024 GPUs over 42 days. The issue does not identify files, entry points, a specific optimization, or acceptance criteria, so the intended change and definition of done need clarification before implementation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100