lllyasviel / lllyasviel/FramePack
FramePack training details
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 17.3k
- Forks
- 1.7k
- PR merge metrics
- No merged PRs in 30d
Description
Dear Authors,
I hope this message finds you well. I’m exploring FramePack training based on either Hunyuan or Wan and would appreciate your guidance on the following:
1、Base Model Architecture:
Is the base model designed for Text-to-Video (T2V) or Image-to-Video (I2V) tasks?
2、Parameter Fine-Tuning Strategy:
When freezing most parameters, is it sufficient to fine-tune only the PatchEmbedForCleanLatents modules (proj, proj_2x, proj_4x)?
Or would you recommend fine-tuning all parameters or selectively unfreezing the first few DiT blocks for better performance?
3、Training Convergence:
Approximately how many training steps are typically needed for convergence under standard settings (e.g., dataset size, batch size)?
Thank you for your time and insights! Looking forward to your response.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points. Start by reviewing the FramePack training context around Hunyuan, Wan, and PatchEmbedForCleanLatents, then clarify the expected task with maintainers. Done would require an agreed, documented answer covering the base architecture, fine-tuning strategy, and convergence steps.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100