Implementation and the description in the paper
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 61
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
thank you for your hard work you’ve put into this project. I really appreciate the effort and dedication that has gone into maintaining Tora.
I have a question regarding the difference between the model architecture described in the paper and the current implementation.
In the paper, the usage of cross-attention mechanism for both the S-DiT-B block and T-DiT-B block is mentioned.
However, from reviewing the code, it doesn’t seem like that it is actually implemented.
If I’m missing something, please let me know.
Thanks again!
Contributor guide
No contributing guide indexed for this repository
Research direction
Compare the paper's descriptions of cross-attention in the S-DiT-B and T-DiT-B blocks with the corresponding Python implementation. Determine whether the mechanism is present and, if not, what change would reconcile the implementation with the paper; completion requires resolving or documenting the discrepancy.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100