MoonshotAI / MoonshotAI/FlashKDA
What is the motivation behind replacing the fp16 Neumann inverse with an 8x8 fp32 forward substitution
Nobody has claimed this yet.
- Dominant language
- Cuda
- Stars
- 1.3k
- Forks
- 122
- PR merge metrics
- No merged PRs in 30d
Description
Hi, I noticed that matrix inversion in kernel1 has changed from Neumann inverse to forward substitution. There appears to be little difference in data precision error between these two methods. May I ask what the motivation behind this change is?
Does forward substitution offer better performance?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating kernel1 and the change from the fp16 Neumann inverse to 8x8 fp32 forward substitution. Compare the two approaches' precision and performance in the available kernel benchmarks, then document the motivation and any measured trade-offs.
Written by the indexing model from the issue text.
Assessment
- Domain
- performance
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100