MoonshotAI / MoonshotAI/FlashKDA

What is the motivation behind replacing the fp16 Neumann inverse with an 8x8 fp32 forward substitution

Open
#34 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Cuda
Stars
1.3k
Forks
122
PR merge metrics
No merged PRs in 30d

Description

Hi, I noticed that matrix inversion in kernel1 has changed from Neumann inverse to forward substitution. There appears to be little difference in data precision error between these two methods. May I ask what the motivation behind this change is?

Does forward substitution offer better performance?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating kernel1 and the change from the fp16 Neumann inverse to 8x8 fp32 forward substitution. Compare the two approaches' precision and performance in the available kernel benchmarks, then document the motivation and any measured trade-offs.

Written by the indexing model from the issue text.

Assessment

Domain
performance
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.