deepseek-ai / deepseek-ai/FlashMLA

does flash_mla_with_kvcache work only in paged mode?

Open
#126 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
12.9k
Forks
1.2k
Avg merge
4h 20m
Merged PRs (30d)
2

Description

Is is possible to use it with kcache (and q, k) first dimension being the batch size ?
if not any plan to support the non paged mode ?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the flash_mla_with_kvcache entry point and its documented tensor layouts, focusing on the paged-mode requirement. Determine whether batch-first q, k, and kcache inputs are supported; done means establishing the supported behavior or defining the scope for non-paged support.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.