AnswerDotAI / AnswerDotAI/cold-compress

MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Open
#29 0 comments 0 reactions 2 assignees Claimed by @griff4692 View on GitHub
Dominant language
Python
Stars
153
Forks
16
PR merge metrics
No merged PRs in 30d

Description

Implement this [paper](https://arxiv.org/abs/2407.02490).

Similar to `class KVCacheFastGen` in that it involves a profiling step.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.