InternLM / InternLM/lmdeploy

[Feature] how to open window attention in qwen-14B?

Open
#638 4 comments 0 reactions 0 assignees View on GitHub
backlog
Dominant language
Python
Stars
8.1k
Forks
748
Avg merge
6d 2h
Merged PRs (30d)
54

Description

### Motivation

I know I can change `/path/to/turbomind-style/triton_models/weights/config.ini` to open NTK-aware interpolation and LogN attention scaling.But where I can open window attention?

If I use NTK-aware interpolation and LogN attention scaling to extend the context from 2k to 16k, Is there any config needing to change for getting the speed as fast as the 2k context?

### Related resources

_No response_

### Additional context

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.