[Feature] how to open window attention in qwen-14B?
Open
backlog
- Dominant language
- Python
- Stars
- 8.1k
- Forks
- 748
- Avg merge
- 6d 2h
- Merged PRs (30d)
- 54
Description
### Motivation
I know I can change `/path/to/turbomind-style/triton_models/weights/config.ini` to open NTK-aware interpolation and LogN attention scaling.But where I can open window attention?
If I use NTK-aware interpolation and LogN attention scaling to extend the context from 2k to 16k, Is there any config needing to change for getting the speed as fast as the 2k context?
### Related resources
_No response_
### Additional context
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.