google / google/gemma.cpp

Default flash attention retains redundant BF16 KV projection history

Abierto
#1,029 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
C++
Estrellas
7k
Forks
660
Merge medio
20 h 43 min
PR fusionados (30 d)
33

Descripción

## Summary

Default flash attention retains two BF16 representations of KV history.
The sequence-major projection buffer remains allocated and populated after
its keys and values have been transposed into the buffers attention reads.
This increases resident memory as context length grows.

## Affected path

- Backend: default `--attention_impl flash`.
- Confirmed with Gemma 3 270M, 1B, and 4B text inference.
- Measured baseline: `ffc1abc05abdf11d875d25d36c8859553ddf2641`.
- The retained-history path is also present on current `dev` (`b68def1`).
- T5 and DeepSeek use their legacy buffers differently.

## What happens

1. `ComputeQKV` writes BF16 projections into sequence-major `kv_cache`.
2. It applies normalization and positional encoding with BF16 rounding.
3. It transposes the results into `k_cache` and `v_cache`.
4. `FlashAttention` consumes those transposed buffers.
5. The original projection buffer still retains every layer and position.

References:
- [Projection and transpose path](https://github.com/google/gemma.cpp/blob/ffc1abc05abdf11d875d25d36c8859553ddf2641/gemma/attention.cc#L208-L322).
- [Full-sequence cache allocations](https://github.com/google/gemma.cpp/blob/ffc1abc05abdf11d875d25d36c8859553ddf2641/gemma/kv_cache.cc#L275-L297).

## Reproduction and observed cost

Run default-flash Gemma 3 270M with a 32,736-token text prompt,
sequence capacity 32,768, prefill batch 4,096, and 16 decode tokens.
Inspect the allocated cache extents and peak process RSS during inference.

The projection buffer occupies 578 MiB including row padding.
The transposed K/V buffers occupy another 576 MiB.
Measured peak process RSS is 1,778.59 MiB on this workload.
This is active retained data, not merely unused virtual address space.

Environment: Linux, Intel i5-12400F, six pinned threads,
Release AVX2/Haswell build, no oneDNN, and approximately 15.5 GiB RAM.

## Expected behavior

Projection intermediates should not retain a second complete KV history
after attention's persistent representation has been produced.
Required context, existing BF16 rounding, and model behavior must be preserved.

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Lee primero las líneas 208-322 de gemma/attention.cc y las líneas 275-297 de gemma/kv_cache.cc para seguir la proyección, la transposición y las asignaciones del caché. Reproduce la carga de trabajo de Gemma 3 con default-flash usando el prompt y la configuración de secuencia indicados; después, inspecciona las extensiones del caché y el RSS máximo. Se considera terminado cuando no se conserva el historial KV redundante, mientras la longitud del contexto, el redondeo BF16 y el comportamiento del modelo permanecen sin cambios.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
cpp
Área
machine-learning, performance
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Activo
Claridad
Bien especificado
Aptitud para principiantes
68/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.