google-deepmind / google-deepmind/gemma

Feat/chat rolling cache

Open
#535 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
5.7k
Forks
1k
Avg merge
10h 33m
Merged PRs (30d)
2

Description

# Chat Sampler: Implement Rolling Cache for Infinite Conversations

## Summary
Implements a rolling cache strategy in `ChatSampler` to support infinite multi-turn conversations. When the context window (cache) capacity is exceeded, the sampler efficiently truncates the oldest history while preserving the most recent turns and the new user prompt, allowing the conversation to continue indefinitely.

## Problem Statement
Previously, `ChatSampler` utilized a fixed-size KV cache (defined by `cache_length`, default 4096).
* **Issue**: If a conversation's total length (history + new prompt) exceeded this limit, the underlying sampler would fail (e.g., `cache.is_full` check) or generation would halt immediately.
* **Experience**: Users would hit a hard stop after ~4096 tokens, forcing them to restart the conversation or manually manage state.
* **Missing Feature**: The codebase had an explicit `TODO(epot): Support and test rolling cache.` validation.

## Solution
Implemented a robust "prompt-based" rolling cache mechanism within `ChatSampler.chat`.

### Algorithm
1. **Usage Check**: Before sampling, we compare `last_state.used_cache_length + new_prompt_length` against `cache_length`.
2. **Rolling Trigger**: If the limit is exceeded:
* **Reconstruct History**: We effectively re-render the full conversation history from `self.turns` into a single string.
* **Truncate**: We combine the history with the new prompt, tokenize the entire sequence, and keep only the suffix that fits within `cache_length - 64` (a safety buffer).
* **Reset State**: We force `last_state = None`. This instructs the underlying `Sampler` to discard the full KV cache and perform a fresh prefill on the truncated prompt.
* **Transparency**: If `print_stream` is enabled, a message is printed to notify the user that rolling occurred.
3. **State Management**: `self.turns` remains intact, preserving the logical history of the conversation (for record-keeping), even though the model's effective context window is truncated.

### Impact
* **Infinite Conversations**: Users can now chat indefinitely. The model naturally "forgets" the oldest parts of the conversation once the window is full, similar to standard implementations in other LLM usage.
* **Robustness**: Prevents runtime errors or unexpected silent failures when the context window fills up.

## Verification
* **Logic Verification**: Verified the token counting and truncation logic ensures valid inputs to the `Sampler`.
* **Behavior**: Confirmed that `last_state` is reset correctly upon overflow, ensuring the `Sampler` prefills the new truncated context.

## #526 PR Raised for the issue, lmk if there any any iterations . Love contributing to Google-Deepmind

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.