googleapis / googleapis/python-aiplatform

Cache works differently with vertexai.preview vs google.genai

Offen
#5,271 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
api: vertex-ai
Vorherrschende Sprache
Python
Sterne
905
Forks
465
Ø Merge
1 T. 13 Std.
Gemergte PRs (30 T.)
44

Beschreibung

I am trying to understand why caching works differently between these two libraries. vertexai.preview seems to be agnostic to the model used for creating cache however, genai runs into an error when cache is loaded with a different model than the one used to create it. Does the caching work differently under the hood and is the behaviour of one library more correct than the other?

In vertexai.preview, the model can be loaded with a cache using the following and it works fine.
```python
from vertexai.preview import caching
from vertexai.preview.generative_models import GenerativeModel

cached_content = caching.CachedContent.create(
model_name = "gemini-1.5-pro-002",
contents = ["The sky is blue." * 2000])
model = GenerativeModel(model_name="gemini-2.0-flash-001").from_cached_content(cached_content)
response = model.generate_content("What colour is sky?")
print(response.text)
>> Blue
```

However, a similar code using genai results in an error.
```python
from google.genai import Client, types

client = Client()
cached_content = client.caches.create(
model="gemini-1.5-pro-002",
config=types.CreateCachedContentConfig(
contents=["The sky is blue." * 2000]))
response = client.models.generate_content(
model="gemini-2.0-flash-001",
contents="What colour is the sky?",
config=types.GenerateContentConfig(cached_content=cached_content.name,))
print(response.text)

>> google.genai.errors.ClientError: 400 INVALID_ARGUMENT. {'error': {'code': 400, 'message': 'The model in the inference request gemini-2.0-flash-001 does not match the model in the cached content...'}}
```

Beitragsleitfaden

Beitragsleitfaden öffnen

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.