ollama / ollama/ollama-python

How to set token budget for Gemma4 with Ollama

Open
#651 7 comments 7 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
10.5k
Forks
1.2k
Avg merge
4m
Merged PRs (30d)
1

Description

The Gemma4 documentation by Google suggests setting the token budget to adjust the resolution quality of images for more accurate OCR tasks

https://ai.google.dev/gemma/docs/capabilities/vision/image#variable_resolution_token_budget

However, I do not see any configuration options to set it when running AsyncClient.chat(). Adding max_soft_tokens to options does not seem to do anything.

response = await AsyncClient(host=OLLAMA_HOST).chat(
            model=MODEL,
            messages=messages,
            format=self.result_model.model_json_schema(),
            options={
                "num_ctx": 32 * 1024,
                "temperature": 0.0,
                "max_soft_tokens": 560,    # does nothing
            }
        )

By default, Gemma4 models runs at 280 token budget which is not enough for my OCR task; I am testing with Gemma4-e4b.

To verify, I had a vehicle license plate image uploaded to the my ollama endpoint, the result had a missing letter at the end. Then I swapped to using pure transformers library to load an unquantized gemma4-e4b that also produces the same result. However, the transformers AutoProcessor library had an option to set max_soft_tokens where I set it to 560 and it produced the correct expected result.

// ollama default (q4_k_m) and unquantized gemma4
{
    "license_plate_number": "YRSGNB",    // expected YRSGNBY
    "license_plate_state": "California"
}

// unquantized gemma4 with max_soft_token=560
{
    "license_plate_number": "YRSGNBY",    // expected YRSGNBY
    "license_plate_state": "California"
}

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at AsyncClient.chat() and trace how the options dictionary is handled and forwarded, focusing on whether max_soft_tokens is supported. Reproduce the supplied Gemma4 OCR comparison with the default budget and a requested 560-token budget; the work is done when the setting is supported or its limitation is clearly established.

Written by the indexing model from the issue text.

Assessment

Tech stack
ollama, python
Domain
ai, api
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.