googleapis / googleapis/python-aiplatform
Configurable cache path for List token (local tokenizer file path)
- 主要语言
- Python
- 星标
- 905
- 派生
- 465
- 平均合并
- 1 天 13 小时
- 30 天内合并 PR
- 44
描述
**Is your feature request related to a problem? Please describe.**
Currently Google aiplatform loads tokenizer by either downloading from github or reading from a tmpdir cache path which is not configurable at a library level. https://github.com/googleapis/python-aiplatform/blob/main/vertexai/tokenization/_tokenizer_loading.py#L136-L147
Can we make it configurable like how TikToken or NLK does it ? https://github.com/openai/tiktoken/blob/main/tiktoken/load.py#L34-L42 With a env variable like `VERTEX_TOKENIZER_CACHE_DIR` ?
**Describe the solution you'd like**
Our org does not allow network download of file on our deployment servers, so we need to uplaod the file to a fixed read only directory on the server. Being able to configure the path for that server would be useful.
**Describe alternatives you've considered**
I tried setting the TMPDIR env variable, but that works at python global level for all libraries and does not seem very configurable.
贡献指南
评估
这个 Issue 还没有评估数据。