NVIDIA-NeMo / NVIDIA-NeMo/Export-Deploy

Move operations like chat templates to the server side for VLM deployment

Open
#423 1 comment 0 reactions 1 assignee View on GitHub

@meatybobby is already working on this.

Since Oct 1, 2025.

enhancement
Dominant language
Python
Stars
42
Forks
18
Avg merge
1d 2h
Merged PRs (30d)
8

Description

Is your feature request related to a problem? Please describe.
In multi-modal deployment for accuracy evaluation it was noticed that there are some inconsistencies between LLMs and VLMs deployment in pytriton. The biggest different is that for LLMs the chat template is applied on the server side (here), while for VLMs there's no such method and everything need to happen on the client side (here). Would it be possible to move this processing to the server side for VLMs too? Without this we don't have OpenAI-like api and it cannot be used the server for evaluation.

Solution:
The operations like chat template should be moved to the server side.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.