NVIDIA-NeMo / NVIDIA-NeMo/Curator

Support Sentence Transformer Models for text (non embedding)

Open
#1,265 3 comments 0 reactions 2 assignees View on GitHub

@praateekmahajan is already working on this.

Since Jan 20, 2026.

enhancement
Dominant language
Python
Stars
1.8k
Forks
328
Avg merge
4d 5h
Merged PRs (30d)
30

Description

Is your feature request related to a problem? Please describe.
Currently we have custom code to do mean pooling / last token. However that doesn't work for models like google/embeddinggemma-300m which do linear projection before doing mean pooling (and the linear projection is not available on AutoModel).
The code will be simpler if in our load_model we do SentenceTransformers.load_model(...)

I like how vLLM does it in their test cases which is is_sentence_transformer=??? to correctly load either from AutoModel or SentenceTransformer (https://github.com/vllm-project/vllm/blob/11857a00b0a59183286eab393df4e13b20efec3a/tests/conftest.py#L279)

Describe the solution you'd like
A clear and concise description of what you want to happen.

Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.

Additional context
Add any other context or screenshots about the feature request here.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.