NVIDIA-NeMo / NVIDIA-NeMo/Curator
Support Sentence Transformer Models for text (non embedding)
@praateekmahajan is already working on this.
Since Jan 20, 2026.
- Dominant language
- Python
- Stars
- 1.8k
- Forks
- 328
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 30
Description
Is your feature request related to a problem? Please describe.
Currently we have custom code to do mean pooling / last token. However that doesn't work for models like google/embeddinggemma-300m which do linear projection before doing mean pooling (and the linear projection is not available on AutoModel).
The code will be simpler if in our load_model we do SentenceTransformers.load_model(...)
I like how vLLM does it in their test cases which is is_sentence_transformer=??? to correctly load either from AutoModel or SentenceTransformer (https://github.com/vllm-project/vllm/blob/11857a00b0a59183286eab393df4e13b20efec3a/tests/conftest.py#L279)
Describe the solution you'd like
A clear and concise description of what you want to happen.
Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.
Additional context
Add any other context or screenshots about the feature request here.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.