NVIDIA-NeMo / NVIDIA-NeMo/RL

Indifferent prompt/generation sequence length

Open
#796 0 comments 1 reaction 1 assignee Claimed by @ZhiyuLi-Nvidia View on GitHub
bug
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

**Describe the bug**

Nemo_rl uses the same length indifferently for:
* max_total_sequence_length
* max_input_seq_length
* max_model_len
- this might be inaccurate since “In vLLM, max_model_len refers to the maximum context length of the model, which includes both the input prompt tokens and the generated output tokens.”
* max_new_tokens in SamplingParam

**Steps/Code to reproduce bug**

Please list *minimal* steps or code snippet for us to be able to reproduce the bug.

A helpful guide on on how to craft a minimal bug report http://matthewrocklin.com/blog/work/2018/02/28/minimal-bug-reports.

**Expected behavior**

A clear and concise description of what you expected to happen.

**Environment overview (please complete the following information)**

- Environment location: [Bare-metal, Docker, Cloud(specify cloud provider - AWS, Azure, GCP, Collab)]
- Method of install: [pip install or from source]. Please specify exact commands you used to install.
- If method of install is [Docker], provide `docker pull` & `docker run` commands used

**Environment details**

If NVIDIA docker image is used you don't need to specify these.
Otherwise, please provide:
- OS version
- PyTorch version
- Python version

**Additional context**

Add any other context about the problem here.
Example: GPU model

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.