facebookresearch / facebookresearch/fairseq2

SamplingSequenceGenerator not respecting `max_gen_len`

Open
#1,078 2 comments 0 reactions 0 assignees View on GitHub
bug generation
Dominant language
Python
Stars
1.1k
Forks
144
Avg merge
4d 1h
Merged PRs (30d)
1

Description

**Describe the bug:**
When specifying `max_gen_len`, the `SamplingSequenceGenerator` can potentially generate more than `max_gen_len` for all batched sequences whose prompt length is shorter than the longest prompt length in batch.

In [this line](https://github.com/facebookresearch/fairseq2/blob/main/src/fairseq2/generation/_sampling/_generator.py#L702) the sequences are forced to be EOS with relation to `self._max_seq_len`. However, `self._max_seq_len` is calculated as the length of the longest prompt + `max_gen_len` (unless `max_seq_len` is specified). Therefore, any sequence in a batch whose prompt length ($N$) is shorter than the max prompt length in a batch ($M$) can potentially generate $M - N$ + `max_gen_len` (unless they naturally generate EOS beforehand).

**Describe how to reproduce:**
```
tokens, padding_mask = pad_seqs([torch.Tensor([4]).long(), torch.Tensor([1,5,2,6,8,4]).long()])
sampler = TopKSampler(k=1)
generator = SamplingSequenceGenerator(model, sampler, max_gen_len=10)
res = generator(prompt_seqs=tokens, prompt_padding_mask=padding_mask)

In [48]: res.hypotheses[0][0].seq.shape
Out[48]: torch.Size([15]) # max_gen_len + 5

In [49]: res.hypotheses[1][0].seq.shape
Out[49]: torch.Size([10]) # max_gen_len
```

**Describe the expected behavior:**
```
In [48]: res.hypotheses[0][0].seq.shape
Out[48]: torch.Size([10]) # max_gen_len

In [49]: res.hypotheses[1][0].seq.shape
Out[49]: torch.Size([10]) # max_gen_len
```
**Environment:**
fairseq2==0.4.0

Contributor guide

Open the contributing guide

Research direction

Start at src/fairseq2/generation/_sampling/_generator.py line 702 and inspect how self._max_seq_len is applied to batched prompts. Reproduce the issue with the provided pad_seqs, TopKSampler, and max_gen_len=10 example. Done means both output sequences have length 10 rather than the shorter-prompt sequence reaching length 15.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.