facebookresearch / facebookresearch/LoRe
Possible data formatting issue affecting reproduction of PRISM results
- Dominant language
- Python
- Stars
- 18
- Forks
- 10
- PR merge metrics
- No merged PRs in 30d
Description
Hi, thanks for releasing this repository.
When trying to reproduce the reported results, I found that the accuracies are consistently close to 1.0. After checking the code, I noticed a potential data formatting issue that may significantly affect model training and evaluation.
In the data construction code, `chosen_utterance` and `rejected_utterance` are handled inconsistently:
```python
# the single resporse string
chosen_utterance = turn["chosen_utterance"][0]
entry['extra_info']['chosen_utterance'] = chosen_utterance
# all rejected response strings
rejected_utterance = turn["rejected_utterance"]
entry['extra_info']['rejected_utterance'] = rejected_utterance
```
Later, they are both wrapped as assistant messages:
```python
chosen = [{"content": entry["extra_info"]["chosen_utterance"], "role": "assistant"}]
rejected = [{"content": entry["extra_info"]["rejected_utterance"], "role": "assistant"}]
```
As a result, chosen_utterance is a string, while rejected_utterance is a list of strings.
When computing embeddings or feeding the data into the model, rejected_utterance is implicitly converted into a string representation of a list (e.g., `"['...']"`), which always contains the tokens `[` and `]`.
This introduces a trivial and spurious feature that the model can easily exploit to distinguish rejected samples from chosen ones, without learning any meaningful semantic preference. In practice, this may explain why the accuracy is consistently very high (often close to 1.0).
I hope my understanding is helpful. Please feel free to let me know if I have misunderstood anything or if there is additional context I might be missing. Thanks again for sharing your work.
Contributor guide
Assessment
This issue has not been assessed yet.