facebookresearch / facebookresearch/LoRe
Chosen and rejected responses are encoded differently in PRISM/prepare.py — is this intended?
- Dominant language
- Python
- Stars
- 18
- Forks
- 10
- PR merge metrics
- No merged PRs in 30d
Description
While working with the PRISM pipeline I noticed that chosen and rejected responses appear to be formatted differently before embedding, and standardizing this seems to reduce LoRe performance to that of the base reward model.
## What I'm seeing
In `PRISM/prepare.py` (lines ~340–345), the chosen response is formatted as a string, while the rejected response is formatted as a list. The list is later cast to a string at the embedding step, which appears to introduce stray characters (brackets and quotes) into the rejected text but not the chosen text.
Concretely:
```python
# the single resporse string
chosen_utterance = turn["chosen_utterance"][0]
=entry['extra_info']['chosen_utterance'] = chosen_utterance
# all rejected response strings
rejected_utterance = turn["rejected_utterance"]
entry['extra_info']['rejected_utterance'] = rejected_utterance
```
And at the embedding call in `PRISM/generate-prism-embeddings.py`:
```python
rejected = [{"content": entry["extra_info"]["rejected_utterance"], "role": "assistant"}]
chosen_conv = prompt + chosen
rejected_conv = prompt + rejected
```
So the embedded rejected string ends up with additional characters, like brackets, that I assume make it easier to detect. If the two sides of each pair are encoded differently, the difference is present in every pair in the same direction, which means it could in principle be picked up as signal rather than the underlying preference.
I made a minimal change — embedding the first element of the list rather than the string cast of the whole list — so that both sides are formatted identically:
```python
# list (their "rejected" text was previously just the stringified "[]")
rejected_list = entry["extra_info"]["rejected_utterance"]
rejected = [{"content": rejected_list[0], "role": "assistant"}] if rejected_list else None
chosen_conv = prompt + chosen
rejected_conv = prompt + rejected if rejected is not None else None
```
With that change, accuracy on seen users / unseen prompts is roughly flat across basis sizes:
| K | As released | With matched encoding |
| --- | --- | --- |
| 0 | 72% | ~56% |
| 1 | 76% | ~56% |
| 10 | 93% | ~56% |
Since `K=0` is the reference model, the learned basis doesn't appear to improve over it once the encoding is matched.
## Reproduction details
Full diff and results: https://github.com/RachelFreedman/LoRe-investigation/tree/encoding
---
Is this formatting difference intentional — for example, is the list form handling something about multi-turn structure that I've missed, or was a different preprocessing path used for the reported experiments? And if it is unintended, I'd be glad to help work out a version that recovers the personalization gains with matched encoding; I've found a few changes that recover some of it and would be happy to share.
Contributor guide
Research direction
Start with PRISM/prepare.py around lines 340–345 and trace the chosen and rejected values into PRISM/generate-prism-embeddings.py. Compare the released path with the encoding branch and its reported results to determine whether the representation difference is intentional or preprocessing inconsistency. Done means the intent is established and, if unintended, matched encoding and its effect on the reported accuracy are documented.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100