facebookresearch / facebookresearch/RAM

Inquiry about the Self-Taught Evaluator

Open
#22 0 comments 0 reactions 2 assignees Claimed by @xianxl View on GitHub
Dominant language
Python
Stars
382
Forks
49
PR merge metrics
No merged PRs in 30d

Description

The paper mentions that over 20,000 instructions categorized as “reasoning” were ultimately selected. However, the number of “reasoning”-related instructions obtained from WildChat far exceeds 20,000. How can the final 20,000 instructions used in the study be selected from this larger set? Following the settings described in the paper, I trained the model for 2 epochs, but my experiments indicate that the peak performance on RewardBench appears to be achieved after about one epoch, after which it begins to decline.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.