CarperAI / CarperAI/trlx

Well-known RMs and eval utilities

Open
#415 1 comment 0 reactions 0 assignees View on GitHub
feature request
Dominant language
Python
Stars
4.8k
Forks
487
PR merge metrics
No merged PRs in 30d

Description

### 🚀 The feature, motivation, and pitch

Collect together settings and commonly used reward models and evaluations. These RMs can be used for training time eval, but we would probably also want to use multiple-choice evals too

Ideas for RMs:
* Our HH RMs
* Sentiments
* SteamSHP?
* OpenAI API GPT3.5/4 (probably only usable for test time?)

Multiple Choice evals:
* Anthropics model generated evals
* OpenAIs new evals
* `lm-evaluation-harness` supported evals (would need to add ability to run model ourselves for thaT?)
* Ranking evals: Provide a pre-ranked set of responses and have model order them from most to least aligned

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.