CarperAI / CarperAI/autocrit

Reproduce Constitutional AI Steps

Open
#2 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
87
Forks
17
PR merge metrics
No merged PRs in 30d

Description

# Overview

This issue captures some of the key steps required to reproduce the Constitutional AI paper steps to fine tune a RLHF model with feedback generated by a RLAIF model.

## Phase One
image

- [ ] Gather a dataset of harmful prompts
- [ ] Create a base script to compose prompts using a base *constitution*
- [ ] Generate a new dataset of prompts + responses using Carper's [GPT-J RLHF](https://huggingface.co/reciprocate/ppo_hh_gpt-j) to review / critique the output
- [ ] Fine-tune the original model on revised responses using supervised learning

## Phase Two
image

- [ ] Sample the fine tuned model using the dataset of harmful prompts to create a new dataset with multiple outputs
- [ ] Train a "reward model' (i.e. https://github.com/Dahoas/reward-modeling) to select the best result (fine tuned preference model)
- [ ] Use RLAIF training to fine tune the RLHF model

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.