Glavin001 / Glavin001/PeakProgrammer

Initial foundation

Open
#1 0 comments 0 reactions 1 assignee Claimed by @Glavin001 View on GitHub
help wanted
Dominant language
Jupyter Notebook
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

- [ ] Pluggable fine-grained reward functions
- [ ] Reward
- [ ] Penalty
- [ ] Completion-wise feedback
- [ ] Sentence/Sequence-wise feedback
- [ ] Token-wise feedback

## Resources

- https://github.com/allenai/FineGrainedRLHF
- https://huggingface.co/teknium/Replit-v2-CodeInstruct-3B
- https://github.com/CarperAI/trlx/blob/main/examples/experiments/grounded_program_synthesis/train_trlx.py#L35-L53
- https://wandb.ai/carperai/summarize_RLHF/reports/Implementing-RLHF-Learning-to-Summarize-with-trlX--VmlldzozMzAwODM2

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.