Glavin001 / Glavin001/PeakProgrammer
Initial foundation
Open
help wanted
- Dominant language
- Jupyter Notebook
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
- [ ] Pluggable fine-grained reward functions
- [ ] Reward
- [ ] Penalty
- [ ] Completion-wise feedback
- [ ] Sentence/Sequence-wise feedback
- [ ] Token-wise feedback
## Resources
- https://github.com/allenai/FineGrainedRLHF
- https://huggingface.co/teknium/Replit-v2-CodeInstruct-3B
- https://github.com/CarperAI/trlx/blob/main/examples/experiments/grounded_program_synthesis/train_trlx.py#L35-L53
- https://wandb.ai/carperai/summarize_RLHF/reports/Implementing-RLHF-Learning-to-Summarize-with-trlX--VmlldzozMzAwODM2
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.