huggingface / huggingface/alignment-handbook

Cannot apply "run_dpo.py" on a trained Axolotl model

Open
#105 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
5.7k
Forks
490
Avg merge
2m
Merged PRs (30d)
1

Description

After using Axolotl to SFT my mistral7b model I tried to align it using DPO
At some point in the code (in the DPOTrainer initialization) the code freezes and stops after timeout is reached.
When trying to run the script on the base model (https://huggingface.co/TokenBender/pic_7B_mistral_Full_v0.2) it works well.
Attaching a screenshot of the part where it freezes.

image

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.