EleutherAI / EleutherAI/w2s

Hellaswag result not reproducing

Open
#3 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
25
Forks
4
PR merge metrics
No merged PRs in 30d

Description

Hi,

Thanks a lot for the codebase! It works quite smoothly and I was able to reproduce results across many datasets using the default sft_config provided, but not on hellaswag. I've attached my plot (weak_ft, strong_ft are floor and ceil respectively).

Specifically, I'm simply running `python run.py --dataset="$arg_1"` using the default Qwen-1.5-0.5b and Llama-3-8b models as the weak, strong pair. Are the configurations used to produce the main weak to strong plot different (specifically across datasets) than the ones provided by default in `sft_config.py`. If so, could you please provide detailed config files for reproduction, especially if there are any changes needed for Hellaswag?

Thanks!
![image](https://github.com/user-attachments/assets/90b126a0-47ea-415a-8319-361ed555e7b6)

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.