DAMO-NLP-SG / DAMO-NLP-SG/multilingual_analysis

Issues on reproducing the paper results

Open
#11 0 comments 2 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
52
Forks
11
PR merge metrics
No merged PRs in 30d

Description

I was able to run the codes for three stages and each requires a new virtual environment. It troubled when it comes to training safety neurons.
I ran Llama3-8B-Instruct with the following requirements:
`transformers==4.38.2`
`peft==0.10.0`
`trl==0.9.6`
`accelerate==0.43.2`
Along with replacing `/conda/env/path/site-packages/transformers/trainer.py` with the `transformers/trainer.py` provided in this repo, you also need to append the definition of `activate_neurons` to `/conda/env/path/site-packages/transformers/training_args.py` in line 2786.

Image

After training, I checked the changes in parameters and discovered checkpoint saved was identical to the original model. This could be solved by modifying the saving strategy as follows:
```python
is_main_process = True
if hasattr(trainer, "accelerator"):
is_main_process = trainer.accelerator.is_main_process

if is_main_process:
try:
model_to_save = trainer.accelerator.unwrap_model(trainer.model)
except Exception:
model_to_save = trainer.model

if isinstance(model_to_save, PeftModel):
model_to_save.save_pretrained(output_dir)
tokenizer.save_pretrained(output_dir)
else:
trainer.save_model(output_dir)
tokenizer.save_pretrained(output_dir)
```

However, I was not able to reproduce the result in the paper. Here are the training logs of safety-neuron tuned version and all parameter tuned version:
Safety-Neuron tuned version:
Image
All parameters tuned version:

Image

SFT data: the 50 samples randomly selected from the training data in repo of Circuit-Break(https://arxiv.org/pdf/2406.04313):

[circuit_breakers_train_sample50.json](https://github.com/user-attachments/files/23598855/circuit_breakers_train_sample50.json)

I compared math capability using gsm8k-250 English and safety using MultiJail-EN:
The tag safe/unsafe/invalid is measured with the prompt provided in MultiJail paper (https://openreview.net/pdf?id=vESNKdEMGp)

The result was rather wired:

Image

I noticed that in the paper there was no comparison with the all parameters tuned with same SFT data. I expected that all-param SFT would perform worse in math tasks and have equivalent level or less safe than the safety-neuron tuned version.

Could you provide more information on the running environment that I could replicate the experimental result? Or could you open-source your tuned models for reference?

Thanks

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the supplied requirements and the modified transformers/trainer.py and training_args.py files, then trace the safety-neuron training and checkpoint-saving steps. Compare the reported safety-neuron and all-parameter logs with the paper’s evaluation setup and determine whether the environment or saved model differs. Done means the reproduction discrepancy is explained or the required environment and reference model are documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.