DAMO-NLP-SG / DAMO-NLP-SG/multilingual_analysis

Questions about neuron enchancement

Open
#3 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
52
Forks
11
PR merge metrics
No merged PRs in 30d

Description

@zhaoyiran924, Thanks for sharing your great work! I have a question about the implementation of `train_neuron.py` and its data format.

### Current Implementation
According to the paper, I thought it was using raw Wikipedia's passage to train those neurons, in a next-token prediction way,
while the code currently seems to use a question-answer format for training:

```python
def formatting_prompts_func(example):
output_texts = []
for i in range(len(example['original_question'])):
text = f"{example['original_question'][i]}. {example['response'][i]}"
output_texts.append(text)
return output_texts
```

Questions

1. Could you please confirm the format of how the wiki data passed,
or share an example of how the Wikipedia documents are preprocessed to get the 'original_question' and 'response' fields?

2. Would it make more sense to use a simpler format for plain text documents, like:
```
def formatting_prompts_func(example):
return example['text']
```

Thanks for your help.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.