huggingface / huggingface/diffusers
eos_token_id for Textual Inversion
- Vorherrschende Sprache
- Python
- Sterne
- 34.5k
- Forks
- 7.3k
- Ø Merge
- 3 T. 3 Std.
- Gemergte PRs (30 T.)
- 91
Beschreibung
### Describe the bug
Hi, I implemented textual inversion follwowing this [link](https://huggingface.co/docs/diffusers/v0.32.2/en/training/text_inversion), but I think there is something wrong with `eos_token_id` in stable-diffusion-v1-5 text encoder [config](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/blob/main/text_encoder/config.json).
The config file is like this, this means `eos_token_id == 2`:
```
{
| "_name_or_path": "openai/clip-vit-large-patch14",
| "architectures": [
| "CLIPTextModel"
| ],
| "attention_dropout": 0.0,
| "bos_token_id": 0,
| "dropout": 0.0,
| "eos_token_id": 2,
| "hidden_act": "quick_gelu",
| "hidden_size": 768,
| "initializer_factor": 1.0,
| "initializer_range": 0.02,
| "intermediate_size": 3072,
| "layer_norm_eps": 1e-05,
| "max_position_embeddings": 77,
| "model_type": "clip_text_model",
| "num_attention_heads": 12,
| "num_hidden_layers": 12,
| "pad_token_id": 1,
| "projection_dim": 768,
| "torch_dtype": "float32",
| "transformers_version": "4.22.0.dev0",
| "vocab_size": 49408
| }
```
but in transformers modeling_clip.py,
```python
if self.eos_token_id == 2:
# The `eos_token_id` was incorrect before PR #24773: Let's keep what have been done here.
# A CLIP model with such `eos_token_id` in the config can't work correctly with extra new tokens added
# ------------------------------------------------------------
# text_embeds.shape = [batch_size, sequence_length, transformer.width]
# take features from the eot embedding (eot_token is the highest number in each sequence)
# casting to torch.int for onnx compatibility: argmax doesn't support int64 inputs with opset 14
pooled_output = last_hidden_state[
torch.arange(last_hidden_state.shape[0], device=last_hidden_state.device),
input_ids.to(dtype=torch.int, device=last_hidden_state.device).argmax(dim=-1),
]
```
I think this means current code is not compatible with textual inversion (because we just get the embedding of newly added token with token id 49408, not the eos token.
I might be wrong, but it will be really helpful for giving me any comments.
Thank you.
### Reproduction
accelerate launch textual_inversion.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATA_DIR \
--learnable_property="object" \
--placeholder_token="" \
--initializer_token="dog" \
--resolution=512 \
--train_batch_size=1 \
--gradient_accumulation_steps=4 \
--max_train_steps=3000 \
--learning_rate=5.0e-04 \
--scale_lr \
--lr_scheduler="constant" \
--lr_warmup_steps=0 \
### System Info
diffusers 0.32.0.dev0
### Who can help?
_No response_
Beitragsleitfaden
Rechercherichtung
Start with transformers modeling_clip.py and the textual_inversion.py reproduction command. Compare pooling behavior for the stable-diffusion-v1-5 text encoder before and after adding the placeholder token, checking whether the selected position is the EOS token. Done means the compatibility question is resolved and the demonstrated textual inversion workflow has verified behavior.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python, pytorch
- Bereich
- machine-learning
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 35/100