[bug] Issue in CLIP model code for ch16
- Dominant language
- Jupyter Notebook
- Stars
- 2k
- Forks
- 624
- Avg merge
- 7h 55m
- Merged PRs (30d)
- 1
Description
### Enter the chapter number
Ch16
### Enter the page number
674
### What is the cell's number in the notebook
cell 39 in clip section
### Enter the environment you are using to run the notebook
Colab
### Describe your issue
get_image_features and get_text_features used to return a plain torch.FloatTensor, but now they output a BaseModelOutputWithPooling. We cannot use .norm with this object, leading to an error in this line:
```python
image_features /= image_features.norm(dim=1, keepdim=True)
```
with the error being:
```python
AttributeError: 'BaseModelOutputWithPooling' object has no attribute 'norm'
```
### If you found a workaround, describe it here
I have found a workaround after taking a look at different discussions in huggingface
```python
output = self.clip_model.get_image_features(**inputs)
if not isinstance(output, torch.Tensor):
output = output.pooler_output
```
adding the following line fixes the issue
```python
image_features = image_features.pooler_output
```
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.