explosion / explosion/spaCy

Italian & Spanish NER shouldn't extract "Google" or "Facebook"?

Open
#13,551 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
33.9k
Forks
4.7k
Avg merge
3m
Merged PRs (30d)
1

Description

Extracting entities from news articles I've realized this behavior:

![image](https://github.com/explosion/spaCy/assets/158814524/3a92b3aa-4359-462a-983d-e27abd393c85)

These words are present in articles but are not extracted by the models.

Does anyone know the reason?

## Info about spaCy

- **spaCy version:** 3.7.5
- **Platform:** Linux-6.1.85+-x86_64-with-glibc2.35
- **Python version:** 3.10.12
- **Pipelines:** es_core_news_lg (3.7.0), it_core_news_lg (3.7.0)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.