adobe / adobe/spacecat-task-processor

brand-profile agent resolves the WRONG Wikidata entity for same-named brands (not fixed by #359)

Ouverte
#360 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
JavaScript
Étoiles
1
Forks
1
Merge moyen
6 h 28 min
PR mergées (30 j)
13

Description

## Summary

The `brand-profile` agent resolves the **wrong Wikidata entity** for brands whose name collides with an unrelated same-named entity, because `findWikidataId` searches on the **brand name alone** and ignores the strongest disambiguator we already have: the **site domain** (e.g. the TLD) and the scraped site content.

PR #359 (adobe/spacecat-task-processor#359) anchors the Wikipedia *article* fetch to the resolved QID and guards on `wikibase_item === QID`. That fixes the case where the QID is correct but the article lookup went astray. It does **not** fix the case where the resolved **QID itself is wrong** — the agent then faithfully anchors to the wrong entity and still emits contaminated `products` / `sub_brands` / `competitors`.

This is the sibling of the data-quality report adobe-rnd/llmo-data-retrieval-service#3200.

## Two distinct failure classes

| Class | Cause | Example | Fixed by #359? |
|---|---|---|---|
| **A — QID correct, article wrong** | old step grabbed a same-named *article* despite a correct QID | Lovesac → `Q6690181` (LoveSac, furniture) ✅ | **Yes** — regen now returns furniture |
| **B — QID wrong** | entity *resolution* picked a same-named unrelated entity | Capella → `Q12970` (**the star Capella**) ❌ | **No** — regen still returns star data |

## Reproduction (Class B — verified on prod after #359)

`capella.edu` is Capella University (an online university; `.edu`). Its resolved QID is `Q12970`, which is the **star Capella** (constellation Auriga), not the university.

Before #359 (name-search article): products were a children's TV series ("Cappelli & Company", songs, cassettes).
After #359 regen: products became **stars** — `Capella Aa`, `Capella Ab`, `Capella H` (category "Star (primary giant in Capella system)"), with `products_metadata.brand_wikidata_id: "Q12970"` and `wikipedia_verified: true`. Still wrong — the guard passed because the article does match the (wrong) QID.

## Root cause

`src/agents/brand-profile/services/wikipedia.js` → `findWikidataId(brandName)`:
- calls `wbsearchentities` with `search: brandName`,
- returns the first result whose description contains a company term, else the first result.

There is no use of the **site domain** or scraped content to disambiguate. `capella.edu` (a `.edu`) is an unambiguous "this is a university, not a star" signal that is discarded.

## Prevalence (fleet scan)

Of **851** brands that have a stored products QID, ~**247** have a stored QID whose Wikidata entity does not match the brand (Class B; upper bound — includes some acronym false-positives such as AIB→"Allied Irish Banks", which is actually correct). Clear examples:

- AIA (insurance) → `Q25228` = **Anguilla** (a country)
- American Airlines → `Q838581` = **"comitative case"** (grammar)
- AAA Membership → `Q1823` = **Aceh**
- Capella University → `Q12970` = **the star Capella**
- OECD → polycrystalline silicon; Borussia Dortmund → BYD; Logista → Lovisa jewellery

## Blast radius

Same as #3200: `products.items[*].category` is the priority‑1 category source for synthetic-persona prompt generation, and the profile also feeds Semrush, Citation Attempt / Strategy Chat, and GSC. A wrong entity contaminates all of them.

## Suggested fixes

1. **Disambiguate with the site domain + scraped content.** Pass the base URL (esp. the registrable domain / TLD) and the industry/vertical into `findWikidataId`; prefer the candidate whose Wikidata sitelink/official-website (`P856`) or description matches the domain/industry. A `.edu` should never resolve to a star.
2. **Post-resolution guard.** After picking a QID, verify its Wikidata description/`instance of` (`P31`) is consistent with the brand's known vertical; reject on clear mismatch.
3. **Fail safe on low confidence.** When no candidate matches the domain/industry, return **no data** rather than a wrong entity (consistent with #359's "no data rather than wrong data" principle).

## Links
- Original data-quality report: adobe-rnd/llmo-data-retrieval-service#3200
- Partial fix (Class A): adobe/spacecat-task-processor#359

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Commencez dans src/agents/brand-profile/services/wikipedia.js, au niveau de findWikidataId(brandName), puis examinez comment brand-profile fournit le domaine du site et le secteur ou la verticale récupéré(e). Comparez le flux de résolution avec le correctif partiel de PR #359. Le travail est terminé lorsque les marques portant le même nom sont résolues vers un QID cohérent avec le domaine et le contenu, tandis que les correspondances à faible confiance ne renvoient aucune donnée au lieu de contaminer le profil.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
javascript
Domaine
backend, data, search
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Active
Clarté
Plutôt claire
Accessibilité débutants
52/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.