adobe / adobe/spacecat-task-processor

brand-profile agent resolves the WRONG Wikidata entity for same-named brands (not fixed by #359)

Aperta
#360 1 commento 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
JavaScript
Stelle
1
Fork
1
Merge medio
6h 28m
PR unite (30g)
13

Descrizione

## Summary

The `brand-profile` agent resolves the **wrong Wikidata entity** for brands whose name collides with an unrelated same-named entity, because `findWikidataId` searches on the **brand name alone** and ignores the strongest disambiguator we already have: the **site domain** (e.g. the TLD) and the scraped site content.

PR #359 (adobe/spacecat-task-processor#359) anchors the Wikipedia *article* fetch to the resolved QID and guards on `wikibase_item === QID`. That fixes the case where the QID is correct but the article lookup went astray. It does **not** fix the case where the resolved **QID itself is wrong** — the agent then faithfully anchors to the wrong entity and still emits contaminated `products` / `sub_brands` / `competitors`.

This is the sibling of the data-quality report adobe-rnd/llmo-data-retrieval-service#3200.

## Two distinct failure classes

| Class | Cause | Example | Fixed by #359? |
|---|---|---|---|
| **A — QID correct, article wrong** | old step grabbed a same-named *article* despite a correct QID | Lovesac → `Q6690181` (LoveSac, furniture) ✅ | **Yes** — regen now returns furniture |
| **B — QID wrong** | entity *resolution* picked a same-named unrelated entity | Capella → `Q12970` (**the star Capella**) ❌ | **No** — regen still returns star data |

## Reproduction (Class B — verified on prod after #359)

`capella.edu` is Capella University (an online university; `.edu`). Its resolved QID is `Q12970`, which is the **star Capella** (constellation Auriga), not the university.

Before #359 (name-search article): products were a children's TV series ("Cappelli & Company", songs, cassettes).
After #359 regen: products became **stars** — `Capella Aa`, `Capella Ab`, `Capella H` (category "Star (primary giant in Capella system)"), with `products_metadata.brand_wikidata_id: "Q12970"` and `wikipedia_verified: true`. Still wrong — the guard passed because the article does match the (wrong) QID.

## Root cause

`src/agents/brand-profile/services/wikipedia.js` → `findWikidataId(brandName)`:
- calls `wbsearchentities` with `search: brandName`,
- returns the first result whose description contains a company term, else the first result.

There is no use of the **site domain** or scraped content to disambiguate. `capella.edu` (a `.edu`) is an unambiguous "this is a university, not a star" signal that is discarded.

## Prevalence (fleet scan)

Of **851** brands that have a stored products QID, ~**247** have a stored QID whose Wikidata entity does not match the brand (Class B; upper bound — includes some acronym false-positives such as AIB→"Allied Irish Banks", which is actually correct). Clear examples:

- AIA (insurance) → `Q25228` = **Anguilla** (a country)
- American Airlines → `Q838581` = **"comitative case"** (grammar)
- AAA Membership → `Q1823` = **Aceh**
- Capella University → `Q12970` = **the star Capella**
- OECD → polycrystalline silicon; Borussia Dortmund → BYD; Logista → Lovisa jewellery

## Blast radius

Same as #3200: `products.items[*].category` is the priority‑1 category source for synthetic-persona prompt generation, and the profile also feeds Semrush, Citation Attempt / Strategy Chat, and GSC. A wrong entity contaminates all of them.

## Suggested fixes

1. **Disambiguate with the site domain + scraped content.** Pass the base URL (esp. the registrable domain / TLD) and the industry/vertical into `findWikidataId`; prefer the candidate whose Wikidata sitelink/official-website (`P856`) or description matches the domain/industry. A `.edu` should never resolve to a star.
2. **Post-resolution guard.** After picking a QID, verify its Wikidata description/`instance of` (`P31`) is consistent with the brand's known vertical; reject on clear mismatch.
3. **Fail safe on low confidence.** When no candidate matches the domain/industry, return **no data** rather than a wrong entity (consistent with #359's "no data rather than wrong data" principle).

## Links
- Original data-quality report: adobe-rnd/llmo-data-retrieval-service#3200
- Partial fix (Class A): adobe/spacecat-task-processor#359

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia in src/agents/brand-profile/services/wikipedia.js, in findWikidataId(brandName), poi esamina come brand-profile fornisce il dominio del sito e il settore o la verticale estratti. Confronta il flusso di risoluzione con la correzione parziale in PR #359. Il lavoro è completo quando i brand con lo stesso nome vengono risolti in un QID coerente con il dominio e il contenuto, mentre le corrispondenze a bassa confidenza non restituiscono dati invece di contaminare il profilo.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
javascript
Ambito
backend, data, search
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
52/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.