adobe / adobe/spacecat-task-processor
brand-profile agent resolves the WRONG Wikidata entity for same-named brands (not fixed by #359)
- Lingua principale
- JavaScript
- Stelle
- 1
- Fork
- 1
- Merge medio
- 6h 28m
- PR unite (30g)
- 13
Descrizione
## Summary
The `brand-profile` agent resolves the **wrong Wikidata entity** for brands whose name collides with an unrelated same-named entity, because `findWikidataId` searches on the **brand name alone** and ignores the strongest disambiguator we already have: the **site domain** (e.g. the TLD) and the scraped site content.
PR #359 (adobe/spacecat-task-processor#359) anchors the Wikipedia *article* fetch to the resolved QID and guards on `wikibase_item === QID`. That fixes the case where the QID is correct but the article lookup went astray. It does **not** fix the case where the resolved **QID itself is wrong** — the agent then faithfully anchors to the wrong entity and still emits contaminated `products` / `sub_brands` / `competitors`.
This is the sibling of the data-quality report adobe-rnd/llmo-data-retrieval-service#3200.
## Two distinct failure classes
| Class | Cause | Example | Fixed by #359? |
|---|---|---|---|
| **A — QID correct, article wrong** | old step grabbed a same-named *article* despite a correct QID | Lovesac → `Q6690181` (LoveSac, furniture) ✅ | **Yes** — regen now returns furniture |
| **B — QID wrong** | entity *resolution* picked a same-named unrelated entity | Capella → `Q12970` (**the star Capella**) ❌ | **No** — regen still returns star data |
## Reproduction (Class B — verified on prod after #359)
`capella.edu` is Capella University (an online university; `.edu`). Its resolved QID is `Q12970`, which is the **star Capella** (constellation Auriga), not the university.
Before #359 (name-search article): products were a children's TV series ("Cappelli & Company", songs, cassettes).
After #359 regen: products became **stars** — `Capella Aa`, `Capella Ab`, `Capella H` (category "Star (primary giant in Capella system)"), with `products_metadata.brand_wikidata_id: "Q12970"` and `wikipedia_verified: true`. Still wrong — the guard passed because the article does match the (wrong) QID.
## Root cause
`src/agents/brand-profile/services/wikipedia.js` → `findWikidataId(brandName)`:
- calls `wbsearchentities` with `search: brandName`,
- returns the first result whose description contains a company term, else the first result.
There is no use of the **site domain** or scraped content to disambiguate. `capella.edu` (a `.edu`) is an unambiguous "this is a university, not a star" signal that is discarded.
## Prevalence (fleet scan)
Of **851** brands that have a stored products QID, ~**247** have a stored QID whose Wikidata entity does not match the brand (Class B; upper bound — includes some acronym false-positives such as AIB→"Allied Irish Banks", which is actually correct). Clear examples:
- AIA (insurance) → `Q25228` = **Anguilla** (a country)
- American Airlines → `Q838581` = **"comitative case"** (grammar)
- AAA Membership → `Q1823` = **Aceh**
- Capella University → `Q12970` = **the star Capella**
- OECD → polycrystalline silicon; Borussia Dortmund → BYD; Logista → Lovisa jewellery
## Blast radius
Same as #3200: `products.items[*].category` is the priority‑1 category source for synthetic-persona prompt generation, and the profile also feeds Semrush, Citation Attempt / Strategy Chat, and GSC. A wrong entity contaminates all of them.
## Suggested fixes
1. **Disambiguate with the site domain + scraped content.** Pass the base URL (esp. the registrable domain / TLD) and the industry/vertical into `findWikidataId`; prefer the candidate whose Wikidata sitelink/official-website (`P856`) or description matches the domain/industry. A `.edu` should never resolve to a star.
2. **Post-resolution guard.** After picking a QID, verify its Wikidata description/`instance of` (`P31`) is consistent with the brand's known vertical; reject on clear mismatch.
3. **Fail safe on low confidence.** When no candidate matches the domain/industry, return **no data** rather than a wrong entity (consistent with #359's "no data rather than wrong data" principle).
## Links
- Original data-quality report: adobe-rnd/llmo-data-retrieval-service#3200
- Partial fix (Class A): adobe/spacecat-task-processor#359
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia in src/agents/brand-profile/services/wikipedia.js, in findWikidataId(brandName), poi esamina come brand-profile fornisce il dominio del sito e il settore o la verticale estratti. Confronta il flusso di risoluzione con la correzione parziale in PR #359. Il lavoro è completo quando i brand con lo stesso nome vengono risolti in un QID coerente con il dominio e il contenuto, mentre le corrispondenze a bassa confidenza non restituiscono dati invece di contaminare il profilo.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- javascript
- Ambito
- backend, data, search
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Attiva
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 52/100