cancervariants / cancervariants/therapy-normalization

Capture "pref_name" from ChEMBL

Open
#421 0 comments 0 reactions 0 assignees View on GitHub
ChEMBL enhancement priority:low
Dominant language
Python
Stars
15
Forks
3
PR merge metrics
No merged PRs in 30d

Description

As of v34

```
Pref_name curation. Progress has been made towards standardising drug and clinical candidate pref_names (in MOLECULE_DICTIONARY.PREF_NAME) whereby an approved drug name (FDA /EMA) is assigned in the first instance, if available. If not available, the USAN is assigned, followed by the INN name, respectively. If the USAN/INN name assignment is ambiguous, the FDA GSRS preferred name is used. A company research code, or Clinical Trial intervention name, is assigned if no standardised name is available. For virtual parent compounds, progress has been made towards assigning a distinct pref_name (typically based on the FDA GSRS preferred name) that differs from the child compound name.
```

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue provides ChEMBL v34 curation guidance but names no repository file, test, or entry point. Start by locating the existing ChEMBL ingestion or molecule-dictionary handling, then define how PREF_NAME should be captured and what validation demonstrates completion.

Written by the indexing model from the issue text.

Assessment

Domain
data, databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.