bigscience-workshop / bigscience-workshop/biomedical

Create dataset loader for CASI (Clinical Abbreviation Sense Inventory)

Open
#669 1 comment 0 reactions 0 assignees View on GitHub
New Dataset
Dominant language
Python
Stars
505
Forks
117
PR merge metrics
No merged PRs in 30d

Description

## Adding a Dataset
- **Name:** *CASI (Clinical Abbreviation Sense Inventory)*
- **Description:** *For our comprehensive sense inventory for clinical abbreviations and acronyms, a total of 440 most frequently used abbreviations and acronyms were selected from 352,267 dictated clinical notes.*
- **Task:** *Word Sense Disambiguation*
- **Paper:** *https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3932450/*
- **Data:** *https://conservancy.umn.edu/handle/11299/137703*
- **License:** *Type of license; please provide public for new datasets*
- **Motivation:** *Clinical dataset used to build zero shot prompts in "Large Language Models are Zero-Shot Clinical Information Extractors" (Agrawal et al. 2022)*

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.