bigscience-workshop / bigscience-workshop/biomedical

Proposal to add the MedSTS dataset

Open
#356 1 comment 0 reactions 0 assignees View on GitHub
High New Dataset Private Semantic Textual Similarity
Dominant language
Python
Stars
505
Forks
117
PR merge metrics
No merged PRs in 30d

Description

## Adding a Dataset
- **Name:** MedSTS
- **Description:** 1,068 sentence pairs annotated by two medical experts with semantic similarity scores of 0-5 (low to high similarity).
- **Task:** STS
- **Paper:** https://arxiv.org/abs/1808.09397
- **Data:** (must be asked by email to Mayo Clinic)
- **License:** (probably depends on agreement with Mayo Clinic)
- **Motivation:** one of the largest clinical text similarity dataset

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.