huggingface / huggingface/paperswithcode-feedback

[Dataset Request] Add AI4Privacy pii-masking-300k — PII detection/masking benchmark

Open
#30 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
10
Forks
1
PR merge metrics
No merged PRs in 30d

Description

Hi! I'd like to request indexing the **AI4Privacy pii-masking-300k** dataset and a leaderboard for **PII Detection / Masking**.

### Dataset
- **HF:** [`ai4privacy/pii-masking-300k`](https://huggingface.co/datasets/ai4privacy/pii-masking-300k)
- **Size:** ~225K+ examples (multilingual; English subset commonly used)
- **Labels:** 27+ PII types (names, emails, phones, addresses, etc.); token-level spans + masked text
- **Task:** PII Detection / Named-entity–style span extraction
- **Primary metric:** token/span-level F1

### Suggested results to seed the leaderboard
Verified from the model cards. ⚠️ **These are not directly comparable** — they were trained/evaluated on different versions of the ai4privacy family and report different metrics. Please seed them as separate entries keyed to the correct dataset version:

| Model | Metric | Value | Dataset version | License | Source |
|---|---|---|---|---|---|
| `Isotonic/deberta-v3-base_finetuned_ai4privacy_v2` | overall token F1 | **97.57%** (P 97.22 / R 97.92) | pii-masking-**200k** | CC-BY-NC-4.0 | [card](https://huggingface.co/Isotonic/deberta-v3-base_finetuned_ai4privacy_v2) |
| `iiiorg/piiranha-v1-detect-personal-information` (mDeBERTa) | multiclass F1 | **93.12%** (headline metric on card is 99.44% *accuracy*) | pii-masking-**400k** | CC-BY-NC-ND-4.0 | [card](https://huggingface.co/iiiorg/piiranha-v1-detect-personal-information) |

Note both are self-reported model-card numbers, not from an independent leaderboard. More can be imported from `pwc-archive/evaluation-tables` if present.

### Note
One of the most widely used open PII benchmarks; a leaderboard here would be valuable for the privacy/NER community. It may be worth structuring per dataset version (200k / 300k / 400k) since models target different releases. Thanks!

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no repository files, tests, or entry points. Start by finding how datasets and leaderboard results are indexed, then verify the pii-masking-300k metadata and determine how the 200k, 300k, and 400k versions should remain separate. Done means the requested dataset and leaderboard structure are represented without treating the self-reported model-card metrics as directly comparable.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface
Domain
data, machine-learning, security
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.