explosion / explosion/spaCy

spancat training not working with span group other than "sc"

Open
#13,090 3 comments 0 reactions 0 assignees View on GitHub
feat / spancat feat / training training
Dominant language
Python
Stars
33.9k
Forks
4.7k
Avg merge
3m
Merged PRs (30d)
1

Description

## How to reproduce the behaviour

When `span_key` in `[components.spancat_singlelabel]` or `[components.spancat]` sections is other than `"sc"`, training output looks like this:
```
ℹ Pipeline: ['sentencizer', 'tok2vec', 'spancat_singlelabel']
ℹ Set annotations on update for: ['sentencizer']
ℹ Initial learn rate: 0.001

E # LOSS TOK2VEC LOSS SPANC... SENTS_F SENTS_P SENTS_R SPANS_SC_F SPANS_SC_P SPANS_SC_R SCORE
--- ------ ------------ ------------- ------- ------- ------- ---------- ---------- ---------- ------
0 0 0.00 19.33 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 200 5.46 409.10 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 400 10.51 96.83 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 600 10.37 77.10 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 800 8.99 99.86 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 1000 9.52 100.14 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 1200 6.18 62.30 100.00 100.00 100.00 0.00 0.00 0.00 0.50

```
I thought I was getting insane :)
Seems like a bug at the evaluation step, I didn't investigate further, sorry.
Other parts seem to be working:
```
python -m spacy debug data config.cfg
```
is happy with non-"sc" value.

Also, the labels get picked up from the training dataset and show up correctly in `meta.json`.

After I changed `span_key` back to the default "sc", I got this:
```
E # LOSS TOK2VEC LOSS SPANC... SENTS_F SENTS_P SENTS_R SPANS_SC_F SPANS_SC_P SPANS_SC_R SCORE
--- ------ ------------ ------------- ------- ------- ------- ---------- ---------- ---------- ------
0 0 0.00 19.33 100.00 100.00 100.00 96.81 96.81 96.81 0.98
0 200 6.64 421.19 100.00 100.00 100.00 99.32 99.32 99.32 1.00
0 400 8.43 82.30 100.00 100.00 100.00 99.43 99.43 99.43 1.00
0 600 9.34 71.43 100.00 100.00 100.00 99.45 99.45 99.45 1.00
0 800 9.28 107.05 100.00 100.00 100.00 99.59 99.59 99.59 1.00
0 1000 8.78 89.08 100.00 100.00 100.00 98.95 98.95 98.95 0.99
```

Also, not sure why sentence metrics show up - I am using non-trainable simple sentencizer. It's not that important, obviously.

## Your Environment

- **spaCy version:** 3.6.1
- **Platform:** macOS-12.4-x86_64-i386-64bit
- **Python version:** 3.10.11
- **Pipelines:** en_core_web_sm (3.6.0)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.