spancat training not working with span group other than "sc"
- Dominant language
- Python
- Stars
- 33.9k
- Forks
- 4.7k
- Avg merge
- 3m
- Merged PRs (30d)
- 1
Description
## How to reproduce the behaviour
When `span_key` in `[components.spancat_singlelabel]` or `[components.spancat]` sections is other than `"sc"`, training output looks like this:
```
ℹ Pipeline: ['sentencizer', 'tok2vec', 'spancat_singlelabel']
ℹ Set annotations on update for: ['sentencizer']
ℹ Initial learn rate: 0.001
E # LOSS TOK2VEC LOSS SPANC... SENTS_F SENTS_P SENTS_R SPANS_SC_F SPANS_SC_P SPANS_SC_R SCORE
--- ------ ------------ ------------- ------- ------- ------- ---------- ---------- ---------- ------
0 0 0.00 19.33 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 200 5.46 409.10 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 400 10.51 96.83 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 600 10.37 77.10 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 800 8.99 99.86 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 1000 9.52 100.14 100.00 100.00 100.00 0.00 0.00 0.00 0.50
0 1200 6.18 62.30 100.00 100.00 100.00 0.00 0.00 0.00 0.50
```
I thought I was getting insane :)
Seems like a bug at the evaluation step, I didn't investigate further, sorry.
Other parts seem to be working:
```
python -m spacy debug data config.cfg
```
is happy with non-"sc" value.
Also, the labels get picked up from the training dataset and show up correctly in `meta.json`.
After I changed `span_key` back to the default "sc", I got this:
```
E # LOSS TOK2VEC LOSS SPANC... SENTS_F SENTS_P SENTS_R SPANS_SC_F SPANS_SC_P SPANS_SC_R SCORE
--- ------ ------------ ------------- ------- ------- ------- ---------- ---------- ---------- ------
0 0 0.00 19.33 100.00 100.00 100.00 96.81 96.81 96.81 0.98
0 200 6.64 421.19 100.00 100.00 100.00 99.32 99.32 99.32 1.00
0 400 8.43 82.30 100.00 100.00 100.00 99.43 99.43 99.43 1.00
0 600 9.34 71.43 100.00 100.00 100.00 99.45 99.45 99.45 1.00
0 800 9.28 107.05 100.00 100.00 100.00 99.59 99.59 99.59 1.00
0 1000 8.78 89.08 100.00 100.00 100.00 98.95 98.95 98.95 0.99
```
Also, not sure why sentence metrics show up - I am using non-trainable simple sentencizer. It's not that important, obviously.
## Your Environment
- **spaCy version:** 3.6.1
- **Platform:** macOS-12.4-x86_64-i386-64bit
- **Python version:** 3.10.11
- **Pipelines:** en_core_web_sm (3.6.0)
Contributor guide
Assessment
This issue has not been assessed yet.