DoReCo / DoReCo/doreco

segmentation mistakes and lost items in some of the Komnzo texts

Open
#16 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
2
Forks
0
PR merge metrics
No merged PRs in 30d

Description

I have noticed mistakes of the following type: words are "lost" on the mb@-tier, gl@-tier, ps@-tier, and doreco-m-tier. These words are correctly included on the word tier (tx@) and the phonetic tier (ph@).

two examples from the same text (tci20100905a):
- 0061_doreco_komn1238_tci20100905a: _gardame_ appears on the wd@-tier, but is then lost on all dependent tiers
- 0065_doreco_komn1238_tci20100905a: same as above with _gäwkarä_

There are many more such mistakes in the same text. I cannot tell how common this kind of mistake is in the doreco data, but I have seen it in at least one other text (tci20100905a):
- 0044_doreco_komn1238_tci20100905a: _kabe_
- 0045_doreco_komn1238_tci20100905a: _boba_
- 0046_doreco_komn1238_tci20100905a: _ʔetfəth_ and _kafar_ (see screenshot below)

The problem seems to be caused by a missing space in the wd@-tier, e.g. "ʔetfəthmənzen" on the wd@-tier when it should be "ʔetfəth mənzen" (screenshot attached below).

PS: please contact me (via email) if you need help/advice with the Komnzo data.

Image

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.