DoReCo / DoReCo/doreco

DoReCo 1.2

Open
#11 0 comments 0 reactions 1 assignee Claimed by @LuPaschen View on GitHub
documentation
Dominant language
No language data
Stars
2
Forks
0
PR merge metrics
No merged PRs in 30d

Description

# DoReCo version 1.2: Summary of changes

This posts gives an overview of the changes made from DoReCo 1.1 to DoReCo 1.2, released 16 December 2022.

### Improvements and other changes

- Fixed an issue where units adjacent to a \<\\> label were lost or appended to other units, creating overlaps; this also resolves an issue with start time codes of words appearing smaller than start time codes of the first corresponding phone (#10)
- wd_ID, mb_ID, and ph_ID columns in csvs now include leading zeros (e.g. w000012 instead of w12)
- Added one file to the Bora extended set
- Added proper tokenization to the English (Southern England) datset
- Substantially reworked the annotation of infixes, reduplication and other glosses in Gorwaa
- Substantially reworked the annotation of infixes and reduplication in Movima
- Improved time alignments in various files, including 09_Areal_History_pt1 (Pnar), Nah_02 (Teop), mc_english_kent02_b (English), MZP_KH_NARR_130907_JGD-CAIMAN (Movima), CJP-TXT-AN-00000-13 (Cabécar)
- Corrected one annotation in Arapaho based on feedback from corpus creator (thanks Andrew Cowell!)
- Added gloss abbreviation lists for 14 datasets
- Various minor changes to transcriptions and labels across several datasets

### Bug fixes

- Removed residual \ items from Gorwaa, Nǁng, Pnar, Texistepec Popoluca, and Yurakaré
- Fixed an issue where some speaker codes in the supplementary metadata table did not match the speaker codes in the _ph and _wd csvs
- Fixed an issue causing phones to be listed backwards in trin1278 ph.csv (#7)
- Corrected some entries in the corpus-wide metadata table (#8)
- Fixed an issue where the character û was not correctly encoded in Goemai (#9)
- Fixed an issue whereby some .eaf files did not include a link to a .wav file in their header
- Removed special characters from labeled items in Texistepec Popoluca (e.g. “\<\Bweenu,\>” -> \<\Bweenu\>)
- Fixed an issue whereby í was sometimes incorrectly represented as ï in some files from the Cabécar dataset
- Repaired a broken tier structure in one file from the Nǁng dataset

### Website functionality

- The "Sort by" function on the "Languages" site now works properly
- The webiste now provides separate links for core and extended datasets that lead to different file bundles (core vs. extended) (#6)

### Remaining issues

We are aware of a couple of remaining issues that we are planning to address in a future update.

- An issue causing empty units on the wd tier in one file from the Sanzhi Dargwa dataset
- A small number of overlaps around \<\\> labels in Evenki
- A small number of units with a start time of -1 in Evenki
- A modest number of overlaps (unrelated to \<\\>) in some Bora extended files
- Inconsistencies with the treatment of certain clitics in the Savosavo dataset

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.