DoReCo 1.2
- Dominant language
- No language data
- Stars
- 2
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
# DoReCo version 1.2: Summary of changes
This posts gives an overview of the changes made from DoReCo 1.1 to DoReCo 1.2, released 16 December 2022.
### Improvements and other changes
- Fixed an issue where units adjacent to a \<\\> label were lost or appended to other units, creating overlaps; this also resolves an issue with start time codes of words appearing smaller than start time codes of the first corresponding phone (#10)
- wd_ID, mb_ID, and ph_ID columns in csvs now include leading zeros (e.g. w000012 instead of w12)
- Added one file to the Bora extended set
- Added proper tokenization to the English (Southern England) datset
- Substantially reworked the annotation of infixes, reduplication and other glosses in Gorwaa
- Substantially reworked the annotation of infixes and reduplication in Movima
- Improved time alignments in various files, including 09_Areal_History_pt1 (Pnar), Nah_02 (Teop), mc_english_kent02_b (English), MZP_KH_NARR_130907_JGD-CAIMAN (Movima), CJP-TXT-AN-00000-13 (Cabécar)
- Corrected one annotation in Arapaho based on feedback from corpus creator (thanks Andrew Cowell!)
- Added gloss abbreviation lists for 14 datasets
- Various minor changes to transcriptions and labels across several datasets
### Bug fixes
- Removed residual \ items from Gorwaa, Nǁng, Pnar, Texistepec Popoluca, and Yurakaré
- Fixed an issue where some speaker codes in the supplementary metadata table did not match the speaker codes in the _ph and _wd csvs
- Fixed an issue causing phones to be listed backwards in trin1278 ph.csv (#7)
- Corrected some entries in the corpus-wide metadata table (#8)
- Fixed an issue where the character û was not correctly encoded in Goemai (#9)
- Fixed an issue whereby some .eaf files did not include a link to a .wav file in their header
- Removed special characters from labeled items in Texistepec Popoluca (e.g. “\<\Bweenu,\>” -> \<\Bweenu\>)
- Fixed an issue whereby í was sometimes incorrectly represented as ï in some files from the Cabécar dataset
- Repaired a broken tier structure in one file from the Nǁng dataset
### Website functionality
- The "Sort by" function on the "Languages" site now works properly
- The webiste now provides separate links for core and extended datasets that lead to different file bundles (core vs. extended) (#6)
### Remaining issues
We are aware of a couple of remaining issues that we are planning to address in a future update.
- An issue causing empty units on the wd tier in one file from the Sanzhi Dargwa dataset
- A small number of overlaps around \<\\> labels in Evenki
- A small number of units with a start time of -1 in Evenki
- A modest number of overlaps (unrelated to \<\\>) in some Bora extended files
- Inconsistencies with the treatment of certain clitics in the Savosavo dataset
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.