google / google/crossmodal-3600
Tokenization method
Open
- Dominant language
- HTML
- Stars
- 10
- Forks
- 2
- PR merge metrics
- No merged PRs in 30d
Description
Hey, thanks for the great work! I was wondering if you could elaborate on the tokenization method used for the caption/tokenized field in XM3600. This would include the normalization strategy and neural segmentation for languages without word boundaries. I apologize if I have overlooked this information.
Thanks and best regards,
Julian
Contributor guide
Assessment
This issue has not been assessed yet.