google / google/crossmodal-3600

Tokenization method

Open
#4 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
HTML
Stars
10
Forks
2
PR merge metrics
No merged PRs in 30d

Description

Hey, thanks for the great work! I was wondering if you could elaborate on the tokenization method used for the caption/tokenized field in XM3600. This would include the normalization strategy and neural segmentation for languages without word boundaries. I apologize if I have overlooked this information.

Thanks and best regards,
Julian

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.