AnswerDotAI / AnswerDotAI/RAGatouille

Question: Guidance on using language_code for custom, non-linguistic data

Open
#273 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
4k
Forks
276
PR merge metrics
No merged PRs in 30d

Description

Hi there, thanks for this amazing library! It seems perfect for a project I'm working on.

I am training a ColBERT model from scratch using a base BERT model that was pre-trained on my own custom, non-linguistic data.

My data is highly structured and symbolic. A typical sequence looks like this: 0_common 1_rare [MASK] 0_low_frequency. It's essentially a space-delimited list of custom tokens, not a natural language.

I'm looking at the RAGTrainer documentation and I have a question about the language_code parameter. The docs explain that this is used to get relevant processing utilities for languages like English ('en') or Japanese ('ja'), which makes perfect sense for things like sentence splitting.

My question is: What is the best practice for the language_code parameter when working with custom, non-linguistic data like mine?

Should I:

Omit the language_code parameter entirely?
Explicitly set it to language_code=None?
Is there a specific value (e.g., 'any', 'raw', 'none') that signifies raw/generic processing?
My main goal is to ensure that RAGatouille does not apply any language-specific normalization or sentence-splitting logic to my data, and that it's treated simply as a sequence of space-delimited tokens.

Any guidance on this would be greatly appreciated. Thank you!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the RAGTrainer documentation and inspect how the language_code parameter is described and handled. Determine the documented behavior for custom, non-linguistic, space-delimited data, including whether omission and None are supported. Done means the recommended usage and raw-processing behavior are clearly documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
documentation
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.