AOSSIE-Org / AOSSIE-Org/OpenVerifiableLLM
[FEATURE]: Deterministic Dataset Tokenization with Merkle Verification
- Ngôn ngữ chính
- Python
- Star
- 18
- Fork
- 31
- Merge trung bình
- 1 phút
- Pull request đã merge (30 ngày)
- 2
Mô tả
### Feature and its Use Cases
#### Problem
Currently the tokenizer training pipeline introduced in **PR #17** trains and saves the tokenizer configuration but does **not tokenize the actual Wikipedia dataset**.
This leaves a verification gap between the **verified preprocessing pipeline** and the **model training stage**.
Because the tokenized dataset is not verified, the following risks exist:
* The dataset could be tokenized using a different tokenizer than claimed
* Tokenized outputs could be modified before training
* There is no cryptographic linkage between preprocessing and training
---
#### Proposed Solution
Implement a **deterministic dataset tokenization pipeline** that:
1. Uses the trained tokenizer to tokenize the cleaned Wikipedia dataset
2. Saves the tokenized dataset as a deterministic artifact
3. Computes a **Merkle root over tokenized chunks**
4. Adds the tokenized dataset hash to the verification manifest
This ensures the tokenized dataset is **cryptographically tied to the tokenizer configuration and preprocessing outputs**.
---
#### Builds On
* **PR #17** — BaseTokenizer, BPETokenizer, SentencePieceTokenizer
---
#### Outcome
This will extend the verification pipeline by adding a **verifiable tokenized dataset layer**, creating the following chain:
```
Raw Dataset → Processed Dataset → Tokenizer Config → Tokenized Dataset → Model Training
```
Happy to implement this if the approach looks good.
### Additional Context
_No response_
### Code of Conduct
- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.