AOSSIE-Org / AOSSIE-Org/OpenVerifiableLLM

[FEATURE]: Deterministic Dataset Tokenization with Merkle Verification

Open
#61 1 comment 0 reactions 1 assignee Claimed by @Varshiniputtabakula View on GitHub
enhancement
Dominant language
Python
Stars
18
Forks
31
Avg merge
1m
Merged PRs (30d)
2

Description

### Feature and its Use Cases

#### Problem

Currently the tokenizer training pipeline introduced in **PR #17** trains and saves the tokenizer configuration but does **not tokenize the actual Wikipedia dataset**.

This leaves a verification gap between the **verified preprocessing pipeline** and the **model training stage**.

Because the tokenized dataset is not verified, the following risks exist:

* The dataset could be tokenized using a different tokenizer than claimed
* Tokenized outputs could be modified before training
* There is no cryptographic linkage between preprocessing and training

---

#### Proposed Solution

Implement a **deterministic dataset tokenization pipeline** that:

1. Uses the trained tokenizer to tokenize the cleaned Wikipedia dataset
2. Saves the tokenized dataset as a deterministic artifact
3. Computes a **Merkle root over tokenized chunks**
4. Adds the tokenized dataset hash to the verification manifest

This ensures the tokenized dataset is **cryptographically tied to the tokenizer configuration and preprocessing outputs**.

---

#### Builds On

* **PR #17** — BaseTokenizer, BPETokenizer, SentencePieceTokenizer

---

#### Outcome

This will extend the verification pipeline by adding a **verifiable tokenized dataset layer**, creating the following chain:

```
Raw Dataset → Processed Dataset → Tokenizer Config → Tokenized Dataset → Model Training
```

Happy to implement this if the approach looks good.

### Additional Context

_No response_

### Code of Conduct

- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.