AOSSIE-Org / AOSSIE-Org/OpenVerifiableLLM
[FEATURE]: Implement Deterministic Dataset Encoding Pipeline and Verifiable LLaMA Model (Model Architecture)
- 主要語言
- Python
- 星號
- 18
- 分支
- 31
- 平均合併
- 1 分鐘
- 30 天內合併 PR
- 2
描述
# deterministic dataset encoding pipeline and a minimal implementation of a LLaMA architecture.
* The dataset is currently processed into wiki_clean.txt and a tokenizer has been trained. We need to implement a memory-efficient script to encode the entire dataset into binary format for training.
* Implement a minimal PyTorch LLaMA-style architecture to ensure deterministic behavior and full control over initialization.
* Read the dataset in chunks to avoid high memory usage.
* Use the trained tokenizer (BPE/SentencePiece) to convert text into token IDs.
* Stream token IDs into a binary dataset file (.bin, uint16 or similar).
* Compute a SHA256 hash of the resulting file.
# Verification Criteria
* Running the dataset encoding pipeline twice should produce identical binary files and SHA256 hashes.
* Initializing the model twice with the same seed should produce identical parameter hashes.
must output the exact same initial parameter hashes.
### Additional Context
_No response_
### Code of Conduct
- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates
貢獻指南
評估
這個 Issue 還沒有評估資料。