AOSSIE-Org / AOSSIE-Org/OpenVerifiableLLM

[FEATURE]: Complete SentencePiece tokenizer — encode, decode, load and input validation

Đang mở
#51 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
enhancement
Ngôn ngữ chính
Python
Star
18
Fork
31
Merge trung bình
1 phút
Pull request đã merge (30 ngày)
2

Mô tả

## Summary

This issue proposes completing the SentencePieceTokenizer introduced in #17. The class currently supports training only. It is missing **encode**(), **decode**(), **load**(), **input validation,** special tokens, and directory creation


## Background

PR #17 introduced a modular tokenizer architecture with SentencePieceTokenizer as one of the two implementations.

However the current SentencePieceTokenizer only covers the training step

A tokenizer that cannot encode or decode text cannot be used in any downstream pipeline task.


## Problems in Current Implementation

### 1. No encode() or decode()

### 2. No load()

### 3. Missing special tokens in training

### 4. No input validation

### 5. save_path directory never created

## Why This Matters

SentencePiece is used by major modern LLMs:

Without a complete SentencePieceTokenizer, the project cannot support cryptographic verification of pipelines
built on any of these models — which represent the majority of current open source LLMs.


## Scope

This is a change touching only:
- sentencepiece_tokenizer.py

No overlap with any existing open PRs.

## Related

- PR #17 — feat: deterministic tokenizer training and config hashing

### Additional Context

_No response_

### Code of Conduct

- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.