OpenEuroLLM / OpenEuroLLM/Taskboard

Convert the current tokenizer to handle the qwen chat-template faithfully.

Open
#370 5 comments 0 reactions 1 assignee View on GitHub

@Apsod is already working on this.

Since Aug 17, 2026.

4.6 post-training
Dominant language
No language data
Stars
3
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Goal

The tokenizer used for prelude (and flagship?) does not include <think> </think>, et.c.
The goal of this task is to convert the tokenizer used for prelude (and flagship) to support whatever
special tokens that is used in the template that we will use during post-training (Qwen3).

Deliverable scope
A post-training adapted tokenizer with special tokens include.d.

Dependencies
None

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.