baidu / baidu/lac

想做大规模切词怎么加速,比如对20G文本切词

Open
#140 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
4k
Forks
592
PR merge metrics
No merged PRs in 30d

Description

想做大规模切词怎么加速,比如对20G文本切词

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points. Start by locating the current large-text tokenization path and measuring it against the 20G-text use case; a complete result would need a defined acceleration approach and benchmark criteria.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.