firecrawl / firecrawl/llmstxt-generator

Feature Request: Add Token Count Estimation to Output

Open
#1 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
TypeScript
Stars
536
Forks
70
PR merge metrics
No merged PRs in 30d

Description

I'd like to propose adding an estimated token count to the generated output. This would help users know if their generated text fits within their LLM's context window limits.

**Proposed Feature:**
- Add an estimated token count at the beginning of both `llms.txt` and `llms-full.txt` files
- Display format could be something like:
`Estimated Tokens: 12,345`

**Why this would be useful:**
- Helps users immediately know if the generated text will fit their LLM's context window
- Prevents trial-and-error when loading large text files into LLMs
- Makes it easier to split content into appropriate chunk sizes if needed

**Implementation Suggestions:**
- Could use libraries like `tiktoken` or a simple character-based approximation
- Token count could be placed in a header section or metadata block at the start of the file

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.