Photorealistic derivative dataset based on olmOCR-mix-1025
- Linguagem predominante
- Python
- Estrelas
- 19.5k
- Forks
- 1.6k
- Métricas de merge de PRs
- Nenhum PR com merge em 30d
Descrição
Hi allenai team,
I’ve created a photorealistic version of olmOCR-mix-1025:
[olmOCR-mix-1025-Photoreal](https://huggingface.co/datasets/AlroWilde/olmOCR-mix-1025-Photoreal)
The pages are enhanced instead of clean digital PDFs, while keeping the original filenames so the existing annotations remain usable.
I hope this can help improve robustness in future olmOCR models.
If anyone needs similar datasets (specific effects or application scenarios), feel free to use it.
Happy to discuss if useful.
Guia de contribuição
Direção de pesquisa
Review the linked Hugging Face dataset and compare its photorealistic pages and preserved filenames with olmOCR-mix-1025. The issue does not name a repository file, test, entry point, or acceptance criterion; first clarify whether maintainers want evaluation, documentation, or integration, and consider the work complete only once that scope and success measure are agreed.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Stack de tecnologia
- python
- Domínio
- data, machine-learning
- Tipo de issue
- Funcionalidade
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Status de atividade
- Pouca atividade
- Clareza
- Precisa de esclarecimento
- Facilidade para iniciantes
- 25/100