Photorealistic derivative dataset based on olmOCR-mix-1025
Abierto
- Lenguaje dominante
- Python
- Estrellas
- 19.5k
- Forks
- 1.6k
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
Hi allenai team,
I’ve created a photorealistic version of olmOCR-mix-1025:
[olmOCR-mix-1025-Photoreal](https://huggingface.co/datasets/AlroWilde/olmOCR-mix-1025-Photoreal)
The pages are enhanced instead of clean digital PDFs, while keeping the original filenames so the existing annotations remain usable.
I hope this can help improve robustness in future olmOCR models.
If anyone needs similar datasets (specific effects or application scenarios), feel free to use it.
Happy to discuss if useful.
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.