AOSSIE-Org / AOSSIE-Org/PictoPy
Feat: On-device OCR to extract and copy text from images
- Lingua principale
- Python
- Stelle
- 284
- Fork
- 680
- Merge medio
- 7g 5h
- PR unite (30g)
- 4
Descrizione
### Describe the feature
### Feature
Add a local, privacy-preserving Optical Character Recognition (OCR) feature that allows users to extract text from photos and copy it directly to their clipboard.
### Proposed Implementation & Architecture
#### 1. Backend (On-Device Inference via ONNX)
- **Model**: Integrate a lightweight, offline ONNX OCR model .
- Register the model under a new or optional feature tier so users can manage/download it locally.
- Implement a dedicated model class extending `ONNXSessionBase` for thread-safe inference.
- Index extracted text in SQLite so users can search their photo library by text found inside images.
#### 2. Frontend (Viewer & UX)
- **Viewer Action**: Add a "Scan Text" icon button .
- **Extraction Flow**:
- Triggering the action shows a subtle loading state while inference runs.
- A modal dialog displays the recognized text with line formatting preserved.
- A primary **"Copy to Clipboard"** button with immediate feedback (e.g. "Copied!").
- Handles cases where no text was found.
- Prompts the user to download/enable the OCR model from Settings if not yet installed.
### Add ScreenShots
Hi @rohan-pandeyy can i work in this feature.
### Record
- [x] I agree to follow this project's Code of Conduct
- [x] I want to work on this issue
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Start by locating the viewer action, ONNXSessionBase, model registration and feature-tier settings, and SQLite indexing paths. Review how the viewer should handle loading, modal text display, no-text results, clipboard feedback, and an unavailable model. Done means offline OCR extracts selectable text, supports copying, and indexes it for photo-library search.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python, sqlite
- Ambito
- computer-vision, databases, desktop, machine-learning
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Attiva
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 28/100