GoogleCloudPlatform / GoogleCloudPlatform/dcm2bq

Add ability to extract text from embedded images within PDF

Open
#24 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
JavaScript
Stars
9
Forks
5
PR merge metrics
No merged PRs in 30d

Description

This issue has no description.

Contributor guide

Open the contributing guide

Research direction

The issue has no body and names no files, tests, or entry points. Start by locating the existing PDF and text-extraction path in the JavaScript service; done should mean that text contained in an embedded PDF image is extracted in a reproducible case.

Written by the indexing model from the issue text.

Assessment

Tech stack
google-cloud, javascript
Domain
backend
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.