google / google/turbinia

Create job for processing image with Apache Tika

Open
#457 0 comments 0 reactions 0 assignees View on GitHub
enhancement new-task
Dominant language
Python
Stars
794
Forks
172
PR merge metrics
No merged PRs in 30d

Description

https://github.com/apache/tika

"Apache Tika(TM) is a toolkit for detecting and extracting metadata and structured text content from various documents using existing parser libraries."

Contributor guide

Open the contributing guide

Research direction

Start with the linked Apache Tika project to understand its image metadata and text-extraction capabilities, then inspect Turbinia's existing job patterns and entry points. The issue does not name files, tests, expected outputs, or supported image formats, so those requirements must be clarified before implementation and validation are defined.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
security
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.