apache / apache/hop

[Feature Request]: Apache HOP Plugin for Taxonomy-Based Text Extraction and Labelling

Open
#4,418 0 comments 0 reactions 0 assignees View on GitHub
awaiting triage P3 Transforms
Dominant language
Java
Stars
1.5k
Forks
476
Avg merge
18h 32m
Merged PRs (30d)
216

Description

### What would you like to happen?

An Apache HOP plugin that utilises a predefined taxonomy (a categorisation scheme) and automatically extracts and classifies snippets of text around key terms. The plugin would capture a specified number of words around the term, rounded to complete sentences, and label the text according to the taxonomy.

Use Case:
The plugin will be useful for users who need to focus on specific sections of large documents, such as known bias terms, industry-specific terms, product names, or other key phrases. Instead of manually reviewing entire documents, the plugin automatically extracts and labels relevant text segments. This plugin will facilitate semi-supervised learning by using preliminary labelled data to guide the analysis of unlabelled data.

### Issue Priority

Priority: 3

### Issue Component

Component: Transforms

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are named. Start by reviewing Apache Hop’s existing transform and plugin conventions, then clarify how the predefined taxonomy, context-word count, sentence boundaries, and labels should be configured; done means those requirements are implemented as a usable transform.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.