[Feature Request]: Apache HOP Plugin for Taxonomy-Based Text Extraction and Labelling
- Dominant language
- Java
- Stars
- 1.5k
- Forks
- 476
- Avg merge
- 18h 32m
- Merged PRs (30d)
- 216
Description
### What would you like to happen?
An Apache HOP plugin that utilises a predefined taxonomy (a categorisation scheme) and automatically extracts and classifies snippets of text around key terms. The plugin would capture a specified number of words around the term, rounded to complete sentences, and label the text according to the taxonomy.
Use Case:
The plugin will be useful for users who need to focus on specific sections of large documents, such as known bias terms, industry-specific terms, product names, or other key phrases. Instead of manually reviewing entire documents, the plugin automatically extracts and labels relevant text segments. This plugin will facilitate semi-supervised learning by using preliminary labelled data to guide the analysis of unlabelled data.
### Issue Priority
Priority: 3
### Issue Component
Component: Transforms
Contributor guide
Research direction
No files, tests, or entry points are named. Start by reviewing Apache Hop’s existing transform and plugin conventions, then clarify how the predefined taxonomy, context-word count, sentence boundaries, and labels should be configured; done means those requirements are implemented as a usable transform.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100