hackclubnmit / hackclubnmit/tensorflow-hacktoberfest-2020

Better data cleaning and pre processing for the text/caption generating model.

Open
#3 0 comments 0 reactions 0 assignees View on GitHub
enhancement good first issue hacktoberfest
Dominant language
Jupyter Notebook
Stars
2
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Filter out the text appropriately. Since the data is straight from instagram, it has noise.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the Jupyter Notebook code that prepares Instagram captions for the caption-generation model and inspect the current preprocessing steps. Define which Instagram noise should be filtered, apply the cleaning consistently to the training data, and verify that the resulting captions are suitable model inputs.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, python
Domain
data, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.