hackclubnmit / hackclubnmit/tensorflow-hacktoberfest-2020
Better data cleaning and pre processing for the text/caption generating model.
Open
enhancement
good first issue
hacktoberfest
- Dominant language
- Jupyter Notebook
- Stars
- 2
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Filter out the text appropriately. Since the data is straight from instagram, it has noise.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the Jupyter Notebook code that prepares Instagram captions for the caption-generation model and inspect the current preprocessing steps. Define which Instagram noise should be filtered, apply the cleaning consistently to the training data, and verify that the resulting captions are suitable model inputs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- jupyter-notebook, python
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100