gqcpm / gqcpm/wavewatch

webscraping

Open
#22 0 comments 0 reactions 0 assignees View on GitHub
enhancement help wanted
Dominant language
Python
Stars
1
Forks
1
PR merge metrics
No merged PRs in 30d

Description

Currently, this model utilizes a RAG model to get data from saved `knowledge_base` in the `data/` folder. There is currently only data for pleasure point that was manually inputted with multiple sources. I would like this process to be either
1. automated
or
2. manually inputted.

For efficiency sake, I would like this to be initially automated by the use of webscraping. This will get the necessary information for each beach.

Long term: aka if there could be designated locals, to create a singular ideal conditions md file rather than pulling data from multiple sources, it would be much more accurate.

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the RAG model's use of the saved knowledge_base in the data/ folder and the existing Pleasure Point data. Define which beach information should be collected, which sources are acceptable, and how manual input remains supported; done means the project has an agreed process for supplying data for each beach.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.