joshsoftware / joshsoftware/hackathon-team-5
Final Testing and Presentation
Open
- Dominant language
- TypeScript
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
1. Screenshot Collection: Puppeteer captures the browser screen and sends it to the FastAPI backend.
2. Vision Model Prediction: The Llama Vision Model processes the screenshot and predicts: Coordinates (x, y) of the element. Action to be performed (e.g., click, scroll).
3. Action Execution: Puppeteer performs the predicted action in the browser.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.