joshsoftware / joshsoftware/hackathon-team-5
Puppeteer Task Execution
Open
- Dominant language
- TypeScript
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Implement Puppeteer logic to execute actions based on Llama Vision Model outputs.
1. Use Puppeteer to take screenshots of the browser for vision model input.
2. Send screenshots to the FastAPI backend and receive actions with coordinates.
3. Implement Puppeteer functions to perform actions (e.g., click, scroll) based on coordinates.
4. Test navigation workflows with multiple pages (e.g., clicking a "Next" button).
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.