joshsoftware / joshsoftware/hackathon-team-5

Puppeteer Task Execution

Open
#3 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
TypeScript
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Implement Puppeteer logic to execute actions based on Llama Vision Model outputs.

1. Use Puppeteer to take screenshots of the browser for vision model input.
2. Send screenshots to the FastAPI backend and receive actions with coordinates.
3. Implement Puppeteer functions to perform actions (e.g., click, scroll) based on coordinates.
4. Test navigation workflows with multiple pages (e.g., clicking a "Next" button).

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.