anomalyco / anomalyco/opencode

[FEATURE]:Screen vision (screenshots) + browser control tools for the agent

Open
#48,377 1 comment 0 reactions 1 assignee View on GitHub

@jlongster is already working on this.

Since Sep 10, 2026.

Dominant language
TypeScript
Stars
209k
Forks
27.5k
PR merge metrics
PR metrics pending

Description

Feature hasn't been suggested before.
  • I have verified this feature I'm about to request hasn't been suggested before.
Describe the enhancement you want to request

Hi! I'm a user working with an OpenCode agent (Muse Spark) on Windows.
My agent builds GUI programs (tkinter, turtle, pygame) and opens web pages
for me, but it is completely blind:

  • It can launch windows (Start-Process) but never sees them. When a program
    crashes on start, it only finds out via log files. It cannot verify layout,
    colors, or that a drawing looks right — it checks code with tests instead
    of looking with eyes.
  • It can open URLs but cannot click, scroll, fill forms, or read what is on
    the page. "Click that button for me" is impossible.
    Feature request:
  1. A screenshot tool (capture screen / active window) whose output is
    attached as an image the model can actually see.
  2. Optionally: basic browser control (go to URL, click, type, read content).
    Use cases: verifying GUI apps the agent just wrote, guiding users step by
    step through websites ("press the blue button"), visual debugging.
    Thank you!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.