microsoft / microsoft/OmniParser

Do we have finetuned models for more accurate icon classification?

Open
#98 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
25.4k
Forks
2.2k
PR merge metrics
No merged PRs in 30d

Description

The icon detection is supper great, however the icon classification ( parsed_content_list 'content' field is not so accurate )
If more accurate icon detection result we can feed to LLM directly to generate actions.
Image

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the icon classification output in parsed_content_list['content'] and the attached example. Determine whether a finetuned model for more accurate icon classification is available, and clarify what result would be accurate enough to feed directly to an LLM.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
computer-vision, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.