microsoft / microsoft/OmniParser
Do we have finetuned models for more accurate icon classification?
Open
Nobody has claimed this yet.
- Dominant language
- Jupyter Notebook
- Stars
- 25.4k
- Forks
- 2.2k
- PR merge metrics
- No merged PRs in 30d
Description
The icon detection is supper great, however the icon classification ( parsed_content_list 'content' field is not so accurate )
If more accurate icon detection result we can feed to LLM directly to generate actions.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the icon classification output in parsed_content_list['content'] and the attached example. Determine whether a finetuned model for more accurate icon classification is available, and clarify what result would be accurate enough to feed directly to an LLM.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning
- Domain
- computer-vision, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100