On the Issue of Reproducing UITARS 1.5 on OSWorld
- Dominant language
- Python
- Stars
- 11.5k
- Forks
- 877
- PR merge metrics
- No merged PRs in 30d
Description
I attempted to reproduce the scores for UI-TARS-250705 (41.8%) using the script available at [run_multienv_uitars15_v1.py](https://github.com/xlang-ai/OSWorld/blob/main/run_multienv_uitars15_v1.py).
However, I encountered some bugs, for example, when constructing messages using the last 5 images during the reproduction process. After resolving these issues, I was unable to achieve the scores reported on the leaderboard (our reproduction 31.3%).
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with run_multienv_uitars15_v1.py and reproduce the reported UI-TARS-250705 run. Inspect the message construction involving the last 5 images and compare the reproduction result of 31.3% with the reported leaderboard score of 41.8%. Done means identifying and correcting the reproducibility issues sufficiently to explain or match the reported score.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100