bytedance / bytedance/UI-TARS

Is Android World evaluated the same way as OS World SHOWED IN osworld.py?

Open
#143 4 comments 8 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
11.5k
Forks
877
PR merge metrics
No merged PRs in 30d

Description

Is Android World evaluated the same way as OS World?

I have tested UI-TARS-7B-DPO and UI-TARS-72B-DPO on Android World, and they scored 26.7 and 35.7, respectively, compared with 33 and 46.6 in the paper.

my test tips:
i. i use Huawei Ascend NPU
ii. 4 historical images + 1 current image (all uncompressed)
iii. All historical actions and thoughts

Moreover, i used the same way to test UI-TARS-1.5-7B, the score is only 16.9. Then I used two tricks to get the score to 28.4(ALSO FAR FROM THE LASTED 64.2 SCORE!!!):
i. Modify the prompt to srocll down to guide the search for the app list.
ii. Turn off the gear logo at the main page header.

Is this normal?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading osworld.py and comparing its evaluation path with the Android World procedure described in the issue and the paper scores. Check the effects of the listed image history, prompt guidance, and gear-logo settings; the work is done when the score discrepancy and expected evaluation method are explained or documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.