microsoft / microsoft/fara

Performance claims, data table don't make sense (to me)

Open
#34 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
6.2k
Forks
603
PR merge metrics
No merged PRs in 30d

Description

Fara-7B appears to under-perform compared to SoM Agents o3-mini and GPT-4o-0513, and mostly under-performs compared to computer use model OpenAI computer-use-preview. This doesn't seem to match the claim that Fara achieves state-of-the-art results [...] outperforming both comparable-sized models and larger systems. Am I missing something?


Image

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Review the performance claims and the data table referenced in the issue, focusing on the comparisons among Fara-7B, SoM Agents, GPT-4o-0513, and OpenAI computer-use-preview. Verify the underlying benchmark values and determine whether the table or the claim is inaccurate; done means the discrepancy is explained and the affected content is corrected or clarified.

Written by the indexing model from the issue text.

Assessment

Domain
documentation
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.