OpenHands / OpenHands/benchmarks

Difference in models results

Open
#707 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
124
Forks
90
Avg merge
1d 6h
Merged PRs (30d)
1

Description

Please see

  | private endpoint | public endpoint
Issue resolution: | 71.40% | 74.60%
Testing: | 67.00% | 70.40%
Information Gathering: | 63.60% | 74.50%
Greenfield development: | 50.00% | 25.00%

Both endpoints contain the same model Kimi-k2.6 but the private endpoint si consistently returning lower results.

Let's compare conversation and logs to determine why:

Private Issue resolution:"full_archive": "https://results.eval.all-hands.dev/swebench/litellm_proxy-accounts-graham-openhands-deployments-mghcd1dc/25200072711/results.tar.gz",

Public issue resolution "full_archive": "https://results.eval.all-hands.dev/swebench/litellm_proxy-moonshot-kimi-k2-6/25007210109/results.tar.gz",

Private Testing "full_archive": "https://results.eval.all-hands.dev/swtbench/litellm_proxy-accounts-graham-openhands-deployments-mghcd1dc/25328867381/results.tar.gz",
Public Testing "full_archive": "https://results.eval.all-hands.dev/swtbench/litellm_proxy-moonshot-kimi-k2-6/24901879531/results.tar.gz",

Private Information Gathering "full_archive": "https://results.eval.all-hands.dev/gaia/litellm_proxy-accounts-graham-openhands-deployments-mghcd1dc/25223608404/results.tar.gz",
Public Information Gathering "full_archive": "https://results.eval.all-hands.dev/gaia/litellm_proxy-moonshot-kimi-k2-6/25710749383/results.tar.gz",

Private Greenfield "full_archive": "https://results.eval.all-hands.dev/commit0/litellm_proxy-accounts-graham-openhands-deployments-mghcd1dc/25294380732/results.tar.gz",
Public Greenfield "full_archive": "https://results.eval.all-hands.dev/commit0/litellm_proxy-moonshot-kimi-k2-6/25710683155/results.tar.gz",

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by downloading and comparing the private and public result archives linked in the issue for Issue Resolution, Testing, Information Gathering, and Greenfield runs. Compare their conversations and logs to identify why the same Kimi-k2.6 model produces different scores, then document the cause and any corrective action.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.