OpenHands / OpenHands/benchmarks

Investigate persona wording in SWE-Bench prompt template

Open
#688 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
124
Forks
90
Avg merge
1d 6h
Merged PRs (30d)
1

Description

Summary

In the SWE-Bench user instruction template, the first user message starts with:

I have access to a python code repository in the directory /workspace/django/ . You can explore and modify files using the available tools. Consider the following issue description:

That wording appears to switch persona in a slightly odd way: the surrounding setup already frames the agent as the actor with repository/tool access, but the user message says "I have access".

Where this seems to come from

Current source appears to be:

  • benchmarks/swebench/prompts/default.j2

and the same pattern also seems to exist in a few related prompt templates.

Request

Please look into:

  1. whether this persona wording is intentional or accidental
  2. where exactly this phrasing originated and whether it is shared across benchmark families
  3. whether it has any measurable effect on model behavior

My guess is that it probably has no material effect, but I’ve been surprised by prompt phrasing details before, so it seems worth checking.

Why this is mildly concerning

Even if harmless, this may create a subtle mismatch between:

  • the agent/system prompt, which frames the model as the actor using tools
  • the benchmark/user prompt, which momentarily sounds like the user is the one with tool/repo access

That could matter for certain models or prompt variants, especially when doing system-prompt experiments.

This issue was created by an AI assistant (OpenHands) on behalf of the user.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with benchmarks/swebench/prompts/default.j2 and inspect the related prompt templates mentioned in the issue. Trace where the wording originated and compare its use across benchmark families, then check whether existing evaluation or experiment results can show a behavioral effect. Done means the intent, scope, evidence, and a clear recommendation are documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, testing-qa
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.