OpenHands / OpenHands/software-agent-sdk

[Harness Watch] P1 — diagnose confirmed trajectory differences

Open
#4,641 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement evaluation P1
Dominant language
Python
Stars
1.1k
Forks
539
Avg merge
1d 19h
Merged PRs (30d)
137

Description

Parent epic: #4627

Goal

After the deterministic comparison framework is stable, explain selected confirmed differences and propose one human-reviewable experiment. This does not block P0.

Reuse or refactor the existing propose-harness action instead of building a second diagnosis platform.

P1

  • Consume the paired report and only the leading outcome-discordant or resource-divergent trajectories.
  • Run read-only with no repository, branch, push, or issue-write capability.
  • Produce a structured hypothesis, alternatives, evidence and counterevidence with instance/event links, predicted observable change, and one single-variable experiment.
  • Ground quantitative statements in the deterministic report or tool output.
  • Stop at the proposal; a human decides whether to implement or run it.

The agent interprets evidence. It does not calculate statistics, decide significance, or modify repositories.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with parent epic #4627 and the existing propose-harness action, then review how the paired report exposes outcome-discordant or resource-divergent trajectories. Done means a read-only proposal containing hypotheses, alternatives, evidence and counterevidence with instance/event links, predicted observable changes, and one single-variable experiment, without calculating statistics or modifying repositories.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.