Jordan-Hall / Jordan-Hall/browser

[P1][EVAL-03] Outcome benchmarks and anti-wrapper tests

Open
#103 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Programme: #1
Epic: #33

## Objective
Measure whether the product actually delivers a persistent user-owned computing experience that outperforms ordinary browsing/chat assistance on verified outcomes—not merely a better-looking agent demo.

## Scope
- Scenario suites for research, shopping comparison, synthesized reading, social, coding, files, browser automation and native desktop tasks.
- Compare ordinary browsing, chat-based assistance using the same underlying tools/models, and persistent workspace mode.
- Anti-wrapper tests: close chat, restart app, disable inference, change provider/model and continue using the workspace.
- Repeated trials for nondeterministic tasks with confidence intervals and failure taxonomy.
- Human review for relevance, usability, evidence quality and corrective intervention burden.
- Metrics: verified completion, elapsed time, interventions, forced Original-mode handoffs, source/evidence quality, user understanding of authority, latency/cost/energy.
- Held-out private scenarios plus staged low-risk production canaries.

## Evaluation rules
- The model cannot be the sole grader of its own completion.
- Provider/model comparisons use pinned versions/configurations and disclose them.
- Do not use time-spent-in-product as the north-star metric; forced source inspection and healthy provenance use must not be penalized.

## Acceptance criteria
- [ ] Workspace remains directly usable after chat closure/restart with inference disabled for supported deterministic interactions.
- [ ] Provider switch preserves GoalContract, workspace, evidence and artifacts for declared scenarios.
- [ ] Benchmarks report verified completion separately from attempted/self-reported completion.
- [ ] Nondeterministic results include repeated trials, confidence intervals, interventions and failure classes.
- [ ] Comparisons use the same underlying tools/models where the goal is to isolate the interaction/runtime advantage.
- [ ] Release dashboards expose cost, latency, energy and correctness together rather than optimizing a single benchmark score.

## Dependencies
- EVAL-01
- WS-01
- WS-02

**First phase:** P1
**Maturity target:** P7 (continuous)
**Owner:** security-evaluation-release

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the EVAL-01, WS-01 and WS-02 dependencies and defining which declared scenarios they support. Establish the benchmark and human-review entry points for deterministic and nondeterministic trials. Done means the acceptance criteria are measured separately, with repeated-trial confidence intervals, failure classes, provenance checks, and combined cost, latency, energy and correctness reporting.

Written by the indexing model from the issue text.

Assessment

Domain
testing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.