Jordan-Hall / Jordan-Hall/browser
[P1][EVAL-03] Outcome benchmarks and anti-wrapper tests
- Dominant language
- No language data
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Programme: #1
Epic: #33
## Objective
Measure whether the product actually delivers a persistent user-owned computing experience that outperforms ordinary browsing/chat assistance on verified outcomes—not merely a better-looking agent demo.
## Scope
- Scenario suites for research, shopping comparison, synthesized reading, social, coding, files, browser automation and native desktop tasks.
- Compare ordinary browsing, chat-based assistance using the same underlying tools/models, and persistent workspace mode.
- Anti-wrapper tests: close chat, restart app, disable inference, change provider/model and continue using the workspace.
- Repeated trials for nondeterministic tasks with confidence intervals and failure taxonomy.
- Human review for relevance, usability, evidence quality and corrective intervention burden.
- Metrics: verified completion, elapsed time, interventions, forced Original-mode handoffs, source/evidence quality, user understanding of authority, latency/cost/energy.
- Held-out private scenarios plus staged low-risk production canaries.
## Evaluation rules
- The model cannot be the sole grader of its own completion.
- Provider/model comparisons use pinned versions/configurations and disclose them.
- Do not use time-spent-in-product as the north-star metric; forced source inspection and healthy provenance use must not be penalized.
## Acceptance criteria
- [ ] Workspace remains directly usable after chat closure/restart with inference disabled for supported deterministic interactions.
- [ ] Provider switch preserves GoalContract, workspace, evidence and artifacts for declared scenarios.
- [ ] Benchmarks report verified completion separately from attempted/self-reported completion.
- [ ] Nondeterministic results include repeated trials, confidence intervals, interventions and failure classes.
- [ ] Comparisons use the same underlying tools/models where the goal is to isolate the interaction/runtime advantage.
- [ ] Release dashboards expose cost, latency, energy and correctness together rather than optimizing a single benchmark score.
## Dependencies
- EVAL-01
- WS-01
- WS-02
**First phase:** P1
**Maturity target:** P7 (continuous)
**Owner:** security-evaluation-release
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading the EVAL-01, WS-01 and WS-02 dependencies and defining which declared scenarios they support. Establish the benchmark and human-review entry points for deterministic and nondeterministic trials. Done means the acceptance criteria are measured separately, with repeated-trial confidence intervals, failure classes, provenance checks, and combined cost, latency, energy and correctness reporting.
Written by the indexing model from the issue text.
Assessment
- Domain
- testing
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100