Luminous-Dynamics / Luminous-Dynamics/symthaea
RQ-039: Cooperative reasoning and human-complementarity benchmark
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 9
- Forks
- 1
- Avg merge
- 15m
- Merged PRs (30d)
- 8
Description
Parent: #2536
Measure not only autonomous task performance but whether Symthaea improves human/expert outcomes: error catching, hypothesis diversity, calibration, time-to-solution, decision quality, and appropriate disagreement. Include cases where the best behavior is to defer to human/domain expertise.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the parent issue #2536 to understand the surrounding benchmark work. Define evaluation cases and measures for error catching, hypothesis diversity, calibration, time-to-solution, decision quality, appropriate disagreement, and deference to human expertise; done means these cooperative outcomes are measured alongside autonomous performance.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100