mantoshkumar1 / mantoshkumar1/mantoshkumar1.github.io
Establish a human-rated quality benchmark for Ask Mantosh answers
Nobody has claimed this yet.
- Dominant language
- JavaScript
- Stars
- 1
- Forks
- 0
- Avg merge
- 13m
- Merged PRs (30d)
- 2
Description
Problem
The automated evaluator is extensive, but docs/SYSTEM_STATE.md explicitly notes that there is no human-rated response benchmark and that deterministic fixtures do not measure human preference or satisfaction.
Goal
Create a small, repeatable human-review benchmark focused on the site's actual visitor goals: recruiter assessment, engineering credibility, evidence quality, and useful navigation.
Scope
- Define a stable sample of representative questions across key visitor types.
- Create a concise scoring rubric for correctness, evidence grounding, usefulness, concision, and tone.
- Record baseline scores for the current production behavior.
- Define how benchmark results are compared when answer policy, retrieval, or model behavior changes.
- Avoid optimizing for generic chatbot engagement.
Acceptance criteria
- A version-controlled benchmark set and scoring rubric exist.
- Current production behavior has a recorded baseline.
- Reviewers can score answers consistently without needing private information.
- The rubric explicitly rewards evidence-backed portfolio usefulness rather than general-assistant breadth.
- Release documentation explains when human re-evaluation is required.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with docs/SYSTEM_STATE.md and inspect the current production behavior and existing evaluator or fixture coverage before defining the benchmark artifacts. Create a version-controlled question set and rubric, record a baseline, document comparison and reviewer guidance, and explain when human re-evaluation is required in release documentation.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation, testing-qa
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 52/100