llnl / llnl/dmx-learn

Run benchmark review for post-fit and downstream-task behavior

Open
#97 0 comments 0 reactions 0 assignees View on GitHub
agentic-dmx
Dominant language
Jupyter Notebook
Stars
10
Forks
1
Avg merge
1m
Merged PRs (30d)
34

Description

Run a second focused manual review against the benchmark suite to validate
post-fit usage, posterior reweighting, and downstream-task execution behavior.

This issue is also review-oriented and should stay narrow.

Scope:

- run benchmark prompts through the current skill set
- focus only on post-fit usage and downstream-task behavior
- collect and categorize failures

Acceptance criteria:

- a review artifact exists
- post-fit and downstream-task failures are categorized clearly
- follow-up fixes can be opened as separate small issues if needed

Contributor guide

Open the contributing guide

Research direction

Start by running the benchmark prompts through the current skill set, focusing only on post-fit usage and downstream-task execution. Record and categorize failures, then produce a review artifact showing the categorized results and any follow-up fixes that should become separate issues.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, python
Domain
machine-learning, testing-qa
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
65/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.