galaxyproject / galaxyproject/foundry-pattern
Redo Frontier Model Skill Generation Experiment from Blog Post
- Dominant language
- Astro
- Stars
- 1
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
@mvdbeek
>> The blind model spontaneously added more validity caveats than the original.
> I would not (automatically, without checking) call that rigor, in my experience it’s often just being a smart-ass, and sometimes in very embarassing ways. Have a frontier model suggest improvements to a UDT if you want a good laugh. (edited)
A fair point and the whole experiment should be redone with more rigor and more documentation.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue does not name any files, tests, or entry points. Start by locating the original experiment described in the blog post, then define and document a more rigorous methodology; done means the experiment has been rerun and its process and findings are documented.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100