Imageomics / Imageomics/Collaborative-distributed-science-guide
Consider reframing HF Dataset Upload Guide as a more general HF Upload Guide
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
[...] The other remaining point I noticed when reviewing this is that it only talks about dataset repos on HF. But the guidance (or at least most of it?) applies just as much to model repos.
Presumably we don't want to have near-duplicates of this for dataset and model repo guidance. Would the idea be to create a separate model repo guidance that by and large refers to this page asking the reader to replace the concept of "dataset" with "model"? Or would it be better to have a single page that in the (hopefully very few) places where it matters distinguishes between dataset repo and model repo type? (One pending project candidate that needs this guidance in fact needs it for a model repo.)
_Originally posted by @hlapp in https://github.com/Imageomics/Collaborative-distributed-science-guide/issues/65#issuecomment-4327890173_
It could probably be refactored to a general HF upload guide, especially since I already referenced a model vs dataset vs space distinction. I think models are generally more standardized, so it should be simple enough to mostly point to the docs. The key points to note there are about different checkpoints getting their own repositories and then generating a collection.
This last point will take a bit more consideration for implementation, so we have deferred this change to a new PR based on this issue.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the existing HF Dataset Upload Guide and review the discussion in issue 65. Reframe the guidance for HF uploads generally, distinguishing dataset, model, and space repositories where needed; cover separate model checkpoints and generating a collection, then verify that the result avoids duplicate dataset and model guidance.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 58/100