nf-core / nf-core/deepmodeloptim

[nf-tests] Assure reproducibility

Open
#55 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

high priority
Dominant language
Nextflow
Stars
31
Forks
14
PR merge metrics
No merged PRs in 30d

Description

random sampling
There are many random sampling methods, including random.sample, and other low level within library sampling.
Setting random.seed(0) at the very beginning of a script won't work.

set operations
Sets are unordered, consequently everything handled with sets are not gonna follow a certain order, and this is not controllable.
However, set operations are very efficient.

Alternatives?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are named. Begin by identifying every random-sampling and set-operation path covered by nf-tests, then compare reproducibility alternatives for each. Done means the project has a documented, agreed approach that produces repeatable results without removing the efficiency benefits of set operations.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.