IQSS / IQSS/dataverse

Develop external tool infrastructure for running cluster computing jobs from Dataverse

Open
#10,386 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Epic: Compute on Data
Dominant language
Java
Stars
1.1k
Forks
564
Avg merge
2d 2h
Merged PRs (30d)
29

Description

(we may come up with a better title for the issue; also, we may move it to a different repo - since most of this will need to be developed outside of the main Dataverse source tree - ?)

This is to continue/build on top of the proof of concept demo that we put together at MOC. The working assumption is that we'll be able to continue using the MOC facilities for this.

This next phase of the effort is to develop the one step that was skipped in the PoC configuration: an intermediate service sitting between the Dataverse and the actual computing nodes. It is similar in function to the redirecting script we have in place for the Binder service. Except we want it to do much more:

  • Rather than simply redirect to another running node, this service will allow a user to specify the parameters for the computing resources they need and spin up an OpenShift pod (?) to run their computations;
  • The resources will be requested and allocated from the user's own budget using their cluster account;
  • We will need to develop better/cleaner library code for obtaining local storage access points for the files in the dataset on the Dataverse instance ("local" should probably include support both for s3, as in the demo setup, and block access). There was some discussion of submitting this code to be included in pyDataverse.
  • The plan is to start with Python - for jobs similar to the notebook used in the demo presentation - with support for more languages (R is the next logical choice) going forward.

We will likely be adding more granular sub-components and opening child issues for such individual tasks.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the MOC proof-of-concept configuration, the Binder redirecting script, and the notebook used in the demo. Map the proposed intermediate service, cluster-resource and user-budget requirements, and dataset storage access for S3 and block storage; done means the scope and child tasks for Python job support are defined well enough to implement.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, cloud, infrastructure
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.