chaoss / chaoss/wg-data-science

[Project]: Sudden Archival

Open
#45 3 comments 1 reaction 0 assignees View on GitHub
project proposal
Dominant language
Jupyter Notebook
Stars
30
Forks
26
Avg merge
14h 38m
Merged PRs (30d)
2

Description

### Project Name (1 - 3 words)

Sudden Archival

### Description

Can we predict the likelihood of sudden archiving of project otherwise successful in the sense of being used by others with active development?

Projects that are suddenly and unexpectedly archived can create significant issues for the people / organizations who rely on those projects.

Data concerns:

- Can a large dataset of examples be generated? (probably need hundreds or thousands of examples)
- Can metadata metrics be found that seem to correlate with repositories where this occurs?
- How predictive are those metadata metrics in the sense of when they appear in a repository (looking back in time)?
- What metrics can we use to distinguish between projects that are archived suddenly and unexpectedly vs. those that are archived when they've reached the natural end of life?

This is often related to Elephant Factor, which refers to too much control by a single company. These negative events in a project’s life that are strongly associated with too much single company control, meaning they are different than what occurs in a project with a diverse community of contributors and maintainers.

### Related Links

_No response_

Note that we also have a [Project Scope Template doc](https://docs.google.com/document/d/13iLNDfqJ8nuwBGEyJuFutcT7KRNT6JwFrSlJN_5f4o4/edit) that you can use to think about the project details if you find it useful (not required).

### How would you like to be involved in this project?

I am interested in this project, but do not plan to work on it myself

### Additional Notes.

_No response_

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the issue's data concerns and the linked Project Scope Template doc, then define the dataset, repository metadata, and historical signals needed to study sudden versus natural archiving. Done would require an agreed research scope and evaluation criteria for predicting unexpected archival; no implementation files or tests are identified.

Written by the indexing model from the issue text.

Assessment

Domain
analytics, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.