AlexsLemonade / AlexsLemonade/refinebio

Perform Principal Component Analysis on Agilent Two Color Dataset

Open
#211 7 comments 0 reactions 0 assignees View on GitHub
agilent backlog data science dataviz SCIENCE! (A.K.A. research question)
Dominant language
Python
Stars
135
Forks
21
PR merge metrics
No merged PRs in 30d

Description

### Context

We have [the data](https://s3.amazonaws.com/crunch-outputs/random_twocolor.tar.gz)!

Now we need to know: is the data good? Specifically - Can we clearly separate channel one from channel two? Even better - can we automatically classify an experiment as being a reference or loop experiment?

### Problem or idea

![screen shot 2018-04-19 at 1 32 00 pm](https://user-images.githubusercontent.com/139987/39008261-2c83524a-43d6-11e8-9d80-9f16ceb8d489.png)
![screen shot 2018-04-19 at 1 32 30 pm](https://user-images.githubusercontent.com/139987/39008270-320fd24c-43d6-11e8-9c86-720be77fab17.png)

### Solution or next step

Rich relearns how to use Pandas, scikit-learn and Jupyter.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by downloading the referenced random_twocolor dataset and reviewing the requested analysis with Pandas, scikit-learn, and Jupyter. Define the analysis output needed to determine whether the two channels separate and whether reference and loop experiments can be classified; the issue names no files or tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter, pandas, python, scikit-learn
Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.