DesignMatrix should have to_dataframe() method
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 990
- Forks
- 106
- Avg merge
- 7d 34m
- Merged PRs (30d)
- 1
Description
This would be useful, for example, when I really want to be able to use a design matrix as both a raw numpy array and a pandas dataframe.
I suppose I could specify return_type="dataframe" and then get the numpy array from df.values, and it's also not hard to build the dataframe from scratch, but this would be particularly handy for interactive use, where it would provide a useful shortcut (e.g., X.to_dataframe().plot() or X.to_dataframe().head()).
To do this right, the new method would be factored out of build_design_matrices. Roughly speaking, it would look like this:
def to_dataframe(self):
if not have_pandas:
raise PatsyError("pandas.DataFrame was requested, but "
"pandas is not installed")
di = self.design_info
df = pandas.DataFrame(self, columns=di.column_names,
index=di.pandas_index)
df.design_info = di
return df
The main design change would be that DesignInfo (or DesignMatrix) would need to gain a pandas_index attribute, which would keep track of any index from the original data.
If this seems reasonable, I could put together a pull request.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading build_design_matrices and the DesignMatrix and DesignInfo definitions to understand how matrix data and original pandas indexes are currently handled. The work is complete when DesignMatrix exposes the requested dataframe conversion, preserves the relevant index and column metadata, and handles missing pandas as specified in the issue.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- numpy, pandas, python
- Domain
- data
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100