EducationalTestingService / EducationalTestingService/factor_analyzer

get_factor_variance() returns ndarray which is not ordered by variance.

Open
#132 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
6
Forks
1
PR merge metrics
No merged PRs in 30d

Description

When i print the variance of the factors with get_factor_variance(), the variance is not monotonically decreasing. In my understanding each additional factor should explain less and less variance. The variance of my 15 factors looks like this (notice the values marked in bold):
[12.69, 5.32, 2.7, 2.6, 2.26, 2.08, 1.76, **2.54**, **2.4**, 1.69, 1.49, **2.06**, 1.16, 1.15, **1.45**]

````
n_factors = 15
fa = FactorAnalyzer(n_factors, rotation="varimax", method='principal', use_smc=True)
fa.fit(X)
print(pd.DataFrame(fa.get_factor_variance(),index=['Variance (sum of squared loadings)','Proportional Var','Cumulative Var']))
````
I cannot provide you with the data X

Expected behavior: the output should be sorted by decreasing variance

- OS: macOS
- Python: 3.11
- Versions for `factor_analyzer` / `numpy` / `scipy` / `pandas`: newest

Contributor guide

No contributing guide indexed for this repository

Research direction

Start at the get_factor_variance() entry point and inspect how factor variance is calculated and ordered. Use the issue's FactorAnalyzer example with a local dataset to reproduce the non-monotonic output; done means the returned variance values decrease while the proportional and cumulative rows remain consistent.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.