dotnet / dotnet/machinelearning-samples

PermutationFeatureImportance output is confusing.

Open
#850 0 comments 0 reactions 0 assignees View on GitHub
P3 question
Dominant language
PowerShell
Stars
4.7k
Forks
2.7k
Avg merge
2d 22h
Merged PRs (30d)
1

Description

I was running the following sample code

https://github.com/dotnet/machinelearning-samples/blob/7b6303a17be61294bd45df8e6e6b029695c0a8be/samples/csharp/end-to-end-apps/Model-Explainability/TaxiFarePrediction/TaxiFarePredictionConsoleApp/Program.cs

which gives the following PFI output

```
Feature PFI
VendorId | -0.105016
RateCode | -0.105016
PassengerCount | -0.378364
PassengerCount | -0.082747
TripTime | -0.105016
TripTime | -0.105016
TripDistance | -0.105016
TripDistance | -0.105016
PaymentType | -0.107889
FareAmount | -0.107410
Label | -0.096346
VendorIdEncoded | -0.105016
VendorIdEncoded | -0.105016
RateCodeEncoded | -0.105088
RateCodeEncoded | -0.231427
PaymentTypeEncoded | -0.691058
```

the Feature column has duplicated names and rows with the same name even have different PFIs, such as PassengerCount and RateCodeEncoded. I guess it might be due to OneHotEncoding and NormalizeMeanVariance but to end users that seems not making much sense. How should the PFI output be interpreted? Am I missing something obvious? Thanks!

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.