ageron / ageron/handson-ml2

highly biased training/test split? (Chapter 2 - page 51)

Ouverte
#344 1 commentaire 0 réactions 1 personne assignée Réclamée par @ageron Voir sur GitHub
Langage dominant
Jupyter Notebook
Étoiles
30k
Forks
13.1k
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

On page 53, the author mentions

> [...] you can try to use the most stable features to build a unique identifier.

and then proceeds to build an id based on the latitude and longitude. However, several instances in the dataset have the same latitude and longitude and hence the same identifier, and therefore the same hash. Maybe I'm missing something, but doesn't this introduce a very strong algorithmic bias (if that's the right term) in the training set selection, in that instances with the same (latitude, longitude) will either always get placed in the same set (whether training or test)? Shouldn't we be using more features to compute a unique identifier?

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.