ageron / ageron/handson-ml

Couldnt understand the code in chapter-2 while separating test set

未关闭
#567 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Jupyter Notebook
星标
25.6k
派生
12.7k
PR 合并指标
PR 指标待抓取

描述

Hi Mr.Aurélien Géron,

In your book while separating the test set you have written.
def test_set_check(identifier, test_ratio, hash):
return hash(np.int64(identifier)).digest()[-1] < 256 * test_ratio
def split_train_test_by_id(data, test_ratio, id_column, hash=hashlib.md5):
ids = data[id_column]
in_test_set = ids.apply(lambda id_: test_set_check(id_, test_ratio, hash))
return data.loc[~in_test_set], data.loc[in_test_set]

Can you help me to understand how hash helps in separating the test set and avoid the problems mentioned before. In the second book you have used crc32 and the code is as following:
from zlib import crc32
def test_set_check(identifier, test_ratio):
return crc32(np.int64(identifier)) & 0xffffffff < test_ratio * 2**32
How does this equals to the above code?

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。