python / python/cpython

Undocumented differences between defaultdict and dict behavior

未关闭
#124,875 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

docs
主要语言
Python
星标
77.2k
派生
35.9k
PR 合并指标
PR 指标待抓取

描述

Documentation

defaultdict seems to call __getitem__ whenever __setitem__ is called (regardless of if the item was already present), whereas regular dict does not call __getitem__ when __setitem__ is called. The documentation for defaultdict says that defaultdict and dict are basically identical, except in a few narrow cases.

defaultdict is a subclass of the built-in dict class. It overrides one method and adds one writable instance variable. The remaining functionality is the same as for the dict class and is not documented here.

But nothing is mentioned in the docs about this difference in behavior of calling __getitem__ / __setitem__

This comes up when making a child class of either of them if you want to have a preprocessing step that operates on keys before they are used to index into the dictionary, e.g.

from collections import defaultdict

class Item: ...

def preprocess_item(item: Item) -> str:
    return f'::{item.__class__.__name__}@{hex(id(item))}'

class PreprocessingDefaultDict(defaultdict):
    def __getitem__(self, item: Item):
        key = preprocess_item(item)
        return super().__getitem__(key)

    # def __setitem__(self, item: Item, value):
    #     key = preprocess_item(item)
    #     super().__setitem__(key, value)


class PreprocessingDict(dict):
    def __getitem__(self, item: Item):
        key = preprocess_item(item)
        return super().__getitem__(key)

    def __setitem__(self, key: Item, value):
        key = preprocess_item(key)
        super().__setitem__(key, value)


if __name__ == '__main__':
    item = Item()
    
    d1 = PreprocessingDefaultDict(dict)
    d1[item]['a'] = 10  # initial creation of dict at d1[item]
    d1[item]['a'] = 20  # updating already existing dict at d1[item] 
    print(dict(d1)) # wrap in dict so prints the same as d2

    d2 = PreprocessingDict()
    d2[item] = {}
    d2[item]['a'] = 10  # initial creation of dict at d2[item]
    d2[item]['a'] = 20  # updating already existing dict at d2[item]
    print(d2)

Which prints out something like:

{'::Item@0x7fd29b035280': {'a': 20}}
{'::Item@0x7fd29b035280': {'a': 20}}

In this example, I have a preprocessor function I'd like to run on all keys to convert them from objects into strings which can be used in the dictionary. It is not clear from the docs that you need to not override __setitem__ like I have commented out, because defaultdict will always call __getitem__ thus always running the preprocessor. If you override __setitem__ like I have commented out, you will preprocess the item twice, and end up with results like this:

{'::str@0x7fc55b78a930': {'a': 10}, '::str@0x7fc55b78a970': {'a': 20}}
{'::Item@0x7fc55b754890': {'a': 20}}

or this:

{'::str@0x7f3715686930': {'a': 20}}
{'::Item@0x7f3715650860': {'a': 20}}

(I believe the extra element happens because the string from preprocess_item may or may not allocate new memory given an identical input)

I'm not exactly sure what the underlying cause of this difference is. It doesn't seem to be related to the __missing__ method mentioned in the docs, because the behavior I mentioned happens for keys that are not present in the defaultdict as well as for those that are already present (and presumably wouldn't be calling __missing__).

python version

I ran my example in python 3.6 through 3.12, and observed the same behavior in all of them

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

从链接的 collections.defaultdict 文档开始,将其关于共享 dict 功能的表述与 issue 示例中展示的行为进行比较。验证在所报告的 Python 版本中 getitemsetitem 之间的交互,包括与 missing 的关系;当文档准确解释所观察到的差异及其对 subclassing 的影响时,即表示完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
documentation
Issue 类型
文档
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
45/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。