A dict watcher that follows the documented error protocol makes CPython re-enter Python mid-update
Nadie ha tomado este issue todavía.
- Lenguaje dominante
- Python
- Estrellas
- 77.2k
- Forks
- 35.9k
- Métricas de merge de PR
- Métricas de PR pendientes
Descripción
Crash report
What happened?
Doc/c-api/dict.rst tells a PyDict_WatchCallback two things:
The callback may inspect but must not modify dict; doing so could have unpredictable effects, including infinite recursion. Do not trigger Python code execution in the callback, as it could modify the dict as a side effect.
If the callback sets an exception, it must return
-1; this exception will be printed as an unraisable exception usingPyErr_WriteUnraisable.
A callback that does exactly the second thing causes CPython to do exactly the first thing on its behalf. _PyDict_SendEvent (Objects/dictobject.c:8309-8317):
if (cb && (cb(event, (PyObject*)mp, key, value) < 0)) {
PyErr_FormatUnraisable(
"Exception ignored in %s watcher callback for <dict at %p>",
dict_event_name(event), mp);
}
PyErr_FormatUnraisable runs sys.unraisablehook, which is settable from pure Python. So the "do not trigger Python code execution" requirement is discharged by the callback and then violated by the runtime — inside the notify window, while the dict is mid-update.
The docs do say the exception becomes an unraisable. What they don't say is that this happens synchronously, inside the operation, at a point where the calling code has already captured an entry index, an entry pointer, a bound, or a borrowed reference that it will use afterwards.
Thirteen of the fourteen notify sites hold stale state across that window
I drove all fourteen. Thirteen hold state across the notify and all thirteen produce an observable failure; twelve of those are memory corruption or a memory-invariant assertion. One site is safe.
| site | enclosing function | stale value | worst observed |
|---|---|---|---|
:1917 |
insert_combined_dict |
dk_usable bound |
ASan heap-buffer-overflow WRITE, "0 bytes after" a 1520-byte region; release: silent corruption, exit 0 |
:1997 |
_PyDict_InsertSplitValue |
ix, split-table-ness |
SIGSEGV at :1998, ma_values == NULL — 5/5 on each of four builds |
:2003 |
_PyDict_InsertSplitValue |
ix, old_value |
SIGSEGV at :2004 5/5 on four builds; embedded variant asserts in clear_freelist |
:2060 |
insertdict |
ix, old_value |
ASan heap-UAF READ at :2076; heap-buffer-overflow WRITE at :2068; _Py_NegativeRefcount 5/5 |
:2103 |
insert_to_emptydict |
unpublished newkeys |
ma_used desync puts a NULL in a list → SIGSEGV in list_sort_impl |
:3038 |
_PyDict_DelItem_KnownHash_LockHeld |
ix |
debug assert(hashpos >= 0); release Py_DECREF(NULL) |
:3083 |
delitemif_lock_held |
ix, old_value, hash |
ASan SEGV Py_DECREF(NULL); len() → NULL 5/5 release |
:3142 |
clear_lock_held |
keys pointer | ASan use-after-free at dictkeys_decref:496 |
:3307 |
_PyDict_Pop_KnownHash |
ix, old_value |
ASan heap-UAF READ at :3308; SIGSEGV 5/5 release |
:4234 |
dict_dict_merge |
the whole guard set | debug assert in clone_combined_dict_keys; silent data loss 5/5 — no memory corruption reached |
:5051 |
dict_popitem_impl (unicode) |
ep0 |
ASan UAF R+W; popitem() returns a tuple whose [1] is a raw C NULL |
:5066 |
dict_popitem_impl (general) |
ep0, i |
ASan heap-UAF READ at the site, :5067 |
:7510 |
store_instance_attr_lock_held |
values, dict, ix, old_value |
assert(i < size); ASan overflow WRITE + UAF READ; len(), list() and vars() all raise |
:3652 |
dict_dealloc |
(none) | safe — reads ma_values/ma_keys at :3656-3657, after the notify |
Also reached from these paths: delete_index_from_values:2943, whose search loop is bounded only by assert(i < size) — debug SIGABRT, and on plain release it silently loses a dict entry.
Four that are worth calling out individually
:1997/:2003 are in _PyDict_InsertSplitValue, which is PyAPI_FUNC and carries the comment // Exported for external JIT support. The hook that kills it is d.clear() after del obj has detached the dict, which makes clear_lock_held take its set_values(mp, NULL) branch; the split-table store then dereferences NULL. gdb confirms ma_values == 0x0, ma_keys == &empty_keys_struct, ix == 2.
:3083 breaks a promise the file makes in a comment. dictobject.c:3090-3094 tells callers the predicate → deletion sequence is atomic provided the predicate does not re-enter. is_dead_weakref does not re-enter; the notify does. Reachable as _weakref._remove_dead_weakref(dct, key).
:3307 hands freed memory back to Python. *result at :3312 is the dict's own reference, returned straight to the dict.pop() caller after the hook dropped it.
:7510 gives the sharpest Python-visible invariant break. delete_index_from_values does size-- on a size the hook zeroed, and values->size is a uint8_t, so it wraps to 255 while :7527 drives ma_used to -1. After a plain del o.attr, on release-gil-nojit, 5/5:
len(d) -> SystemError: <built-in function len> returned NULL without setting an exception
list(d) -> ValueError: length must be positive
vars(o) -> ValueError: length must be positive
The one I could not corrupt
:4234 (dict_dict_merge, PyDict_EVENT_CLONED) holds every guard established at :4219 and :4228-4232 across the notify, and those guards are genuinely unchecked afterwards — but I could not reach memory corruption in roughly nine hook variants (0/3 under ASan). The two guards whose falsification would give a type-confused dict are out of reach from Python: mp is a plain dict so a hook cannot make mp->ma_values non-NULL, and a combined other cannot be turned back into a split table. What it does give is a 5/5 debug assertion violation in clone_combined_dict_keys and 5/5 silent data loss on all four builds — d.update(src) discards everything the hook wrote to d, with no error. On release, that path also leaves d owning a heap keys object carrying _Py_DICT_IMMORTAL_INITIAL_REFCNT (memcpy'd from empty_keys_struct), which can never be freed and permanently falsifies clear_lock_held's assert(oldkeys->dk_refcnt == 1) at :3148.
Reproducer
The callback is CPython's own — dict_watch_callback_error in Modules/_testcapi/watchers.c:90-98, in its entirety:
static int
dict_watch_callback_error(PyDict_WatchEvent event, PyObject *dict,
PyObject *key, PyObject *new_value)
{
PyErr_SetString(PyExc_RuntimeError, "boom!");
return -1;
}
That is the documented protocol and nothing else: it sets an exception, returns -1, and runs no Python. The re-entry comes entirely from CPython's response to it.
import sys, _testcapi
d = {}
for i in range(200):
d["k%d" % i] = i
fired = []
def hook(unraisable):
if fired:
return
fired.append(1)
d.clear() # re-enter the dict being mutated
sys.unraisablehook = hook
wid = _testcapi.add_dict_watcher(1) # installs dict_watch_callback_error
_testcapi.watch_dict(wid, d)
del d["k100"] # notify -> -1 -> unraisable -> hook -> clear()
print("len now %d" % len(d))
# debug build -- exit 134 (SIGABRT)
python: Objects/dictobject.c:2963: void delitem_common(...): Assertion `hashpos >= 0' failed.
# release build -- exit 1
SystemError: <built-in function len> returned NULL without setting an exception
One per-site reproducer for each of the thirteen is attached, along with negative controls — for :2060, hooks that store immortal small ints or force a resize do not crash it, so the claim is not "any hook corrupts anything". That site needs heap values specifically: with small ints the stale Py_XDECREF(old_value) at :2076 is a no-op because they are immortal, which is why an earlier attempt at this site came back clean.
Registering a watcher requires the C API, so this is not reachable from pure Python alone — but sys.unraisablehook is pure Python, and any extension that installs a conforming watcher hands that reach to arbitrary Python code.
The one safe site, and why
dict_dealloc (:3650-3658) is the only notify site that survives this. It brackets the notify in _PyObject_ResurrectStart/End, but that bracket defends against resurrection — the comment at _PyDict_SendEvent:8309-8312 names that threat and only that threat. Its mutation-safety is incidental: it re-reads ma_keys/ma_values after the notify rather than before.
So the thing to propagate is the ordering, not the bracket.
Notes toward a fix
Three directions, not mutually exclusive:
- Re-validate after the notify at each of the thirteen sites — reload the index, entry pointer, or bound. Correct but repetitive, and easy to regress: this is thirteen independent chances to get it wrong again.
- Defer the unraisable until the dict operation completes, so the callback's error is still reported but not from inside the window.
- Guard the re-entry centrally in
format_unraisable_v(Python/errors.c:1737), which is the single choke point every unraisable site funnels through.
Given that thirteen of fourteen sites are affected, the central options look more attractive than the per-site one.
Objects/typeobject.c:1219 already carries the observation, added in gh-127266:
// Note that PyErr_FormatUnraisable is potentially re-entrant and the watcher
// callback might be too
dictobject.c has no equivalent, and its notify sites were never audited against it.
Related
- CPython's own reference watcher,
dict_watch_callback(watchers.c:31-70), violates both halves of the documented contract:PyUnicode_FromFormat("new:%S:%S", ...)callsPyObject_Str(forbidden bydict.rst:583-584), and it callsPyUnicode_FromFormat+PyList_Appendwith no save/restore (required bydict.rst:599-603). If the example cannot follow the contract, it is worth asking whether the contract is stateable. PyType_WatchCallback(Doc/c-api/type.rst:138-155) documents no error protocol at all, yettypeobject.c:1222treats< 0as "exception set" and reports it the same way.
Found with [cpython-review-toolkit]](https://github.com/devdanzin/cpython-review-toolkit); reduced, reproduced and drafted with AI assistance (Claude Code)
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Output from running 'python -VV' on the command line:
Python 3.16.0a0 (heads/main:a1d580430c8, Jul 18 2026, 20:21:17) [Clang 21.1.8 (6ubuntu1)]
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Línea de trabajo
Lee Objects/dictobject.c alrededor de los catorce puntos de llamada de _PyDict_SendEvent, junto con Python/errors.c:1737 y Modules/_testcapi/watchers.c:90-98. Ejecuta el reproductor proporcionado de sys.unraisablehook y compara el comportamiento de las compilaciones de depuración y de lanzamiento. Se considera terminado cuando el protocolo de errores documentado de los callbacks ya no permite una reentrada insegura durante las actualizaciones de diccionarios, y se han corregido la corrupción y los fallos de invariantes informados.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- c, python
- Área
- backend
- Tipo de issue
- Error
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Tranquilo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 35/100