Segfault in `execute_ex`: `EX(run_time_cache)` is NULL for a trait-copied `private static` method (no opcache, no JIT)
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- C
- Star
- 40.4k
- Fork
- 8.2k
- Merge trung bình
- 2 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 103
Mô tả
Description
We are seeing a reproducible-in-aggregate segfault in the Zend VM on PHP 8.3.31
and 8.3.32 (Ubuntu 24.04, deb.sury.org builds, NTS). It occurs ~36 times/week on
one CI host under a Laravel application served by php -S (via
php artisan serve). Because the built-in server is a single process, each crash
takes the server down and every subsequent request is refused.
The engine executes an op_array whose run_time_cache is NULL, then performs a
run-time cache slot lookup against it.
Faulting instruction and registers, from a core dump (8.3.32):
=> execute_ex+15353: mov (%rdx,%rax,1),%rbx
rdx = 0x0 <- run-time cache base
rax = 0x8 <- cache slot offset
execute_ex+15357: test %rbx,%rbx
execute_ex+15360: je ...
Read directly from the core:
run_time_cache (execute_data + 0x40) = 0x0
Kernel log line (identical every time, same binary offset on every occurrence):
php8.3[3641252]: segfault at 8 ip ...2a3729 sp ... error 4 in php8.3[...+32f000]
error 4 = user-mode read of a non-present page; the faulting address is
literally 0 + 8, i.e. slot 1 of a NULL cache.
The function involved
Frame 0 is always the same PHP function, reached from different callers. Walking
execute_data out of the core (%r14; func at +0x18, prev_execute_data at
+0x30, function_name at +0x08, zend_string.val at +0x18):
#0 resolveTenantId <- always frame 0
#1 {closure} (an Eloquent global scope closure)
#2 {closure} -> callScope -> applyScopes
core A: ... -> get -> first -> resolveRouteBinding -> substituteImplicitBindings
core B: ... -> toBase -> __call -> currentAgencyHasStudio -> {closure}
resolveTenantId() is a private static method defined in a trait, and that
trait is used by ~68 classes — so the method exists as ~68 separate op_arrays,
each with its own ZEND_MAP_PTR run-time cache. It is called on a very hot path
(every scoped query). Its body is small:
private static function resolveTenantId(): int|string|null
{
if (! function_exists('app')) {
return null;
}
$resolver = app()->bound(TenantResolver::class)
? app(TenantResolver::class)
: null;
return $resolver?->currentTenantId();
}
Ruled out
- opcache / JIT —
opcache.enable_cli => Off. opcache.so is loaded as a
module but caches nothing in this SAPI; JIT therefore never engages. (This
distinguishes it from #22158, which is JIT-specific, and #9396, which requires
opcache.) PHP_CLI_SERVER_WORKERS— unset; confirmed absent from the core's
environment, so the experimental forking server is not involved.- pcov / xdebug or any other VM-hooking extension — not among the 65
extensions loaded in the crashed process. - Userland recursion into the same scope — the class queried by the resolver
is exempt from the global scope, so there is no re-entry loop. - PHP 8.3.33 — its NEWS contains only Date/GD/PGSQL/Phar entries; nothing
engine-related, so we do not expect it to change this.
Environment
- PHP 8.3.32 (
8.3.32-1+ubuntu24.04.1+deb.sury.org+1) and 8.3.31, both NTS, both
affected, on two separate hosts - Ubuntu 24.04
- SAPI:
cli-server(php -S) - opcache loaded,
opcache.enable_cli=Off, JIT inactive - 65 extensions loaded
What I have and what I don't
I have core dumps and can extract any additional structure from them on request
(op_array fields, the zend_function at execute_data+0x18, the opline, etc.).
I do not have a minimal reproducer, and I have tried. The crash is frequent
in aggregate (~5/day) but not deterministic per-request, which is consistent with
an allocation- or lifetime-related issue around the run-time cache rather than a
specific input.
Attempts, both run on the same affected binary and host, both negative:
- Structural mimic — one trait with a
private staticmethod calling
function_exists(), a container-ish helper and a nullsafe method call; 68
classes using the trait; each invoked through astatic functionclosure via
call_user_func(to mirror the global-scope indirection); served by
php -S. 400 requests, ~186k calls: no crash. - Same, scaled to 1500 trait-using classes + 1500 free functions to force
ZEND_MAP_PTRtable growth and reallocation across a long-lived
php -Sprocess. 500 requests: no crash.
So the trait-copy structure alone is not sufficient. Something else in the real
workload is required — plausibly overall arena/memory pressure, the volume of
classes autoloaded per request, or live PDO/session activity. I mention the
negative results explicitly so nobody re-runs the same experiment.
I have the core dumps and can extract any additional structure from them on
request, or run a diagnostic/assertion build against the live workload — that is
likely the fastest route to a fix, since the workload reproduces it reliably in
aggregate even though a reduced case does not.
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Bắt đầu từ lỗi execute_ex và kiểm tra execute_data, op_array và NULL run_time_cache được mô tả trong các core dump. Sử dụng bản build chẩn đoán hoặc assertion cho workload php -S bị ảnh hưởng, rồi so sánh trạng thái runtime-cache của phương thức static private được sao chép từ trait qua các lần gọi. Hoàn tất khi tái hiện được việc truy cập cache không hợp lệ và xác định được nguyên nhân đã được xác nhận ở cấp engine cùng với regression test hoặc fix.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- c, php
- Lĩnh vực
- backend, compilers
- Loại issue
- Lỗi
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức độ hoạt động
- Ít trao đổi
- Độ rõ ràng
- Cần làm rõ
- Mức phù hợp với người mới
- 25/100