arrayfire / arrayfire/arrayfire-rust

[BUG] device_info() buffers are 64 bytes; CUDA backend writes 257, corrupting the stack

Abierto Apto para principiantes
#384 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Bug
Lenguaje dominante
Rust
Estrellas
827
Forks
59
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

Description
===========

`device_info()` allocates the buffer sizes recommended by the ArrayFire documentation, but the CUDA backend writes up to 257 bytes into the first one. On the CUDA backend this corrupts up to 193 bytes of the caller's stack on every call.

This is a plausible root cause for several long-standing crash reports here: #106, #285, #311.

[`src/core/device.rs:109-121`](https://github.com/arrayfire/arrayfire-rust/blob/master/src/core/device.rs#L109-L121):

```rust
pub fn device_info() -> (String, String, String, String) {
let mut name: [c_char; 64] = [0; 64];
let mut platform: [c_char; 10] = [0; 10];
let mut toolkit: [c_char; 64] = [0; 64];
let mut compute: [c_char; 10] = [0; 10];
unsafe {
let err_val = af_device_info(
&mut name[0], &mut platform[0], &mut toolkit[0], &mut compute[0],
);
```

These sizes match ArrayFire's documented contract exactly (`docs/details/device.dox:10-16`, "Recommended minimum size is 64 / 10 / 64 / 10" across the four params), and `af_device_info` takes no length arguments, so there is nothing else to go on.

The CUDA backend does not honour it (`src/backend/cuda/platform.cpp:290-307`, identical from 3.8.0 through master):

```cpp
snprintf(d_name, 256, "%s", dev.name);

// Sanitize input
for (int i = 0; i < 256; i++) {
if (d_name[i] == ' ') { // reads d_name[0..256]
if (d_name[i + 1] == 0 || d_name[i + 1] == ' ') {
d_name[i] = 0; // writes d_name[0..255]
} else {
d_name[i] = '_';
}
}
}
```

The sanitize loop does not stop at the NUL terminator, so it runs the full 256 iterations regardless of the actual device-name length, writing `0x00` or `'_'` wherever it reads a `0x20` byte. The overflow therefore is not conditional on having a long GPU name — it happens on every call.

The CPU and OpenCL backends are correct here (`snprintf(..., 64, ...)`, and OpenCL bounds its sanitize loop at `i < 31`), so **this affects the CUDA backend only** — which matches the reported pattern of crashes that disappear when users switch to CPU or OpenCL.

I've filed the backend-side bug upstream as [arrayfire/arrayfire#](https://github.com/arrayfire/arrayfire/issues/3712).

Impact
--------

Writing ~193 bytes past a stack array clobbers the other three buffers, spilled registers, the `/GS` cookie and the return address. On Windows that surfaces as `STATUS_ACCESS_VIOLATION` (0xC0000005) or `STATUS_STACK_BUFFER_OVERRUN` (0xC0000409).

The corruption is deterministic but the *crash* is intermittent, because whether it is fatal depends on which stack bytes happen to contain `0x20` and on the frame layout rustc chose. That also gives a mechanism for #299 (crash only at `opt-level = 3`): more aggressive inlining means less dead stack padding to absorb the overflow.

`examples/helloworld.rs:9` calls `device_info()`, so this is in the first program a new user runs.

Reproducible Code and/or Steps
------------------------------

see above

System Information
------------------

general code issue, system independent

Checklist
---------

- [x] Using the latest available ArrayFire release
- [x] GPU drivers are up to date

Suggested fix
-------------

Independent of any upstream change, and safe against every 3.8.x:

```rust
let mut name: [c_char; 1024] = [0; 1024];
```

(or at least 257)

Happy to open a PR if that's useful.

Disclaimer
----------

Found by Claude Opus 5 while investigating intermittent 0xC0000005 errors on Windows/CUDA. I checked this manually and it seems like a real bug.

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Start in src/core/device.rs:109-121 and inspect the buffers passed by device_info(); compare them with the documented contract and the CUDA behavior described in src/backend/cuda/platform.cpp:290-307. Run examples/helloworld.rs with the CUDA backend to confirm the call is safe, and consider the issue complete when device_info() no longer permits the CUDA backend to overwrite its caller's stack.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
rust
Área
backend
Tipo de issue
Error
Dificultad
2/5
Tiempo estimado
1-3 horas
Estado de actividad
Tranquilo
Claridad
Bien especificado
Aptitud para principiantes
76/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.