microsoft / microsoft/onnxruntime

[QNN EP] ep.context_embed_mode=1 generates an unloadable EPContext model for multi-partition graphs (embed_mode=0 works)

Open
#31,977 1 comment 0 reactions 0 assignees View on GitHub
ep:QNN model:transformer
Dominant language
C++
Stars
21.9k
Forks
4.2k
Avg merge
4d 11h
Merged PRs (30d)
184

Description

### Describe the issue

When generating an EPContext model for the QNN EP with `ep.context_embed_mode=1` (embedded binary), the resulting model **cannot be loaded** if QNN partitions the graph into more than one `EPContext` node.

Only the **first** `EPContext` node receives the `ep_cache_context` attribute. The remaining partitions are emitted with no `ep_cache_context` at all. On load, ORT reads an empty string, treats it as a file path, and fails:

```
onnxruntime.capi.onnxruntime_pybind11_state.InvalidGraph: [ONNXRuntimeError] : 10 : INVALID_GRAPH :
Failed to load from EpContext model. qnn_backend_manager.cc:945
onnxruntime::qnn::QnnBackendManager::GetFileSizeIfValid Context binary does not exist:
```

Note the **empty path after the colon** — this presents as a missing-file problem but is actually a missing-attribute problem.

`ep.context_embed_mode=0` (external `.bin`) works correctly on the same model, so this appears specific to the embedded path.

This is distinct from #18354, which reported multi-partition context cache being unsupported entirely; that was fixed and confirmed. Here `embed_mode=0` works and only `embed_mode=1` is broken.

Practical impact: users hitting this silently fall back to recompiling the graph on **every process start**. In our case that is 1.4–1.9 s per start versus 257–484 ms when caching works, and there is no warning that caching failed.

### To reproduce

Model: `sentence-transformers/all-MiniLM-L6-v2`, ONNX, input shapes fixed to `[1, 128]`, statically quantised to QDQ a16w8 (`QUInt16` activations / `QUInt8` weights). QNN splits this into **7 partitions**.

**Generate:**

```python
import onnxruntime as ort

so = ort.SessionOptions()
so.add_session_config_entry("ep.context_enable", "1")
so.add_session_config_entry("ep.context_file_path", str(ctx_path))
so.add_session_config_entry("ep.context_embed_mode", "1") # <-- the problem

sess = ort.InferenceSession(
str(model), so,
providers=["QNNExecutionProvider"],
provider_options=[{"backend_path": "QnnHtp.dll"}],
)
```

**Then load the generated `ctx_path`** in a fresh session with the QNN EP — it fails with the error above.

**Inspecting the generated model** shows the cause:

```
node QNNExecutionProvider_QNN_..._7_0 ep_cache_context = EMBEDDED blob, 37.12 MB
node QNNExecutionProvider_QNN_..._8_1 (no ep_cache_context attribute)
node QNNExecutionProvider_QNN_..._9_2 (no ep_cache_context attribute)
node QNNExecutionProvider_QNN_..._10_3 (no ep_cache_context attribute)
node QNNExecutionProvider_QNN_..._11_4 (no ep_cache_context attribute)
node QNNExecutionProvider_QNN_..._12_5 (no ep_cache_context attribute)
node QNNExecutionProvider_QNN_..._13_6 (no ep_cache_context attribute)
```

All 7 nodes carry `embed_mode = 1`, `source = QNNExecutionProvider`, `partition_name`, and `max_size = 0`. Only node 0 has `ep_cache_context`.

Dumped with:

```python
import onnx
m = onnx.load(str(ctx_path), load_external_data=False)
for node in m.graph.node:
if node.op_type == "EPContext":
attrs = {a.name for a in node.attribute}
print(node.name, "has ep_cache_context:", "ep_cache_context" in attrs)
```

### Expected behaviour

Either every `EPContext` node should receive its own embedded `ep_cache_context`, or generation should fail loudly rather than emitting a model that cannot be loaded.

### Workaround

Use external mode:

```python
so.add_session_config_entry("ep.context_embed_mode", "0")
```

This writes `_qnn.bin` next to the context `.onnx`. Interestingly, the same six partitions still lack `ep_cache_context` in this mode, but loading succeeds — so the load path tolerates the missing attribute while the embedded path does not.

Measured on the model above:

| | session creation |
|---|---|
| uncached compile | 1414–1932 ms |
| cached load (`embed_mode=0`) | **257–484 ms** |

Output is bit-identical between cached and uncached sessions (embedding L2 norm 41.884804 in both).

### Urgency

Not blocking — `embed_mode=0` is a complete workaround. Reporting because the failure mode is silent and misleading: the error names a missing file with an empty path, which sends you looking for filesystem or packaging problems rather than a code-generation gap.

### Platform

Windows

### OS Version

Windows 11 26100 (ARM64), Snapdragon X Elite X1E-80-100, Hexagon NPU driver 30.0.220.3000

### ONNX Runtime Installation

Released Package

### ONNX Runtime Version or Build

onnxruntime-qnn 1.24.4 (ORT 1.24.4), Python 3.13.15, ARM64

### ONNX Runtime API

Python

### Architecture

ARM64

### Execution Provider

Other / Unknown

### Execution Provider Library Version

QNN / QAIRT bundled with onnxruntime-qnn 1.24.4, `QnnHtp.dll` backend

Contributor guide

Open the contributing guide

Research direction

Start by tracing QNN EP context generation and loading for ep.context_embed_mode=1, using the multi-partition EPContext model described in the reproduction. Compare embedded and external modes while inspecting each EPContext node's attributes. Done means every partition loads successfully with its embedded context, or generation fails clearly instead of producing an unloadable model.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
backend, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
65/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.