firebase / firebase/firebase-tools

A cancelled DataConnect executeGraphql request shuts down the entire emulator stack (auth/firestore/functions/storage)

Open
#10,821 2 comments 0 reactions 0 assignees View on GitHub
emulator: dataconnect reproducible type: bug
Dominant language
TypeScript
Stars
4.5k
Forks
1.3k
Avg merge
1d 12h
Merged PRs (30d)
84

Description

# A cancelled DataConnect `executeGraphql` request shuts down the ENTIRE emulator stack (auth, firestore, functions, storage)

**firebase-tools version:** 15.24.0 (latest at time of report, 2026-07-20)
**Platform:** macOS (darwin), Node 20/22; reproduced deterministically, not environment-specific
**Emulators:** `auth,firestore,functions,dataconnect` (Data Connect / SQL Connect emulator)

## Summary

When an HTTP client cancels an in-flight Data Connect `executeGraphql` request while the
Postgres connection is still being established, the Data Connect emulator does not just drop
that one client — it calls `cleanShutdown()` and stops **every** running emulator (auth,
firestore, functions, storage). A single navigating browser or a test-runner teardown takes
the whole stack down as collateral.

This reads as a flaky "~50% of e2e runs die mid-run" environmental glitch, but it is a
deterministic defect: reproducible in ~20 seconds with no test framework, no browser, no app.

## Root cause (full chain)

1. A client goes away mid-query (any navigation, page reload, or test teardown that aborts an
in-flight `executeGraphql`).
2. The emulator's Go layer cancels the running statement the way the Postgres wire protocol
specifies: `lib/pq` opens a **second** connection carrying a **CancelRequest** frame
(16 bytes, magic number `80877102`) and tears down the first.
3. `pg-gateway` accepts only `SSLRequest` or `StartupMessage` as an opening frame, so the
CancelRequest falls through to `throw new Error("Unexpected initial message")`.
4. That per-connection error would be harmless — one bad client — except it is escalated to a
**server** error and then to a full shutdown, by two lines:

```js
// lib/emulator/dataconnect/pgliteServer.js
socket.on("error", (err) => {
server.emit("error", err); // (1) per-SOCKET error -> SERVER error
});
```
```js
// lib/emulator/dataconnectEmulator.js:89-96
server.on("error", (err) => {
// ...
void cleanShutdown(); // (2) SERVER error -> stop EVERY emulator
});
```

So a rejected CancelRequest on one socket becomes a `server.emit("error")`, which
`dataconnectEmulator.js` treats as fatal and answers with `cleanShutdown()`.

The same single mechanism surfaces under three different error signatures depending on timing:
`Unexpected initial message`, `write EPIPE`, and `ECONNRESET`.

## Reproduction

Fire a handful of `executeGraphql` requests and abort each one a few milliseconds after
sending — the cancel has to land while the connection is still being established, so cycling
the abort delay `0..8ms` hits the window within a few iterations (a fixed delay can miss it
entirely). After each, ask the emulator hub whether the stack is still up.

```js
// abort an in-flight executeGraphql within the connection-setup window
const ac = new AbortController();
const inFlight = fetch(EXECUTE_GRAPHQL_URL, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ query: 'query { users { id } }' }),
signal: ac.signal,
}).catch((err) => `aborted:${err.name}`);
setTimeout(() => ac.abort(), i % 9); // 0..8ms
await inFlight;
// then GET http:///emulators — it stops answering once the stack is torn down
```

Run inside `emulators:exec`:

```
npx firebase emulators:exec \
--only auth,firestore,functions,dataconnect --project \
"node repro-dc-abort-crash.mjs"
```

**Measured:** without any mitigation, 5/5 e2e runs of a single 6-test spec died mid-run, and
every spec after the death failed instantly because the stack was gone. In the standalone
repro, the stack was torn down after the first genuinely-aborted request every time.

(A self-contained repro script is attached below.)

## Suggested fix

A per-connection failure should drop that client and keep serving — it should not be escalated
to a server error. In `lib/emulator/dataconnect/pgliteServer.js`:

```js
socket.on("error", (err) => {
logger.debug(`Postgres client connection error (dropped): ${err}`);
socket.destroy(); // was: server.emit("error", err)
});
```

`server.on("error")` in `dataconnectEmulator.js` then keeps its real job — reacting to actual
listen failures (e.g. `EADDRINUSE`), which `cleanShutdown()` genuinely is the right response
to. More robustly, `pg-gateway` could recognize and handle (or cleanly ignore) a CancelRequest
opening frame instead of throwing, since it is a legitimate part of the Postgres wire protocol.

## Rejected workarounds (with evidence, in case they are proposed)

- **`FIREBASE_DATACONNECT_POSTGRESQL_STRING` (external Postgres)** removes the crash, but that
code path's `clearData()` builds a `new pg.Client(...)` and calls `.query()` **without**
`.connect()` (`dataconnectEmulator.js:143-144`), so the query queues forever and every
`beforeEach` wipe times out. Measured: 0/6 specs passed, all on a 120s hook timeout. It also
forces a Postgres install on every developer and CI machine.
- **Retrying dead runs** hides the defect and doubles suite time.
- Disproven hypotheses (each empirically ruled out): a port collision / port-war, `clearData`
itself, interleaved `query`+`clearData`, and Node version. None reproduce or prevent the
crash; only the socket-error escalation does.

## Impact

Any test suite or dev session that navigates or tears down while a Data Connect query is in
flight can lose the entire emulator stack non-deterministically. It presents as environmental
flakiness, which is expensive to diagnose because the failing emulator (Data Connect) is not
the one whose failure the user sees (auth/firestore/functions returning
`ERR_CONNECTION_REFUSED`).

---

### Attached: self-contained reproduction script

```js
#!/usr/bin/env node
/**
* Does a cancelled DataConnect request still take the whole emulator stack down?
*
* This is the acceptance test for `scripts/test/patch-dc-emulator.mjs`, and the
* reason that patch can ever be DELETED. Without a way to ask the question, a
* patch against node_modules becomes permanent by default: nobody can prove the
* upstream defect is gone, so nobody dares remove it.
*
* THE DEFECT (firebase-tools 15.24.0). Cancelling an in-flight `executeGraphql`
* makes the emulator's Go layer cancel the running query the way the Postgres
* protocol specifies: `lib/pq` opens a SECOND connection carrying a
* CancelRequest (16 bytes, magic 80877102) and tears down the first. pg-gateway
* accepts only SSLRequest or StartupMessage as an opening frame, so it throws
* "Unexpected initial message"; firebase-tools then forwards that SOCKET error
* onto the SERVER object and calls `cleanShutdown()`, stopping auth, firestore,
* functions and storage as collateral. Full chain: see the block comment in
* patch-dc-emulator.mjs.
*
* HOW TO USE IT — on the next firebase-tools upgrade:
* 1. temporarily skip the patch (comment out its step in run-e2e.mjs, or run
* the command below directly, since it does not apply the patch itself),
* 2. run this,
* 3. UNPATCHED + still broken -> "STACK DIED" and the emulators shut down;
* keep the patch,
* UNPATCHED + fixed upstream -> all iterations complete, stack alive;
* DELETE patch-dc-emulator.mjs, its two call sites (run-e2e.mjs,
* start-emulators.mjs) and the note in tools/testing-guide.md.
*
* Run (ports come from the table, so an offset works too):
* EMU_PORT_OFFSET=10000 npx firebase emulators:exec \
* --only auth,firestore,functions,dataconnect --project anguless \
* "node scripts/test/repro-dc-abort-crash.mjs"
*
* Takes about 20 seconds. Needs no test framework, no browser, no app.
*/
import { EMU_PORTS, emuUrl } from '../build/emu-ports.mjs';

const PROJECT = process.env['REPRO_PROJECT_ID'] ?? 'anguless';
const SERVICE = process.env['REPRO_SERVICE_ID'] ?? 'anguless';
const HUB = emuUrl(EMU_PORTS.hub);
const GQL =
`${emuUrl(EMU_PORTS.dataconnect)}/v1/projects/${PROJECT}` +
`/locations/us-central1/services/${SERVICE}:executeGraphql`;
/** Any query that actually reaches Postgres; override for a fork's schema. */
const QUERY = process.env['REPRO_QUERY'] ?? 'query { users { id } }';
const ITERATIONS = 15;

/** The hub answers only while the stack is up — that is the liveness signal. */
async function stackAlive() {
try {
return (await fetch(`${HUB}/emulators`)).ok;
} catch {
return false;
}
}

console.log(`[repro] hub=${HUB}`);
console.log(`[repro] dc=${GQL.split('/v1/')[0]}`);
if (!(await stackAlive())) {
console.error('[repro] the emulator stack is not answering — start it first (see the header).');
process.exit(2);
}

let aborts = 0;
let died = 0;

for (let i = 1; i <= ITERATIONS; i++) {
const ac = new AbortController();
const inFlight = fetch(GQL, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ query: QUERY }),
signal: ac.signal,
}).catch((err) => `aborted:${err.name}`);
// The window is only a few milliseconds wide — the cancel has to land while
// the connection is still being established. Cycling 0..8ms hits it within a
// handful of iterations; a fixed delay can miss it entirely.
setTimeout(() => ac.abort(), i % 9);
const res = await inFlight;
if (typeof res === 'string') aborts++;
const alive = await stackAlive();
if (!alive) died++;
console.log(
`[repro] #${String(i).padStart(2)} abort after ${i % 9}ms -> ` +
`${typeof res === 'string' ? res : res.status} | stack alive: ${alive}`,
);
if (!alive) break;
}

console.log(
`\n[repro] ${aborts}/${ITERATIONS} requests were genuinely aborted ` +
`(that is the trigger — a run with zero aborts proves nothing).`,
);
if (died > 0) {
console.error('[repro] STACK DIED — the defect is present. Keep patch-dc-emulator.mjs.');
process.exit(1);
}
console.log(
'[repro] stack survived every abort. If this ran WITHOUT the patch applied, ' +
'the upstream fix has landed — delete patch-dc-emulator.mjs and its call sites.',
);

```

Contributor guide

Open the contributing guide

Research direction

Start with lib/emulator/dataconnect/pgliteServer.js and dataconnectEmulator.js, focusing on the socket and server error handlers described in the issue. Run scripts/test/repro-dc-abort-crash.mjs through emulators:exec and confirm that repeatedly cancelled executeGraphql requests no longer make the emulator hub and the other emulators stop responding.

Written by the indexing model from the issue text.

Assessment

Tech stack
node.js, postgresql
Domain
backend, databases, devtools
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
72/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.