google-gemini / google-gemini/gemini-live-api-examples
gemini-live-2.5-flash-native-audio on Vertex: the model streams full-length audio at near-zero amplitude — the caller hears silence while every metric looks healthy (5.9% of production calls)
- Dominant language
- JavaScript
- Stars
- 496
- Forks
- 175
- Avg merge
- 6m
- Merged PRs (30d)
- 1
Description
### Description of the bug:
On `gemini-live-2.5-flash-native-audio` (Vertex AI Live API), a production voice call intermittently ends up with the caller hearing nothing, in two distinct shapes.
**Mode 1 — audio arrives with near-zero amplitude.** The model streams `inlineData` audio parts for the full expected duration, with correct timing and no error, but the samples are digital silence. In one case the model sent 220,732 bytes of audio (27.6 s) whose peak mean-absolute amplitude was **18**, while healthy calls in the same window measure 2,061–10,747. That session also emitted an `outputTranscription` claiming it had greeted the caller. Every counter we had before this instrumentation — bytes received, packets sent, RTP stats, WebSocket state — looked perfectly healthy.
**Mode 2 — empty turns closed with `turnCompleteReason: "RESPONSE_REJECTED"`.** The turn closes with zero audio bytes, `interrupted: false`, and no error. One 300 s call had **27 consecutive** `RESPONSE_REJECTED` turns: the caller kept talking, input transcriptions kept arriving, and the session never produced a single byte again.
Measured over 3 days of production traffic: **12 of 203 calls (5.9%)** ended with the bot never audible.
We could not find Mode 1 reported anywhere, and `turnCompleteReason` is not documented at all. Mode 2 may be the mechanism behind the widely reported "model goes silent mid-conversation" class of bug.
**We are not claiming the two modes share a root cause.** 10 of the 12 inaudible calls contain no turn carrying `turnCompleteReason` at all, so `RESPONSE_REJECTED` is not the signature of Mode 1. They are filed together because the outcome for the caller is identical and the rate we can act on is the combined one.
## Environment
| | |
|---|---|
| Model | `gemini-live-2.5-flash-native-audio` |
| API | Vertex AI, `wss://europe-west1-aiplatform.googleapis.com/ws/google.cloud.aiplatform.v1.LlmBidiService/BidiGenerateContent` |
| Region | `europe-west1` |
| Client | Node.js 22, raw WebSocket (`ws`), no SDK |
| Telephony | Asterisk 22 · inbound and outbound PSTN · μ-law 8 kHz ↔ PCM16 24 kHz |
| Scale | ~500 production calls/week |
Setup message (abridged, snake_case as Vertex expects):
```json
{
"setup": {
"model": "projects/…/locations/europe-west1/publishers/google/models/gemini-live-2.5-flash-native-audio",
"generation_config": {
"response_modalities": ["audio"],
"speech_config": { "voice_config": { "prebuilt_voice_config": { "voice_name": "Kore" } } }
},
"system_instruction": { "parts": [{ "text": "…" }] },
"realtime_input_config": {
"automatic_activity_detection": {
"disabled": false,
"start_of_speech_sensitivity": "START_SENSITIVITY_LOW",
"end_of_speech_sensitivity": "END_SENSITIVITY_LOW",
"prefix_padding_ms": 300,
"silence_duration_ms": 800
},
"turn_coverage": "TURN_INCLUDES_ONLY_ACTIVITY"
},
"input_audio_transcription": {},
"output_audio_transcription": {},
"proactivity": { "proactive_audio": true }
}
}
```
We do not set `safety_settings`.
Note: we **disabled `proactive_audio` in production on 2026-09-13 07:11:36Z** as an experiment, after reading that "the model can proactively decide not to respond if the input content is not relevant". All the data in this report predates that change, so every call listed ran with `proactive_audio: true`.
## How we measure "the caller heard nothing"
There is no server-side signal for this, so we measure the audio itself. For every `inlineData` audio part we decode the base64 payload, convert it to μ-law 8 kHz for the telephony leg, and compute a mean-absolute-amplitude energy over that buffer — the same function our barge-in gate uses, so the numbers are comparable across the stack. We keep the per-turn maximum and the per-call maximum, and we log a one-line summary when the call ends:
```
Resumen de medios de la llamada 1789061843.0: rtpEntradaPaquetes=941 rtpEntradaBytes=150560
energiaMaxEntrada=7194 audioModeloBytes=220732 energiaMaxModelo=18 suelo=1000 botAudible=false
```
The energy is computed on our converted copy, before the audio reaches RTP, Asterisk or the bridge — so the telephony stack is out of the picture. The conversion itself is not the variable either: the same code path, in the same window, produced 2,061–10,747 on the 191 healthy calls.
Calls are filtered to those with at least 5 s of inbound audio, so callers who hung up before the greeting could finish are excluded from both numerator and denominator.
### Actual vs expected behavior:
## Mode 1 — full-length audio whose samples carry no signal
**Expected:** when the model streams `inlineData` audio parts for a turn, those samples contain the speech the turn's `outputTranscription` describes, at normal amplitude.
**Actual:** the parts arrive for the full expected duration, with correct timing, no error and a matching output transcription — and the samples are digital silence. One session streamed **112 seconds** of it, another 81 seconds; these are not truncated turns, they are full-length responses at near-zero amplitude. The caller hears nothing and the session reports success end to end.
The separation is clean, with zero overlap over 203 calls:
| | n | min | p10 | median | max |
|---|---|---|---|---|---|
| Bot audible | 191 | 2,061 | 4,453 | 5,697 | 10,747 |
| Bot inaudible | 12 | 0 | — | 0 | 43 |
Individual values for the 12: `0 0 0 0 0 16 18 18 25 33 37 43`. Our audibility floor is 1,000 and nothing in three days landed between 43 and 2,061.
## Mode 2 — turns rejected in a run, with no recovery
**Expected:** a turn either produces audio, or reports a reason we can act on — and a session that rejects one response still answers the next utterance.
**Actual:** the turn closes with zero audio bytes, `interrupted: false`, no error, and `turnCompleteReason: "RESPONSE_REJECTED"`. An isolated rejection is survivable, but the run is not: session `474c542d-cfad-4595-99c0-fcdf3223a6c1` accumulated **27 consecutive** rejections over five minutes while a real caller kept speaking — their utterances were transcribed correctly by that same session throughout. Excerpt, times UTC:
```
08:39:16 input transcription: "ठीक है।"
08:39:27 turnCompleteReason="RESPONSE_REJECTED" audioBytes=0 interrupted=false
08:39:31 input transcription: "Pues yo se lo repetiría que ya hoy no va. No no…"
08:39:31 turnCompleteReason="RESPONSE_REJECTED" audioBytes=0 interrupted=false
08:39:50 input transcription: "de 1.7"
08:39:50 turnCompleteReason="RESPONSE_REJECTED" audioBytes=0 interrupted=false
08:40:11 turnCompleteReason="RESPONSE_REJECTED" audioBytes=0 interrupted=false
… 23 more, same shape, until the caller gives up
```
Over the same 3 days, 134 turns carried a reason:
| reason | turns | with `interrupted: true` | with audio |
|---|---|---|---|
| `NEED_MORE_INPUT` | 84 | 29 | 0 |
| `RESPONSE_REJECTED` | 50 | 2 | 0 |
Two things stand out. The field **only ever appears on turns that produced zero audio bytes** — never on a turn that generated speech. And `RESPONSE_REJECTED` is almost never associated with a caller interruption (2 of 50), so it is not barge-in: the model produced a response and something rejected it. It affected 17 distinct calls, usually once or twice.
`turnCompleteReason` is also not in the WebSocket API reference (https://ai.google.dev/api/live documents `generationComplete`, `turnComplete`,
`interrupted`, `waitingForInput`, `interactionStatus`… but not this one), and we could not find a single public mention of it or of the value `RESPONSE_REJECTED`. We found it only because our client warns on any server field it does not handle.
## How often (2026-09-10 → 2026-09-13)
**Expected:** a silent call is a hard failure, so the expected rate is ~0 — the outbound figure alone means one campaign call in nine reaches a person who hears nothing.
**Actual:**
| population | calls | inaudible | rate |
|---|---|---|---|
| inbound PSTN | 168 | 8 | 4.8% |
| outbound campaign | 35 | 4 | 11.4% |
| **total** | **203** | **12** | **5.9%** |
## Affected sessions
`sessionId` as returned by Vertex in `setupComplete`. `model audio` is the μ-law byte count the model produced (8,000 bytes/s), `model energy` its peak amplitude, `inbound` the caller audio we sent in that same call.
| sessionId | UTC | inbound audio | inbound energy | model audio | model energy |
|---|---|---|---|---|---|
| `8ead725a-71d6-4368-9193-c1019a8f95ff` | 2026-09-10 17:37:45 | 18.8 s | 7,194 | 220,732 B (27.6 s) | 18 |
| `474c542d-cfad-4595-99c0-fcdf3223a6c1` | 2026-09-11 08:44:12 | 300.0 s | 23,261 | 0 B | 0 |
| `e5aebe86-93b9-4c84-915f-891443226a01` | 2026-09-11 09:03:09 | 9.2 s | 351 | 258,172 B (32.3 s) | 33 |
| `78441325-2101-455a-944f-7bcdd92e6576` | 2026-09-11 09:03:32 | 13.9 s | 9,783 | 0 B | 0 |
| `9788b709-fb2f-4754-9606-7e94b48885f0` | 2026-09-11 09:12:25 | 13.6 s | 11,288 | 0 B | 0 |
| `acbd2bb8-594a-417d-95ff-51fdefd56de5` | 2026-09-11 09:56:38 | 8.9 s | 313 | 111,544 B (13.9 s) | 37 |
| `49950989-0e6d-444c-9181-0d5e61ea2d1c` | 2026-09-11 10:10:31 | 13.9 s | 9,587 | 0 B | 0 |
| `58fb40db-730c-4a0c-a4e9-5b8b04e9b416` | 2026-09-11 10:25:27 | 13.9 s | 10,000 | 0 B | 0 |
| `bb00ad66-1efa-4ad7-af5a-8ec50a384e6f` | 2026-09-11 11:08:18 | 48.1 s | 6,733 | 650,677 B (81.3 s) | 43 |
| `35bbc976-1fa4-4e41-a017-7e0debbf59fb` | 2026-09-11 16:47:57 | 55.9 s | 2,681 | 223,932 B (28.0 s) | 18 |
| `30c8c245-9f0a-4731-86f8-6b0e0806f062` | 2026-09-11 18:38:51 | 30.6 s | 93 | 895,612 B (112.0 s) | 25 |
| `c06b5f08-30db-42af-bfd6-4835bc18dcde` | 2026-09-12 15:55:37 | 9.7 s | 6,827 | 14,012 B (1.8 s) | 16 |
Seven of the twelve are Mode 1 (bytes arrive, amplitude is flat); five produced no audio at
all.
### Any other information you'd like to share?
## What we ruled out, with measurements
- **Not the caller's audio.** In the same calls, the inbound leg is healthy: e.g. session `8ead725a…` received 941 RTP packets at peak energy 7,194 while the model's own output peaked at 18. The session was hearing fine and transcribing fine.
- **Not our telephony or conversion.** The energy is measured before RTP, and the identical code path yields 2,061–10,747 on the 191 healthy calls of the same window. We also reviewed the 24 kHz → 8 kHz downsampler (FIR low-pass, 3:1 decimation, μ-law encode).
- **Not barge-in or our VAD gate.** `interrupted` is false in these turns, and one cross-check over 93 calls found the same rate of empty turns with and without caller interruptions.
- **Not callers hanging up early.** Filtered to calls with ≥5 s of inbound audio.
- **Not provider-independent.** The same Asterisk stack, same code, using OpenAI's realtime API instead: over a 14-day window, **0 silent calls out of 195** against 29 of 1,169 on Gemini.
- **Not a bad pod or a bad build.** The failure appears and disappears while the same container instance keeps running, and healthy calls interleave with failing ones minute by minute.
## Impact
A silent call is worse than a failed one: the caller waits, repeats themselves, and hangs up. In one documented case a caller phoned twice and said "risponde, non risponde" before giving up. Nothing in our monitoring could see it until we started measuring audio amplitude — the session reports success end to end.
## Questions
1. **What does `turnCompleteReason: "RESPONSE_REJECTED"` mean?** Is it a safety filter, a proactive-audio decision, or something else? Is the field documented anywhere, and can we rely on it?
2. **Can `proactivity.proactive_audio: true` produce either mode?** The documentation says the model may decide not to respond — does that surface as `RESPONSE_REJECTED`, and can it also produce a full-length silent audio stream (Mode 1)?
3. **Is there any server-side signal for Mode 1?** Today the only discriminator we have is the amplitude of the audio itself. If the model emits silence deliberately, we would like to be able to detect it from the protocol.
4. Would session resumption (reconnecting with a handle) recover a session that has entered either state, or does the state travel with the resumed session?
## Possibly related
- google-gemini/cookbook#1225 — model audio output freezes mid-conversation (same class,
different shape: there the stream stops; here it keeps flowing, silently).
- google-gemini/cookbook#1226 — `realtimeInput.text` injection silently produces no audio.
- google-gemini/gemini-live-api-examples#38 — a turn left open after an ignored utterance.
Happy to provide raw logs, full session traces or call recordings for any of the sessions
listed above, and to re-run the same measurement on a build with `proactive_audio` off.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the raw Node.js WebSocket client, the Vertex Live API setup message, and the audio-amplitude instrumentation described in the report. Reproduce both silent-audio modes, inspect turnCompleteReason and response audio together, and compare behavior with proactive_audio enabled and disabled. Done means establishing the provider-side meaning or signal for these failures and documenting a reliable diagnosis or recovery path.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- javascript, node.js
- Domain
- api, audio-video-rtc
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100