microsoft-foundry / microsoft-foundry/foundry-samples

Foundry local: Python live audio transcription exits with code 1 after session.append(), no Python traceback

Open
#788 1 comment 0 reactions 4 assignees View on GitHub

@brandom-msft is already working on this.

Since Jul 5, 2026.

bug
Dominant language
Bicep
Stars
445
Forks
494
Avg merge
11h 35m
Merged PRs (30d)
38

Description

Summary

When using the Foundry Local Python SDK live audio transcription API with nemotron-speech-streaming-en-0.6b, the process exits immediately after the streaming session starts and audio data is appended.

The microphone can be opened and read successfully with PyAudio, but the live transcription sample exits with code 1 right after printing the "LIVE TRANSCRIPTION ACTIVE" banner. No Python traceback is emitted.

The same failure also happens when using synthetic PCM audio, so this does not appear to be caused only by the microphone device or PyAudio capture settings.

Environment

  • OS: Windows
  • Python environment: conda
  • Python SDK package versions:
    • foundry-local-sdk==1.2.3
    • foundry-local-core==1.2.3
    • onnxruntime-core==1.26.0
    • onnxruntime-genai-core==0.14.1
  • Model:
    • Alias: nemotron-speech-streaming-en-0.6b
    • Resolved model: nemotron-speech-streaming-en-0.6b-cuda-gpu:1
  • Audio format:
    • PCM signed 16-bit little-endian
    • Mono
    • 16000 Hz
  • Microphone capture library:
    • PyAudio

Reproduction Code

Minimal live transcription flow:

import math
import struct
import time

from foundry_local_sdk import Configuration, FoundryLocalManager

config = Configuration(
    app_name="foundry_local_samples",
    log_level="info",
    model_cache_dir="C:\\Users\\shiji\\.foundry\\cache\\models",
)

FoundryLocalManager.initialize(config)
manager = FoundryLocalManager.instance
manager.download_and_register_eps()

model = manager.catalog.get_model("nemotron-speech-streaming-en-0.6b")
model.load()

audio_client = model.get_audio_client()
session = audio_client.create_live_transcription_session()
session.settings.sample_rate = 16000
session.settings.channels = 1
session.settings.language = "en"

session.start()
print("session started")

# Push synthetic 16-bit PCM audio.
sample_rate = 16000
duration = 2
total_samples = sample_rate * duration
pcm_bytes = bytearray(total_samples * 2)

for i in range(total_samples):
    t = i / sample_rate
    sample = int(32767 * 0.5 * math.sin(2 * math.pi * 440 * t))
    struct.pack_into("<h", pcm_bytes, i * 2, sample)

chunk_size = (sample_rate // 10) * 2

for offset in range(0, len(pcm_bytes), chunk_size):
    end = min(offset + chunk_size, len(pcm_bytes))
    session.append(bytes(pcm_bytes[offset:end]))
    time.sleep(0.1)

print("audio pushed")

session.stop()
model.unload()
print("done")

Actual Behavior

The process exits with code 1 shortly after audio is appended.
Example output:

[foundry-local] | INFO | Native libraries found ...
[foundry-local] | INFO | Foundry.Local.Core initialized successfully ...
done.
✓ Session started
Pushing synthetic audio (440Hz sine, 2s)...
Command exited with code 1

When using microphone input, the output is similar:
✓ Session started

LIVE TRANSCRIPTION ACTIVE
Speak into your microphone.
Press Ctrl+C to stop.

Command exited with code 1

No Python exception or traceback is printed.

Expected Behavior

The process should continue running after session.append(...), and transcription results should be available from:

for result in session.get_stream():
    print(result.content[0].text)

If the audio format, model, execution provider, or streaming state is invalid, the SDK should raise a Python exception such as FoundryLocalException instead of terminating the process.

Possible Root Cause

The failure appears to happen in the native live audio streaming path used by Python SDK:

session.append(pcm_data)

which eventually calls the native command:

audio_stream_push

Additional Diagnostics

I verified that microphone capture itself works independently of Foundry Local:

import time
import pyaudio

RATE = 16000
CHANNELS = 1
CHUNK = RATE // 10

p = pyaudio.PyAudio()
stream = p.open(
    format=pyaudio.paInt16,
    channels=CHANNELS,
    rate=RATE,
    input=True,
    input_device_index=1,
    frames_per_buffer=CHUNK,
)

print("mic stream opened")
start = time.time()
chunks = 0

while time.time() - start < 5:
    data = stream.read(CHUNK, exception_on_overflow=False)
    chunks += 1

print(f"read {chunks} chunks successfully")

stream.stop_stream()
stream.close()
p.terminate()

This succeeds: mic stream opened
read 50 chunks successfully

Since the Python process exits without a Python traceback, this may be a native crash or unhandled native-side failure in [execute_command_with_binary] / audio_stream_push, possibly specific to the CUDA model variant: nemotron-speech-streaming-en-0.6b-cuda-gpu:1

Request

Could you please check whether live audio transcription with nemotron-speech-streaming-en-0.6b-cuda-gpu:1 is currently supported through the Python SDK on Windows?

It would also be helpful if the SDK could surface native audio_stream_push failures as Python exceptions instead of terminating the process.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.