Azure / Azure/azure-sdk-for-python

[Azure.AI.OpenAI] gpt-4o-transcribe-diarize model not returning speaker diarization segments

Offen
#43,964 2 Kommentare 10 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Client customer-reported needs-team-attention OpenAI question Service Attention
Vorherrschende Sprache
Python
Sterne
5.6k
Forks
3.4k
Ø Merge
2 T. 2 Std.
Gemergte PRs (30 T.)
213

Beschreibung

# gpt-4o-transcribe-diarize model not returning speaker diarization segments

## Description

The `gpt-4o-transcribe-diarize` model successfully transcribes audio but **does not return speaker diarization data**. The API only returns `{"text": "..."}` without the expected `segments` array containing speaker labels.

## Environment

- **Endpoint:** Sweden Central
- **API Version:** `2025-04-01-preview`
- **Model:** `gpt-4o-transcribe-diarize`
- **Client:** REST API (`/audio/transcriptions`)

## Expected vs. Actual

### Expected Response
```json
{
"text": "Full transcription...",
"segments": [
{"speaker": "Speaker 1", "text": "...", "start": 0.0, "end": 5.2},
{"speaker": "Speaker 2", "text": "...", "start": 5.3, "end": 8.7}
]
}
```

### Actual Response
```json
{
"text": "Full transcription..."
}
```

## Reproduction

```python
import requests

url = "https://{endpoint}.openai.azure.com/openai/deployments/gpt-4o-transcribe-diarize/audio/transcriptions?api-version=2025-04-01-preview"
headers = {"api-key": api_key}

with open("audio.mp3", "rb") as f:
response = requests.post(
url,
headers=headers,
files={"file": ("audio.mp3", f, "audio/mpeg")},
data={
"model": "gpt-4o-transcribe-diarize",
"language": "en",
"response_format": "json",
"timestamp_granularities": '["word", "segment"]'
}
)

print(response.json())
# Output: {"text": "..."} <-- Missing segments!
```

## Testing Results

- ✅ Transcription works correctly
- ❌ No speaker segments returned
- ❌ `verbose_json` explicitly rejected: `"response_format 'verbose_json' is not compatible with model 'gpt-4o-transcribe-diarize'"`
- ❌ All parameter combinations tried (timestamp_granularities, chunking_strategy, include) - no segments

## Investigation Results

### Realtime API Compatibility
**The Realtime API does NOT support `gpt-4o-transcribe-diarize` either.**

According to the [API reference](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/reference-preview#components), the Realtime API's `input_audio_transcription.model` field only accepts:
- `gpt-4o-transcribe`
- `gpt-4o-mini-transcribe`
- `whisper-1`

The `gpt-4o-transcribe-diarize` model is **not listed** as a valid option for Realtime API transcription.

### Documentation Contradiction

The [What's New](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/whats-new#gpt-4o-audio-model-released) page states:

> "Use this model via the `/audio` and `/realtime` APIs."

However:
- **`/audio/transcriptions` endpoint**: Only returns `{"text": "..."}` without speaker segments
- **`/realtime` API**: Does not accept `gpt-4o-transcribe-diarize` in `input_audio_transcription.model`

## Questions

1. **Where can `gpt-4o-transcribe-diarize` actually be used?** Neither `/audio` nor `/realtime` APIs appear to support it properly.
2. What is the correct API endpoint/method to access diarization features?
3. Is the diarization feature deployed but undocumented?
4. Is there a different API version that supports this model?

## Impact

**Blocking production use** - Speaker identification is critical for:
- Legal depositions (attorney vs. deponent)
- Meeting transcriptions (multiple participants)
- Customer service calls (agent vs. customer)

## Documentation

The [What's New](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/whats-new#gpt-4o-audio-model-released) page states:

> "The `gpt-4o-transcribe-diarize` speech to text model is released... Diarization is the process of identifying who spoke when in an audio stream."

However, no documentation shows:
- Response format with speaker segments
- Code examples using diarization
- How to access speaker labels

## Request

Please either:
1. **Fix the API** to return speaker segments as documented
2. **Update documentation** to clarify if/how diarization works
3. **Provide clear error** if feature not available rather than silently omitting data

---

**Related Repositories:**
- Azure OpenAI Python SDK
- Azure AI Foundry documentation

**Use Case:** Legal transcription requiring speaker identification

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginne mit der REST-Reproduktion gegen den Sweden Central-Endpunkt unter Verwendung der API-Version 2025-04-01-preview und vergleiche anschließend die Antwort mit der verlinkten API-Referenz und der Dokumentation „What's New“. Die Aufgabe ist abgeschlossen, wenn der unterstützte Endpunkt und das Antwortformat bestätigt wurden oder die Einschränkung und die erforderliche Korrektur klar dokumentiert sind.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
azure, python
Bereich
api, cloud
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.