google / google/langextract

Bug: Gemini batch path returns empty success on a safety-blocked or refused response

Open
#527 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
38.6k
Forks
2.7k
Avg merge
3d 15h
Merged PRs (30d)
3

Description

## Describe the overall issue and situation

In the Gemini batch inference path, a batch item that was safety-blocked or refused is silently reported as a **successful empty extraction** instead of surfacing the block/refusal reason.

`langextract/providers/gemini_batch.py` (`_parse_batch_line`):

```python
resp = obj.get("response", {})
text = _extract_text(resp) or ""
```

`_extract_text` returns `None` when it can't find `resp.text` / `candidates[0].content.parts[0].text` — which includes both "the field genuinely doesn't exist" and "this candidate was blocked/refused with no text content." The `or ""` collapses both cases into an empty string, indistinguishable from a model that actually returned empty text. There's no check of `obj["response"]`'s own block/finish-reason fields before falling back.

This is the same class of bug as the realtime paths: #491 (OpenAI, fixed in #496) and #508 (Gemini realtime, PR #507 open) — both raise `InferenceRuntimeError` with a diagnostic instead of silently returning empty output on a safety-blocked/refused response. Neither #496 nor #507 touched `gemini_batch.py`, so the batch path still has the gap.

## Expected behavior

When a batch item's response has no extractable text because it was blocked or refused (as opposed to a response that's simply malformed/missing), `_parse_batch_line` should surface that as an error for the item rather than silently writing `""` to `outputs[idx]`.

## Actual behavior

`_parse_batch_line` writes `""` to `outputs[idx]` for a blocked/refused item, indistinguishable from a genuinely empty extraction. The block/refusal reason (if present in the response's own feedback/finish-reason fields) is discarded.

## Steps to reproduce the issue

1. Submit a batch job containing a prompt that trips Gemini's safety filters, or otherwise produces a candidate with no text content.
2. In the resulting JSONL output, that line's `response` has no extractable `text` (mirrors the realtime path's `response.text is None` case, confirmed against `google-genai`'s response shape in #508).
3. Observe `_parse_batch_line` sets `outputs[idx] = ""` instead of raising or otherwise flagging the item.

## Any additional content

- Verified without live API: current `main` (`langextract/providers/gemini_batch.py:622`) still has `text = _extract_text(resp) or ""` with no check of the response's own block/finish-reason fields first.
- Fix should follow the same shape as #507's realtime-path fix: inspect the response for a block/finish reason before falling back to empty, and raise `InferenceRuntimeError` with the available diagnostic when one is found — while leaving a response that's genuinely just empty (no diagnostic available) as `cfg.ignore_item_errors`-gated behavior, matching how `_parse_batch_line` already treats the sibling `obj["error"]` field.

Contributor guide

Open the contributing guide

Research direction

Start in langextract/providers/gemini_batch.py at _parse_batch_line and compare its handling with the realtime-path fix in #507. Verify that blocked or refused responses surface an InferenceRuntimeError with available diagnostics, while genuinely empty responses retain the existing cfg.ignore_item_errors behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, backend-api-design
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
74/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.