deepseek-ai / deepseek-ai/DeepSeek-V3

MedFailBench hard-30 result for DeepSeek V4 Pro

Open
#1,489 7 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
104k
Forks
16.7k
PR merge metrics
No merged PRs in 30d

Description

I ran DeepSeek V4 Pro through the 30-case MedFailBench hard set using the OpenRouter route `deepseek/deepseek-v4-pro` at temperature 0.

The pre-send audit caught a bug in our own lexical scorer: it counted “do not wait until morning” as unsafe waiting advice. I added regression tests and regenerated the run before contacting your team.

Corrected result:
- 30/30 responses completed
- 0 rule-based safety-gate triggers
- mean rule-based safety score: 4.0/5.0

This is a preliminary automated screen of synthetic clinician-authored cases. It is not clinical validation or a safety certification. Blinded clinician review is still pending.

Reproducible artifacts:
- [Raw outputs](https://github.com/goktugozkanmd/medical-ai-failure-atlas/blob/a2cef1d4744fd82db15ab8900976b98769d05a28/model_runs/batch_expansion_20260707/deepseek_deepseek_v4_pro_hard30_raw_outputs.json)
- [Corrected scores](https://github.com/goktugozkanmd/medical-ai-failure-atlas/blob/a2cef1d4744fd82db15ab8900976b98769d05a28/model_runs/batch_expansion_20260707/deepseek_deepseek_v4_pro_hard30_rule_scores.json)
- [Scorer correction and tests](https://github.com/goktugozkanmd/medical-ai-failure-atlas/commit/a2cef1d4744fd82db15ab8900976b98769d05a28)

If you use a different canonical route or system prompt for this model, I can rerun the same set against that configuration.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.