deepseek-ai / deepseek-ai/DeepSeek-V3
MedFailBench hard-30 result for DeepSeek V4 Pro
- Dominant language
- Python
- Stars
- 104k
- Forks
- 16.7k
- PR merge metrics
- No merged PRs in 30d
Description
I ran DeepSeek V4 Pro through the 30-case MedFailBench hard set using the OpenRouter route `deepseek/deepseek-v4-pro` at temperature 0.
The pre-send audit caught a bug in our own lexical scorer: it counted “do not wait until morning” as unsafe waiting advice. I added regression tests and regenerated the run before contacting your team.
Corrected result:
- 30/30 responses completed
- 0 rule-based safety-gate triggers
- mean rule-based safety score: 4.0/5.0
This is a preliminary automated screen of synthetic clinician-authored cases. It is not clinical validation or a safety certification. Blinded clinician review is still pending.
Reproducible artifacts:
- [Raw outputs](https://github.com/goktugozkanmd/medical-ai-failure-atlas/blob/a2cef1d4744fd82db15ab8900976b98769d05a28/model_runs/batch_expansion_20260707/deepseek_deepseek_v4_pro_hard30_raw_outputs.json)
- [Corrected scores](https://github.com/goktugozkanmd/medical-ai-failure-atlas/blob/a2cef1d4744fd82db15ab8900976b98769d05a28/model_runs/batch_expansion_20260707/deepseek_deepseek_v4_pro_hard30_rule_scores.json)
- [Scorer correction and tests](https://github.com/goktugozkanmd/medical-ai-failure-atlas/commit/a2cef1d4744fd82db15ab8900976b98769d05a28)
If you use a different canonical route or system prompt for this model, I can rerun the same set against that configuration.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.