guardrail 误判问题
Open
Nobody has claimed this yet.
- Dominant language
- TypeScript
- Stars
- 4.7k
- Forks
- 719
- Avg merge
- 4h 30m
- Merged PRs (30d)
- 83
Description
反馈标题:guardrail 误判
反馈内容:希尔伯特、皮亚诺、戴德金、柯西的理论是实质公理论还是形式公理论,这个问题触发guardrail 误判,显示违规内容。但切换模型之后正常回答。
期望效果:正常回答学术类问题
我的 UID:019eed3e-af8c-784b-9928-b5e2132e275d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are identified in the report. Reproduce the false positive with the cited Chinese academic question, compare behavior across models, then trace the guardrail handling to identify where this input is rejected. Done means the question receives a normal academic response without weakening protection for genuinely unsafe content.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai, security
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100