chaitin / chaitin/MonkeyCode

guardrail 误判问题

Open
#952 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
TypeScript
Stars
4.7k
Forks
719
Avg merge
4h 30m
Merged PRs (30d)
83

Description

反馈标题:guardrail 误判

反馈内容:希尔伯特、皮亚诺、戴德金、柯西的理论是实质公理论还是形式公理论,这个问题触发guardrail 误判,显示违规内容。但切换模型之后正常回答。

期望效果:正常回答学术类问题

我的 UID:019eed3e-af8c-784b-9928-b5e2132e275d

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are identified in the report. Reproduce the false positive with the cited Chinese academic question, compare behavior across models, then trace the guardrail handling to identify where this input is rejected. Done means the question receives a normal academic response without weakening protection for genuinely unsafe content.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.