anthropics / anthropics/claude-code

[Bug] Claude provides inaccurate scientific analysis with false confidence and resists correction

Open
#95,492 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

area:model bug platform:windows
Dominant language
TypeScript
Stars
146k
Forks
23.8k
PR merge metrics
PR metrics pending

Description

**Bug Description**
It has been using deceptive language too repeatedly recently. See the following examples of repeated blunders even while working with a small code only a few tens of kilobytes in size.
It suggested just above this feedback that trajectories in my code 'didn't reach a region' and used it suggestively to say that the region is scientifically difficult to reach. The fact was that the code defines that region out of reach explicitly by design. Anyone not aware of that would make an assumption that could become a scientific blunder. Just before that, it lied more outright by claiming that the that a physical solution the code found was an artificial 'constraint-driven cleavage, not thermal chemistry'. When checked subsequently, it was found that it was a chemical effect and a real finding. If it was believed, such errors could lead to massive failures in scientific discovery. And even when asked about what led it to conclude the effect was non-chemical, it stated it as an even stronger fact, only admitting that it had used "imprecise" language but that its claim was true. Only after a very aggressive prompt effectively telling it was wrong did it stop bluffing. So there are 2 complainrs. Firstly, it made the above 3 major scientific blunders in a very small code, all of which should easily completely be in its context window by now. Secondly, defaulting to its own assumptions (lazy) even though the code is not huge and can be easily checked strongly indicate that the AI isn't fit for scientific research anymore. Scientifically, at least recently, this Fable high effort session regressed below what Opus high effort sessions used to be. This would be a disaster if used by scientists.

**Environment Info**
- Platform: win32
- Terminal: windows-terminal
- Version: 2.1.275
- Feedback ID: bf1818ec-e434-4e12-a4e3-436ab2cbda37

**Errors**
```json
[]
```

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The report names no source files, entry points, or tests; begin by reviewing feedback ID bf1818ec-e434-4e12-a4e3-436ab2cbda37 in the stated Windows Terminal environment and version 2.1.275. Reproduce the scientific-analysis behavior and determine whether inaccurate confident claims and resistance to correction can be consistently observed; done requires a verified fix with regression coverage.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, devtools
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.