anthropics / anthropics/claude-code

[BUG] Byte corruption of Latin-1/Windows-1252 single byte characters

未关闭
#93,954 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
area:tools bug has repro platform:macos
主要语言
Python
星标
145k
派生
23.1k
PR 合并指标
PR 指标待抓取

描述

### Preflight Checklist

- [x] I have searched [existing issues](https://github.com/anthropics/claude-code/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug) and this hasn't been reported yet
- [x] This is a single bug report (please file separate reports for different bugs)
- [x] I am using the latest version of Claude Code

### What's Wrong?

## PHP files with special characters (¢, ½, •, etc.) - byte-level editing only

Several PHP files store special characters (¢, ½, •, ü, smart quotes/dashes, etc.) as raw
single-byte Latin-1/Windows-1252 characters, not UTF-8 - matching this site's own
`charset=iso-8859-1` convention (which real browsers, per the WHATWG spec, actually render as
Windows-1252 for the 0x80-0x9F range). A raw byte like `0xA2` (¢) is not valid UTF-8 on its own.

**Never edit such a file with a normal text tool - including Claude Code's own Edit/Write -
without first checking for non-ASCII bytes.** Any tool that reads a file as UTF-8 text and
writes it back the same way will silently replace every byte it can't decode with the Unicode
replacement character on save - and that corruption hits every OTHER special character
elsewhere in the file too, not just the line actually being edited. Check first:
`grep -Pn '[\x80-\xff]' ` or
`python3 -c "print(any(b>=0x80 for b in open('','rb').read()))"`. If a file has any such
bytes, edit it via a byte-level script (Python `open(path,'rb')`/`'wb'`, no text-mode decode) -
never Edit/Write, on any machine, in any session.

Confirmed 2026-09-12: this exact failure mode corrupted 162 characters across 11 files between
2026-05-25 and 2026-09-11, discovered while investigating a garbled ¢ the user noticed. Traced
via git history (each corrupted line diffed against its introducing commit's parent to recover
the exact original byte - contextual guessing would have missed at least one case, a `ü`, not a
currency symbol as the surrounding code's pattern would have suggested). Several of the
introducing commits are directly Claude-Code-authored, including one reproduced live in the same
session that found this (an Edit call meant to fix one line wiped out every other
correctly-restored byte in the file). Treat this as an active, ongoing, cross-machine risk - not
a one-off - until proven otherwise.

### What Should Happen?

Claude should not silently corrupt my files and not notice.

### Error Messages/Logs

```shell

```

### Steps to Reproduce

Create php filie

with a string of any of these characters ¢ ½

Ask claude edit the file to add a comment. View the saved file, these characters have been corrupted

### Claude Model

Sonnet (default)

### Is this a regression?

I don't know

### Last Working Version

_No response_

### Claude Code Version

2.1.269 (Claude Code)

### Platform

Anthropic API

### Operating System

macOS

### Terminal/Shell

Terminal.app (macOS)

### Additional Information

Is the a bug? It seems serious? Should it be reported?

Good questions — let me answer directly: yes, this is a real bug, and yes it should be reported. Silently substituting bytes it can't decode — rather than either preserving them byte-exact or refusing to edit the file with a clear warning — is a genuine data-integrity flaw, not user error. It's especially serious because the damage isn't confined to the line being edited; it corrupts unrelated content anywhere else in the file. Legacy codebases in Latin-1/Windows-1252 (common for older PHP/web projects) are exactly the case this breaks. I'd recommend filing it at https://github.com/anthropics/claude-code/issues — that's the official channel for exactly this kind of report.

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。