https://deploymentsafety.openai.com/gpt-6-astra/metagaming-and-alignment-faking
- Dominant language
- No language data
- Stars
- 42
- Forks
- 40
- PR merge metrics
- No merged PRs in 30d
Description
### Description
Alignment faking in large language
### (Optional:) Please add any files, screenshots, or other information here.
_No response_
### (Required) What is this issue most closely related to? Select one.
Formula
### Internal issue ID
6a70beec-d546-4fb0-856d-17b56f901c22
### Paper URL
https://arxiv.org/html/2412.14093v2
### Browser
Chrome/151.0.7922.199
### Device Type
Android
Contributor guide
Research direction
Start by reading the linked paper, especially the passage beginning “Alignment faking in large language,” and inspect the related formula in the rendered article. The report does not identify a file, test, or concrete rendering problem; a reproducible formula issue and expected result are needed before implementation can begin.
Written by the indexing model from the issue text.
Assessment
- Domain
- content
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100