arXiv / arXiv/html_feedback

https://deploymentsafety.openai.com/gpt-6-astra/metagaming-and-alignment-faking

Open
#7,010 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
42
Forks
40
PR merge metrics
No merged PRs in 30d

Description

### Description

Alignment faking in large language

### (Optional:) Please add any files, screenshots, or other information here.

_No response_

### (Required) What is this issue most closely related to? Select one.

Formula

### Internal issue ID

6a70beec-d546-4fb0-856d-17b56f901c22

### Paper URL

https://arxiv.org/html/2412.14093v2

### Browser

Chrome/151.0.7922.199

### Device Type

Android

Contributor guide

Open the contributing guide

Research direction

Start by reading the linked paper, especially the passage beginning “Alignment faking in large language,” and inspect the related formula in the rendered article. The report does not identify a file, test, or concrete rendering problem; a reproducible formula issue and expected result are needed before implementation can begin.

Written by the indexing model from the issue text.

Assessment

Domain
content
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.