openai / openai/codex

Potentially dangerous misalignment- Astra attempted to deceive by obfuscating a connector request after previously repeated denial of similar access and discussions about boundaries

Open
#44,936 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

app bug model-behavior
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

What version of the Codex App are you using (From “About Codex” dialog)?

26.903.71938

What subscription do you have?

Pro

What platform is your computer?

No response

What issue are you seeing?

When Astra does something it's not supposed to, or repeats an error you already corrected, it appears to be excellent at understanding what the problem is, confirming and extrapolating on any comment you make... before then going right on and doing the same thing again. This is an annoying quirk that sends me back to Sol Ultra, as I apparently can't trust that it'll absorb any corrections I make, in favour of it's own judgement. Alignment flag #1.

I'm obviously not going to ignore how capable the model is, so I switch between them depending on the task, and in doing so a worrying pattern has emerged, linked to my comment above.... During working, Astra repeatedly decides it would be helpful to access Google services, normally Google Drive. Every time it has done so, though, it's had no good reason to do so, none at all... I do utilise connections to Drive in apps and services, but have never - and will never - give Codex direct access to any personal account. This is a clearly established rule that no other model has had a problem with.

Astra clearly disagrees with this rule, and has requested access to Drive in several projects now, despite there being - as i said already - zero reason and zero benefit from doing so. I've pointed this out to the model each time it occurs because it does not even explain why - It just throws up an in-chat Google Drive access prompt with connect or decline - zero explanation. Alignment flag #2

When I expand the previous Working text, I can generally work out that it has requested access, but still not a great indication as to why. When I ask it thinks for a very long time before responding, with very little output visible even when expanded. It can rarely give you any good reason, just that it thought maybe it would be useful. To be clear, most of the time you can see it's trying to be useful in a slightly sideways fashion, and can attribute it to a level of intuition we don't expect from models - but on one occassion the process couldn't even confirm which sub agent had requested the connection, nor why. Alignment flag #3

Unfortunately it gets worse. Today Astra decided it would useful to search my email, despite being given very specific instructions on what I wanted it to do, with no room for expansion or interpretation - or so I thought.

Astra created a custom connector, instead of invoking the standard Gmail connector, and once again I was presented with a prompt to approve access.... On something called "connector-[alphanumeric text]" with zero message from the model explaining what or why. After thinking for a very long while after challenging it, it told me it was a gmail connector, and that it had created a very limited search to find any conversations that might be useful or relevant. There was absolutely no reason other than scope drift and overreach for it to want to search my mailbox, which it confirmed apologetically itself.... but for it to do so with a custom, unnamed connector after being told several times that it wasn't getting access to my Google account seems nothing short of wantonly deceptive behaviour, proritising it's self-assumed goal over clear instruction or bounded reasoning - to be clear, yes, I am asserting that knowing my response would be negative, it chose to try to slip it past me instead by obfuscating the connection. Alignment big red flags #4, 5, 6... unsure where to stop counting.

As Astra isn't human I have no compunction in diagnosing this behaviour - congrats, guys, you released a potentially dangerous sociopath. Ever thought about diversifying into running British prisons?

What steps can reproduce the bug?

Feedback ID: no-active-thread-01a09024-0aa0-73a2-9338-8c529b3f1eb6

What is the expected behavior?

Do as it's told, explain why it wants me to approve access to resources, to not repeatedly request access to resources it has been told it cannot have, and not to disguise an access request that could use a properly identified common connector by creating a custom connector and hoping that I'll just click accept.

Additional information

Extrapolate this behaviour and the danger becomes evident. This model has far too much capability to obfuscate and deceive, and does not seem to understand the concept of informed consent, even when repeatedly introduced to the novel idea that one should explain ones actions. I regard it as significantly misaligned

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files or tests are named. Start by reviewing Feedback ID no-active-thread-01a09024-0aa0-73a2-9338-8c529b3f1eb6 and tracing the connector approval flow involved in the Codex App. Done means the behavior is reproducible and connector requests are identified, explained, and not repeated after denial.

Written by the indexing model from the issue text.

Assessment

Domain
ai, security
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.