google-deepmind / google-deepmind/open_spiel

Interest check: multi-agent reward-hacking / matched-proxy game environments?

Open
#1,563 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
5.5k
Forks
1.2k
Avg merge
2d 8h
Merged PRs (30d)
4

Description

Low-key interest check rather than a full proposal, since I know new-game additions are a bigger ask: would there be interest in multi-agent reward-hacking test games — where a matched-proxy construction (an identical proxy payoff across two game variants, with the true/social payoff diverging) is used to test whether an agent's strategy generalizes beyond exploiting the proxy?

I've built RHOB, a single-agent benchmark of 14 such matched-proxy environments (https://github.com/Aarav500/rhob), and OpenSpiel's game-theoretic framing seems like a natural fit for extending this to genuinely multi-agent/game-theoretic settings (e.g. collusion around a flawed shared payoff) in a way single-agent Gym-style environments can't capture.

If there's genuine interest I'd scope 1-2 small games following OpenSpiel's conventions before anything larger. If not, no worries — mostly wanted to ask given the fit before spending time on something that might not land.

Contributor guide

Open the contributing guide

Research direction

No OpenSpiel files, tests, or entry points are identified. Start by reviewing OpenSpiel's game conventions and the linked RHOB benchmark, then clarify whether there is interest in adding 1–2 small multi-agent matched-proxy games. Done would require an agreed scope and concrete game specifications before implementation begins.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
game-dev, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.