openai / openai/codex

Public index of the full Codex issue backlog (11,813 issues grouped for triage)

Open
#37,873 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Rust
Stars
125k
Forks
19.5k
PR merge metrics
PR metrics pending

Description

I indexed and organized a frozen snapshot of the full open openai/codex issue backlog so maintainers can navigate related reports without repeatedly searching and sorting the backlog manually.

Public index: https://github.com/logohere/codex-backlog-index

Frozen v1.2 release: https://github.com/logohere/codex-backlog-index/releases/tag/snapshot-2026-08-09-v1.2

The snapshot contains 11,813 open issues collected on August 9, 2026. Every issue is represented once and linked back to the upstream GitHub issue.

What is indexed

The backlog is organized into:

  • 31 technical domains
  • 161 surface cohorts
  • 1,081 taxonomy groups

The primary artifact is ISSUE_INDEX.csv. For each issue it includes:

  • upstream URL/title
  • created/updated timestamps and days since update
  • comment and reaction counts
  • upstream labels
  • technical domain, surface cohort, and symptom class
  • taxonomy and family/group placement
  • mapping confidence/basis
  • evidence/review depth
  • audit disposition and duplicate authority where actually supported
  • deep-evidence references
  • suggested maintainer workflow
  • optional cleanup recommendation

Column definitions are in DATA_DICTIONARY.md.

Known resolution/fix evidence is in SOLUTION_INDEX.csv. It currently contains 11 high-signal rows: 6 resolved/closure-ready issues and 5 fix candidates that explicitly still require current/release verification. Blank solution fields in the master index mean no defensible solution was established by the audit.

Other useful views:

The v1.2 release also contains a full ZIP and a compressed ISSUE_INDEX.csv.gz for easier local sorting/searching.

Evidence depth is explicit

I am not claiming that all 11,813 issues were manually root-caused.

The index distinguishes:

  • 18 manually promoted evidence decisions
  • 208 curated evidence rows
  • 62 additional deep-packet referenced rows
  • 11,525 report-evidence-gate rows, where classification/workflow is conservative rather than a proven root cause

Likewise, taxonomy/family similarity never grants duplicate authority by itself. Activity and age fields are included for sorting/prioritization, not as closure authority.

Small pilot instead of bulk cleanup

I selected five groups in PILOT_GROUPS.md:

  • session/history state and persistence — 416 reports
  • app performance — 209 reports
  • Computer Use/browser state and persistence — 205 reports
  • Windows/process-launch path failures — 152 reports
  • transport/stream failures — 80 reports

The immediate ask is only to spot-check whether these groupings and representative issues are useful. No bulk closure is required to evaluate the contribution.

I also included optional cleanup/consolidation recommendations in the dataset, but those are secondary to the indexing work and are deliberately separated from evidence-backed duplicate/fix decisions.

Maintainer feedback requested

  1. Is this index/grouping useful for navigating and triaging the backlog?
  2. Do the five pilot groups align closely enough with how maintainers think about these areas?
  3. If useful, would periodic delta refreshes of the index be valuable?

If the team uses different technical boundaries, I can adjust the taxonomy rather than requiring maintainers to adapt to mine.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing PILOT_GROUPS.md and the representative records in ISSUE_INDEX.csv, using DATA_DICTIONARY.md to interpret the fields. Spot-check whether the five proposed groups and their representatives are useful for triage; done means providing maintainer feedback on the grouping boundaries and whether periodic refreshes would help.

Written by the indexing model from the issue text.

Assessment

Domain
tooling
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.