openai / openai/codex-security

Out-of-range numeric XML entity in a DOCX aborts the whole knowledge base and the scan

Open Beginner friendly
#40 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

area:cli bug priority:p2
Dominant language
TypeScript
Stars
10.8k
Forks
801
Avg merge
1d 8h
Merged PRs (30d)
257

Description

Summary

decodeXml passes the parsed value of a numeric XML entity straight to String.fromCodePoint with no magnitude bound. Any entity above U+10FFFF raises RangeError, which propagates out of extractDocx and fails the whole prepareKnowledgeBase call, aborting the scan before it starts.

https://github.com/openai/codex-security/blob/f22d4a36f26d16287bcdfd707b369116e02a08c3/sdk/typescript/src/knowledge-base.ts#L206-L224

The regex admits #\d+ and #x[\da-f]+ with no upper limit, so � and � both reach String.fromCodePoint and throw.

Affected version and environment

  • Released package: @openai/codex-security@0.1.1
  • Confirmed on current main at f22d4a36f26d16287bcdfd707b369116e02a08c3
  • macOS 26.5.2 (build 25F84)
  • Node.js v24.5.0
  • Bun 1.3.14

Steps to reproduce

Build a minimal .docx (a zip containing word/document.xml) whose body text contains an out-of-range numeric entity, then pass it to prepareKnowledgeBase:

const xml =
  '<?xml version="1.0" encoding="UTF-8"?>' +
  '<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">' +
  "<w:body><w:p><w:r><w:t>Boundary &#x110000; case.</w:t></w:r></w:p></w:body>" +
  "</w:document>";
await writeFile(path, zipSync({ "word/document.xml": strToU8(xml) }));
await prepareKnowledgeBase([path]);

Observed results across four documents that differ only in body text:

[ordinary text] prepared successfully

[valid entity &#65;] prepared successfully

[out-of-range &#x110000;] FAILED
   message: Cannot extract text from knowledge base DOCX: /.../threat-model.docx
   cause:   RangeError: Arguments contain a value that is out of range of code points

[huge decimal &#99999999999999;] FAILED
   message: Cannot extract text from knowledge base DOCX: /.../threat-model.docx
   cause:   RangeError: Arguments contain a value that is out of range of code points

Expected behavior

An entity that cannot represent a code point should be left as literal text (or dropped), the same way an unrecognized named entity is already returned unchanged by the entities[name.toLowerCase()] ?? entity fallback on the line above.

Actual behavior

The whole knowledge base fails to prepare, so --knowledge-base aborts the scan before Codex starts. The reported message — "Cannot extract text from knowledge base DOCX" — suggests the file is unreadable or corrupt, which sends the user looking in the wrong place; the actual RangeError is only visible via the error's cause.

Impact

Modest, and I want to be straightforward about it: Word itself does not emit out-of-range entities, so this is unlikely to fire on a document authored normally. It is reachable through documents produced by conversion tools or hand-edited XML, and by any deliberately malformed file. The consequence is disproportionate to the cause — a single stray entity anywhere in one document blocks the entire scan, including all the other knowledge base files that were fine.

Note that this failure mode is asymmetric with the surrounding code, which is otherwise careful: unsupported file types, symlinks and unreadable entries are all rejected with targeted errors rather than propagating a runtime exception.

Suggested direction

Clamp to the valid range and fall back to the literal entity text:

const codePoint = Number.parseInt(name.slice(hexadecimal ? 2 : 1), hexadecimal ? 16 : 10);
return Number.isFinite(codePoint) && codePoint <= 0x10ffff
  ? String.fromCodePoint(codePoint)
  : entity;

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read sdk/typescript/src/knowledge-base.ts around decodeXml, then trace how extractDocx is called by prepareKnowledgeBase. Reproduce the issue with the minimal DOCX described in the report and verify that out-of-range numeric entities remain literal text while valid entities still decode and the scan completes without aborting.

Written by the indexing model from the issue text.

Assessment

Tech stack
node.js, typescript
Domain
backend, cli
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
80/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.