binary-husky / binary-husky/gpt_academic

Security: transformers cap leaves affected versions reachable in local model bridges

Open
#2,289 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
71.3k
Forks
8.3k
PR merge metrics
No merged PRs in 30d

Description

## Summary

The current dependency cap allows affected `transformers` versions to remain installable from the main requirements file:

- Current dependency line: https://github.com/binary-husky/gpt_academic/blob/d6bde0fa54373309bd05823a49bda8da019d2c77/requirements.txt#L11
- Declared range: `transformers>=4.27.1,<4.42`
- Resolver result checked from that range: `transformers==4.41.2`

That resolved version is below the patched versions for these public advisories:

- CVE-2024-11394 / GHSA-hxxf-235m-72v3, patched in Transformers 4.48.0: https://github.com/advisories/GHSA-hxxf-235m-72v3
- CVE-2025-3264 / GHSA-jjph-296x-mrcr, patched in Transformers 4.51.0: https://github.com/advisories/GHSA-jjph-296x-mrcr

## Source paths to review

The project has several local model bridges that load Transformers models or tokenizers with `trust_remote_code=True`, for example:

- https://github.com/binary-husky/gpt_academic/blob/d6bde0fa54373309bd05823a49bda8da019d2c77/request_llms/bridge_chatglm.py#L36
- https://github.com/binary-husky/gpt_academic/blob/d6bde0fa54373309bd05823a49bda8da019d2c77/request_llms/bridge_chatglm3.py#L30
- https://github.com/binary-husky/gpt_academic/blob/d6bde0fa54373309bd05823a49bda8da019d2c77/request_llms/bridge_chatglm3.py#L35
- https://github.com/binary-husky/gpt_academic/blob/d6bde0fa54373309bd05823a49bda8da019d2c77/request_llms/bridge_chatglm4.py#L31
- https://github.com/binary-husky/gpt_academic/blob/d6bde0fa54373309bd05823a49bda8da019d2c77/request_llms/bridge_chatglm4.py#L36
- https://github.com/binary-husky/gpt_academic/blob/d6bde0fa54373309bd05823a49bda8da019d2c77/request_llms/bridge_deepseekcoder.py#L39

## Local validation notes

With `transformers==4.41.2`, I reproduced the library-level behavior behind the two advisories in a local test environment:

- For GHSA-hxxf-235m-72v3, the TransfoXL tokenizer pickle path was blocked in the default mode, but reached a synthetic local marker payload when trusted remote-code mode was enabled.
- For GHSA-jjph-296x-mrcr, `get_imports()` showed rapidly increasing runtime on crafted input: about 0.037s at size 100, 0.808s at size 500, 2.835s at size 1000, and timeout at size 2000 in the local bounded probe.

This report is scoped to the vulnerable dependency cap plus the current source paths that opt into trusted Transformers model code. It is not a claim of a full unauthenticated application exploit.

## Suggested fix

Please consider removing the `<4.42` cap and requiring a Transformers version at or above 4.51.0 if compatible, or another current compatible release line that includes the fixes above. The optional ChatGLM4 requirements file already has a looser `transformers>=4.44` line, so it may also be worth aligning the main requirements with the maintained path.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with requirements.txt and the listed request_llms/bridge_chatglm.py, bridge_chatglm3.py, bridge_chatglm4.py, and bridge_deepseekcoder.py paths, focusing on their Transformers usage and trusted remote-code settings. Compare the main dependency range with the optional ChatGLM4 requirements, reproduce the resolver result and the two local validation probes described in the report, and confirm that a compatible maintained version addresses both advisories.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, machine-learning, security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
58/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.