binary-husky / binary-husky/gpt_academic

[Bug]: 使用PDF翻译功能出现大量“截断重试”

Open
#1,822 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
71.3k
Forks
8.3k
PR merge metrics
No merged PRs in 30d

Description

### Installation Method | 安装方法与平台

Pip Install (I used latest requirements.txt)

### Version | 版本

Latest | 最新版

### OS | 操作系统

Linux

### Describe the bug | 简述

在进行PDF翻译时成功被DOC2X读取,但在翻译过程中由于未知原因(可能是文章裁切)使得翻译过程出现大量截断重试,且每次截断重试都从头开始,如此循环往复消耗了我大量的token且没有正确的翻译结果。

### Screen Shot | 有帮助的截图

![微信截图_20240523141553](https://github.com/binary-husky/gpt_academic/assets/123653723/cf6131d8-ffd6-4d8e-916c-7fb7cb76f6c4)

### Terminal Traceback & Material to Help Reproduce Bugs | 终端traceback(如有) + 帮助我们复现的测试材料样本(如有)

_No response_

Contributor guide

No contributing guide indexed for this repository

Research direction

No source file, test, traceback, or reproduction document is provided. Start by reproducing PDF translation on Linux with the pip installation and latest requirements.txt, then inspect the translation retry logs to identify why truncation retries restart from the beginning. Done means a PDF completes translation without an endless retry loop or excessive token consumption.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.