AnswerDotAI / AnswerDotAI/byaldi

index() corruption

Open
#65 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
851
Forks
91
PR merge metrics
No merged PRs in 30d

Description

Hi,
I have been trying to run the indexing on a set of 80 pdf documents (~150 pages each) by submitting batch jobs. Since the indexing took longer than expected (8 hours) my session ended abruptly and I get a "ValueError: Expected object or value" when I try to read from_index().

I don't see any method to discard the partially indexed document and continue from the last valid index. This would mean I need to start from the top for another 8+ hours. Is it possible to have some functionality to deal with this situation?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.