neo4j / neo4j/neo4j-graphrag-python

[BUG]: schema extraction uses the whole file without splitting in chunks

Open
#457 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
1.3k
Forks
246
Avg merge
1d 8h
Merged PRs (30d)
11

Description

Before You Report a Bug, Please Confirm You Have Done The Following...
  • I have updated to the latest version of the packages.
  • I have searched for both existing issues and closed issues and found none that matched my issue.
neo4j-graphrag-python's version

1.10.1

Python version

3.12

Operating System

Debian 13

Dependencies
"datasets==3.6.0",
"flask>=3.1.2",
"neo4j>=5.28.2",
"neo4j-graphrag[nlp,ollama,sentence-transformers]>=1.10.1",
"streamlit>=1.52.1",
Reproducible example
PDF_FILE = './some-file.pdf' # a PDF file of big size -- bigger than the model context

kg_builder = SimpleKGPipeline(
    llm=llm,
    driver=neo4j_driver,
    embedder=embedder,
    from_pdf=True,
    text_splitter=text_splitter,
)
await kg_builder.run_async(file_path=PDF_FILE)
Relevant Log Output

JSONDecoder error stating that the output is not in JSON format

Expected Result

I expect the pipeline to finish

What happened instead?

In the schema extraction phase, the pipeline does not split the document(s) in chunks; rather it gives the whole document, despite its size, to the ollama endpoint in a single prompt.

Additional Info

What happens is that the pipeline expects to extract the schema with a single request to OLLAMA giving it the whole document. Instead, it should perform this step chunk-by-chunk.
By giving the whole document in the prompt, the model loses track of the instructions.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with SimpleKGPipeline.run_async and follow the schema extraction path when from_pdf=True and a text_splitter is supplied. Reproduce with a PDF larger than the model context using the reported Ollama setup, then verify schema extraction processes chunks rather than sending the whole document in one prompt and that the pipeline finishes without a JSONDecoder error.

Written by the indexing model from the issue text.

Assessment

Tech stack
ollama, python
Domain
ai, backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.