crewAIInc / crewAIInc/crewAI

[BUG] input_files (PDFFile) are passed as base64 via read_file tool, causing context overflow and inconsistent LLM behavior

Open
#5,930 15 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug no-issue-activity
Dominant language
Python
Stars
58.8k
Forks
8.5k
Avg merge
1d 15h
Merged PRs (30d)
109

Description

Description

When using input_files with PDFFile (or File), CrewAI does not appear to handle the file as a native file input at the provider level.

Instead, the file is processed via the read_file tool and its content is returned as a binary/base64 representation. This content is then indirectly injected into the agent execution context.

As a result:

  • The PDF is effectively treated as inline binary data (base64)
  • The LLM context becomes extremely large
  • Responses become inconsistent or fail due to context overflow
  • The same file is re-processed during agent execution via tools

This makes PDFFile unreliable for large or even medium-sized documents.

Steps to Reproduce
  1. Create a minimal CrewAI setup:
from crewai import Agent, Task, Crew
from crewai_files import PDFFile

agent = Agent(
    role="Document Analyst",
    goal="Extract structured information from PDFs",
    backstory="Expert in document analysis",
    llm="gpt-4o-mini",
)

task = Task(
    description="""
    Read the PDF document {doc}
    Extract the main sections and summarize them precisely.
    """,
    expected_output="Structured list of sections",
    agent=agent,
)

crew = Crew(agents=[agent], tasks=[task], verbose=True)

result = crew.kickoff(
    input_files={
        "doc": PDFFile(source="./src/test_crewai_files/pdfs/sample.pdf")
    }
)

print(result)
  1. Run the flow:
    crewai run
  2. Observe agent execution logs with verbose=True
Expected behavior

The PDF should be:

  • either streamed or parsed externally before being sent to the LLM
  • or converted into structured text chunks (not raw base64)

The agent should receive:

  • structured text
  • or extracted segments
  • NOT a raw base64 PDF representation
Screenshots/Code snippets
from crewai import Agent, Task, Crew
from crewai_files import PDFFile

agent = Agent(
    role="Document Analyst",
    goal="Extract structured information from PDFs",
    backstory="Expert in document analysis",
    llm="gpt-4o-mini",
)

task = Task(
    description="""
    Read the PDF document {doc}
    Extract the main sections and summarize them precisely.
    """,
    expected_output="Structured list of sections",
    agent=agent,
)

crew = Crew(agents=[agent], tasks=[task], verbose=True)

result = crew.kickoff(
    input_files={
        "doc": PDFFile(source="./src/test_crewai_files/pdfs/sample.pdf")
    }
)

print(result)
Operating System

Ubuntu 24.04

Python Version

3.12

crewAI Version

v.1.14.4

crewAI Tools Version

v.1.14.4

Virtual Environment

Venv

Evidence

Verbose output evidence

Tool Execution Started (#3)

Tool: read_file
Args: {'file_name': 'doc'}

Tool Execution Completed (#3)

Tool Completed
Tool: read_file
Output: [Binary file: sample.pdf (application/pdf)]
Base64:
...
Possible Solution

None

Additional context

This issue blocks any production usage of PDF ingestion in multi-step CrewAI flows, because:

  • context size grows linearly with file size
  • multiple tasks re-trigger file expansion
  • sequential workflows amplify token explosion

A safer architecture would:

  • load file once
  • extract structured representation once
  • reuse extracted representation across tasks without re-injecting raw binary

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the flow using input_files with PDFFile and inspect the verbose read_file execution shown in the report. Trace how the file's output enters agent execution and compare it with the expected provider-level or structured-text handling; done means the PDF is not injected as raw base64 and repeated tasks do not re-expand it.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.