apache / apache/stormcrawler

Tika ParserBolt puts no limit on the extracted text and no time limit on the parse

Open
#2,097 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
995
Forks
292
Avg merge
2d 49m
Merged PRs (30d)
62

Description

## What happens
`ParserBolt.execute()` builds its text handler as `new BodyContentHandler(-1)`, which removes the 100000 character limit Tika applies by default, and there is no configuration key to set one. The resulting text is materialised as a single `String`, kept in the `ParseData` and emitted in the tuple alongside the original bytes. The parse itself runs on the executor thread with no timeout, so a document that takes a long time inside a parser holds that thread for as long as it takes.

## Where
`external/tika/src/main/java/org/apache/stormcrawler/tika/ParserBolt.java:216` and `:246`.

```java
ContentHandler textHandler = new BodyContentHandler(-1);
...
tika.getParser().parse(bais, teeHandler, md, parseContext);
text = textHandler.toString();
```

## Why it matters
The only lever an operator has over the size of a parsed document is `http.content.limit`, which defaults to -1 (no limit) in `crawler-default.yaml`. A large document therefore reaches the bolt whole and is held twice, once as bytes and once as text, until the tuple is emitted. Several fetcher threads doing this at the same time is enough to exhaust a worker heap, and the tuples get replayed after the restart. The missing timeout matters less often, but Tika's own guidance for input you do not control is to parse under one.

A caveat on the obvious reading of this code: Tika does apply a compression-ratio check of its own. In 3.3.2 that guard sits inside the container parsers, so a deflated document that expands past the ratio is aborted before this bolt ever sees the text. That is library behaviour and is not covered by the test below, and it does nothing for a plain uncompressed body. What is missing on the StormCrawler side is the character cap and the timeout.

## Reproduction

Save as `external/tika/src/test/java/org/apache/stormcrawler/tika/ParserBoltTextLimitTest.java`.

```java
/*
* Licensed to the Apache Software Foundation (ASF) under one or more
* contributor license agreements. See the NOTICE file distributed with
* this work for additional information regarding copyright ownership.
* The ASF licenses this file to you under the Apache License, Version 2.0
* (the "License"); you may not use this file except in compliance with
* the License. You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/

package org.apache.stormcrawler.tika;

import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
import org.apache.storm.task.OutputCollector;
import org.apache.stormcrawler.Metadata;
import org.apache.stormcrawler.TestUtil;
import org.apache.stormcrawler.parse.ParsingTester;
import org.junit.jupiter.api.Assertions;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;

/**
* ParserBolt builds its text handler with BodyContentHandler(-1), which removes the 100000
* character limit Tika applies by default, and there is no configuration key to put a limit back.
* The whole extracted text is held as one String and emitted in the tuple.
*/
class ParserBoltTextLimitTest extends ParsingTester {

@BeforeEach
void setupParserBolt() {
bolt = new ParserBolt();
setupParserBolt(bolt);
}

@Test
void extractedTextIsBounded() throws IOException {
Map conf = new HashMap<>();
bolt.prepare(conf, TestUtil.getMockedTopologyContext(), new OutputCollector(output));

byte[] content = "word ".repeat(1_000_000).getBytes(StandardCharsets.UTF_8);
parse("https://example.org/big.txt", content, new Metadata());

List> emitted = output.getEmitted();
Assertions.assertEquals(1, emitted.size());
int textLength = emitted.get(0).get(3).toString().length();
System.out.println(
"input bytes: " + content.length + " -> extracted characters: " + textLength);

// Tika's own default for BodyContentHandler is 100000 characters
Assertions.assertTrue(
textLength <= 100_000,
"extracted text should be capped, got " + textLength + " characters");
}
}
```

Run it:

```
mvn -pl external/tika test -Dtest=ParserBoltTextLimitTest
```

It fails on main and becomes the regression test once a cap exists. It feeds a 5 MB text body and asserts the extracted text is capped at Tika's own default of 100000 characters.

```
[INFO] Running org.apache.stormcrawler.tika.ParserBoltTextLimitTest
input bytes: 5000000 -> extracted characters: 5000001
[ERROR] ParserBoltTextLimitTest.extractedTextIsBounded -- Time elapsed: 1.324 s <<< FAILURE!
org.opentest4j.AssertionFailedError: extracted text should be capped, got 5000001 characters ==> expected: but was:
at org.apache.stormcrawler.tika.ParserBoltTextLimitTest.extractedTextIsBounded(ParserBoltTextLimitTest.java:61)
```

## Suggested fix
Read a limit in `ParserBolt.prepare()`, for example `parser.text.maxlength`, and pass it to `BodyContentHandler` in `execute()` instead of -1. Pick a default that is generous but finite and record it in the module documentation. Add a second key for a parse timeout and run `tika.getParser().parse()` on a bounded executor so the bolt can fail the tuple instead of holding the thread. Both changes are visible to operators: documents longer than the cap will be truncated where they were not before, so the default belongs in the release notes.

Contributor guide

Open the contributing guide

Research direction

Start with external/tika/src/main/java/org/apache/stormcrawler/tika/ParserBolt.java at the prepare and execute locations cited in the issue, then run mvn -pl external/tika test -Dtest=ParserBoltTextLimitTest. Done means extracted text is bounded by a configurable finite limit, parsing has a configurable timeout, and the relevant module documentation and release notes describe the operator-visible behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.