testcontainers / testcontainers/testcontainers-java
LogMessageWaitStrategy misses the required log lines when multiple containers are started
Nobody has claimed this yet.
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 1.9k
- Avg merge
- 2d 17h
- Merged PRs (30d)
- 9
Description
I start a custom database container and then I run a second container which executes a bunch of SQL scripts to initialize the database (these scripts usually run in less than a second) and exits. For both of these containers I use LogMessageWaitStrategy to check that (a) the database is ready; (b) the initialization scripts finished running and the initialization was successful.
The problem: the second container almost always (but not absolutely every time) does not pass the startup check, though the required log line is printed to the log.
After a while I came up with a synthetic test case which does not depend on our custom images.
The Dockerfile:
FROM alpine
COPY init /
CMD ["/init"]
The init script:
#!/bin/sh
if [ "$HANG" = 'true' ]
then
echo 'I will start and work hard for a long while.'
else
echo 'I will perform some quick initialization and quit.'
fi
# Emulate some workload before the initialization finishes.
sleep 0.1
echo "Initialization finished."
# Hang to emulate some workload running after initialization.
[ "$HANG" = 'true' ] && sleep inf || true
After building the image with the tag test I run the following test case:
@Test
public void standaloneTest() {
new GenericContainer<>("test")
.withEnv("HANG", "true")
.waitingFor(new LogMessageWaitStrategy()
.withRegEx(".*Initialization finished.*"))
.start();
new GenericContainer<>("test")
.waitingFor(new LogMessageWaitStrategy()
.withRegEx(".*Initialization finished.*"))
.start();
}
The test hangs on the second start() for a while and then fails with Not ready yet exception, though I see the correct log lines reported by Testcontainers just before the exception (Log output from the failed container).
Few things to consider:
- a single container using the test image always passes the startup check successfully disregarding how quickly it starts and how quickly it exits;
- a problem appears only when both containers use
LogMessageWaitStrategy; - increasing the
sleepduration in the script decreases the chances of the issue:sleep 0.3makes the issue less frequent andsleep 1removes it almost completely; Thread.sleep(1000)inserted between thestart()s seems to be a workaround, but it is not reliable and the issue still reproduces, though rarely.
I use Testcontainers 1.14.3. The environment varies; on different machines (all running different flavors of Linux) the issue has a different chance to fire (it seems that the chances are higher on more powerful and fast hardware).
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with LogMessageWaitStrategy and reproduce standaloneTest using the supplied Dockerfile and init script, focusing on two sequential containers with different workloads. Trace how each strategy observes container logs and define done as both startup checks reliably detecting “Initialization finished.” without the intermittent timeout.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- docker, java, shell
- Domain
- devops, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 42/100