spring-projects / spring-projects/spring-batch

Enhance RedisItemReader with batchSize option to optimize N+1 problem via MGET operations

Open
#4,941 4 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

status: waiting-for-triage type: feature
Dominant language
Java
Stars
3k
Forks
2.5k
Avg merge
6d 53m
Merged PRs (30d)
3

Description

Summary

The current RedisItemReader implementation suffers from the classic N+1 problem, executing individual GET commands for each key during batch processing. This creates significant performance bottlenecks due to excessive network round-trips when processing large datasets.

For example, processing 1,000 Redis keys results in 1,000 separate network calls, causing:

  • High latency due to network round-trip time (RTT) multiplication
  • Excessive load on Redis server
  • Poor performance scaling with dataset size

Problem Description

Current Implementation Issue:

@Override
public V read() throws Exception {
   if (this.cursor.hasNext()) {
       K nextKey = this.cursor.next();
       return this.redisTemplate.opsForValue().get(nextKey); // 💀 Individual GET call for each key
   }
   return null;
}
Performance Impact:
  • 1,000 keys → 1,000 network round-trips
  • Total processing time = Network latency × Key count
  • Redis server experiences request flooding
Proposed Solution

Add a batchSize parameter to RedisItemReader and RedisItemReaderBuilder that enables batch processing via Redis MGET operations while maintaining complete backward compatibility.

Key Features:

  • Default batchSize = 1 preserves existing behavior (100% backward compatible)
  • batchSize > 1 enables optimized batch processing using Redis MGET
  • Internal buffering via queue for efficient data handling
  • Significant performance improvement with configurable memory usage

Expected Performance Improvement:

  • 1,000 keys with batchSize = 100: 90% reduction in network calls (10 vs 1,000)
  • 1,000 keys with batchSize = 1000: 99% reduction in network calls (1 vs 1,000)

Implementation Approach
Enhanced RedisItemReader:

  • Add batchSize field with default value of 1
  • Implement conditional logic: single-key mode vs batch mode
  • Use internal Queue for buffering batch results
  • Leverage RedisTemplate.opsForValue().multiGet() for batch operations

Enhanced RedisItemReaderBuilder:

  • Add batchSize(int batchSize) method with validation
  • Pass batchSize to RedisItemReader constructor
  • Comprehensive JavaDoc documentation
Usage Examples:
// Backward compatible (existing behavior)
RedisItemReader<String, Object> reader = new RedisItemReaderBuilder<String, Object>()
    .redisTemplate(template)
    .scanOptions(scanOptions)
    .build(); // Uses batchSize = 1

// Optimized batch processing
RedisItemReader<String, Object> optimizedReader = new RedisItemReaderBuilder<String, Object>()
    .redisTemplate(template)
    .scanOptions(scanOptions)
    .batchSize(100) // Process 100 keys per MGET
    .build();
Benefits
  • Performance Enhancement: Dramatic reduction in network overhead
  • Backward Compatibility: Zero breaking changes to existing code
  • Configurable Optimization: Adjustable batch size based on requirements
  • Memory Efficiency: Controlled memory usage through batch size limits
  • Spring Framework Alignment: Follows Spring's progressive enhancement philosophy

Technical Considerations

Memory Management:

  • Batch size validation (1-1000 range) prevents OOM issues
  • Internal queue management with proper cleanup
  • Null value filtering for memory efficiency

Redis Limitations Awareness:

  • Maintains existing limitations (no restart capability due to SCAN nature)
  • Preserves idempotency requirements
  • Handles duplicate key scenarios appropriately

Additional Context
This enhancement addresses a fundamental performance issue in Spring Batch's Redis integration while maintaining the framework's commitment to backward compatibility. The implementation follows Spring's architectural principles and provides a foundation for future Redis-related optimizations.

Related Issues:

  • Classic N+1 problem pattern in ORM and data access layers
  • Network optimization in distributed systems
  • Batch processing performance in Spring Batch

Environment:

  • Spring Batch version: 5.1+
  • Spring Data Redis: 3.x+
  • Java: 17+

"Terminated by KILL-9 SQUAD 💀"

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating RedisItemReader and RedisItemReaderBuilder, then trace how read() currently obtains keys through redisTemplate.opsForValue().get(). Implement the described batchSize behavior with multiGet and buffering, preserving batchSize=1 behavior; done means validated batch-size input, backward-compatible reads, and coverage for single-key and batched reads.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, redis, spring
Domain
backend, databases
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.