google / google/adk-go

Bug: Memory search fails to match words with punctuation in extractWords

Open
#569 5 comments 0 reactions 0 assignees View on GitHub
bug help wanted
Dominant language
Go
Stars
8.8k
Forks
1k
Avg merge
3d 18h
Merged PRs (30d)
88

Description

## Bug: Memory search fails to match words with punctuation

### Summary
The `extractWords` function in `memory/inmemory.go` uses space-only splitting and doesn't strip punctuation, causing memory search to miss relevant results when text contains punctuation or non-space whitespace.

### Reproduction
**Scenario**: Add a session with content containing punctuation, then search for a word without punctuation.

```go
// Add session with punctuation
session := makeSession("app1", "user1", "sess1", []*session.Event{
{
LLMResponse: model.LLMResponse{
Content: genai.NewContentFromText("The agent works great!", genai.RoleModel),
},
},
})
memSvc.AddSession(ctx, session)

// Search for "great" (without punctuation)
resp, _ := memSvc.Search(ctx, &memory.SearchRequest{
AppName: "app1",
UserID: "user1",
Query: "great",
})

// Expected: 1 memory found
// Actual: 0 memories (because stored token is "great!" not "great")
```

### Root Cause
**File**: `memory/inmemory.go`, line ~155

```go
func extractWords(text string) map[string]struct{} {
res := make(map[string]struct{})

for s := range strings.SplitSeq(text, " ") { // ← Only splits on space
if s == "" {
continue
}
res[strings.ToLower(s)] = struct{}{} // ← Doesn't strip punctuation
}

return res
}
```

**Issues**:
1. **Space-only splitting**: `strings.SplitSeq(text, " ")` doesn't handle tabs, newlines, or multiple spaces
2. **No punctuation normalization**: `"great!"` is stored as-is, won't match `"great"`
3. **Case sensitivity handled but not enough**: Lowercasing happens after punctuation is included

### Impact
- **Search accuracy degraded**: Users searching for "error" won't find memories containing "error." or "error," or "error!"
- **Common patterns affected**:
- Sentences ending with punctuation (`.`, `!`, `?`)
- Comma-separated lists
- Quoted text
- Multi-line responses with `\n` or `\t`

### Proposed Fix
Replace space-only splitting with proper whitespace tokenization and strip punctuation:

```go
func extractWords(text string) map[string]struct{} {
res := make(map[string]struct{})

for _, word := range strings.Fields(text) { // Splits on all whitespace
// Strip punctuation
cleaned := strings.TrimFunc(word, func(r rune) bool {
return !unicode.IsLetter(r) && !unicode.IsNumber(r)
})
if cleaned == "" {
continue
}
res[strings.ToLower(cleaned)] = struct{}{}
}

return res
}
```

**Alternative**: Use a proper tokenizer/stemmer for production-grade search, but the above fix would resolve the immediate issue.

### Test Case to Add
```go
{
name: "match words with punctuation",
initSessions: []session.Session{
makeSession(t, "app1", "user1", "sess1", []*session.Event{
{
LLMResponse: model.LLMResponse{
Content: genai.NewContentFromText("Error: connection timeout! Please retry.", genai.RoleModel),
},
},
}),
},
req: &memory.SearchRequest{
AppName: "app1",
UserID: "user1",
Query: "error timeout retry", // No punctuation
},
wantResp: &memory.SearchResponse{
Memories: []memory.Entry{/* should find the memory */},
},
},
```

### Environment
- **Version**: `main` branch (commit: latest as of 2026-02-16)
- **Go version**: 1.22+

### Additional Context
This is particularly problematic for AI agent memory since LLM responses naturally contain punctuation. The current implementation significantly reduces search recall in real-world usage.

Contributor guide

Open the contributing guide

Research direction

Start in memory/inmemory.go at extractWords and trace how its tokens are used by memory search. Add the described regression case for punctuation and non-space whitespace, then run the existing memory search tests. Done means searches for unpunctuated words match content containing punctuation and whitespace-separated terms.

Written by the indexing model from the issue text.

Assessment

Tech stack
go
Domain
backend, search
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.