apache / apache/hudi

Agentic lakehouse: vector search through the SQL endpoint

Open
#19,261 0 comments 0 reactions 0 assignees View on GitHub
type:devtask
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

### Task Description

**What needs to be done:**
Surface vector/semantic search via the Trino Hudi connector (e.g. an ANN table function or similarity predicate), so one SQL surface serves structured, semantic, and full-text retrieval. Once available in SQL, the agent gateway picks it up through its existing guarded SQL tool (and potentially a dedicated semantic-search tool). Design spike first; RFC if it touches connector/format surfaces.

**Why this task is needed:**
Retrieval for AI workloads shouldn't require a second engine next to the lakehouse.

### Task Type
Other

### Related Issues
**Parent feature issue:** #19256

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by investigating the Trino Hudi connector and the agent gateway's existing guarded SQL tool, both named in the task. Define how vector or semantic search would be exposed through SQL and whether a dedicated tool is needed; produce the design spike and an RFC if connector or format surfaces are affected.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, sql
Domain
backend-api-design, data, databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.