opensearch-project / opensearch-project/data-prepper
Allow for sync from AWS Glue Catalog to AWS OpenSearch
Nobody has claimed this yet.
- Dominant language
- Java
- Stars
- 374
- Forks
- 354
- Avg merge
- 3d 18h
- Merged PRs (30d)
- 8
Description
Is your feature request related to a problem? Please describe.
No, it is a new feature.
Describe the solution you'd like
Currently, the S3 sink to ingest data from S3 requires fixed prefixes and schemas. However larger enterprises have data lakes for large scale data sets (~800M records daily per data set in my case). In AWS, Glue keeps track of all the catalog tables and metadata. If we could set up our catalog as a source and maintain sync, we can remove unnecessary infrastructure and complexity in the indexing process at large scale.
As a source, a user should be able to provide Data Prepper with:
- Glue database name
- Glue table name
- Primary key fields of the table
- Shard count
- Replica count
And should then follow a process similar to this:
Plant UML code for reference:
@startuml
skinparam maxMessageSize 150
autonumber
participant "OSIS" as osi
participant "Glue API" as ga
participant "OpenSearch" as os
participant "S3" as s3
osi -> ga: Get table metadata
osi -> os: Get all indices in the alias
osi -> osi: Compare indices in alias vs table partitions
loop For every missing index
osi -> os: Create index
osi -> os: Set refresh interval to -1
osi -> os: Add index to alias
osi -> s3: Get the data from the partition
osi -> os: Index the data
osi -> os: Set refresh interval back to original value
end
loop For every partition
osi -> s3: Check row count on the partition
osi -> os: Check record count in index
osi -> os: If record mismatch, set refresh interval to -1
osi -> s3: If record mismatch, get the data from the partition
osi -> os: If record mismatch, index the data
osi -> os: If record mismatch, set refresh interval back to original value
osi -> os: Purge deleted records
end
@enduml
As a result, in the OS cluster, we will have:
- An alias called
<database_name>_<table_name> - If the table is partitioned, indices that point to the alias for each partition in this format:
<database_name>_<table_name>_<partition1_value>_<partition2_value>_<partitionN_value>
Describe alternatives you've considered (Optional)
I currently do this:
https://github.com/aws-samples/aws-s3-to-opensearch-pipeline
It's fast and it works, but an out of the box solution would be preferable.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue does not name repository files or tests. Start by mapping the existing S3 sink and the proposed AWS Glue, S3, and OpenSearch flow, then clarify the design and integration points before implementation. Done would include configurable Glue database and table metadata, partition synchronization, and the described OpenSearch alias and index behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, java
- Domain
- backend, cloud, data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100