Metadata Store
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
Hướng nghiên cứu
Không có tệp hoặc bài kiểm thử nào trong repository được nêu tên. Hãy bắt đầu bằng cách đọc phạm vi v1 và epic OpenLineage được liên kết (#856), sau đó so sánh các đường dẫn API được đề xuất và những mối quan ngại về việc lưu trữ PostgreSQL được mô tả ở đây. Công việc được xem là hoàn tất khi thiết kế dịch vụ MVP, các ranh giới API được hỗ trợ, phương án lưu trữ sự kiện và phạm vi triển khai dựa trên Helm đã được thống nhất.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Stackable Metadata Store - v1
[!NOTE]
Yes, I'm aware the name is dull and technical. Anyone with a better idea is welcome to propose one.
[!IMPORTANT]
This is very much under construction
SDP deploys a dozen or more products, some of which hold data and metadata worth knowing about. We want one service that collects as much of the platform's metadata as it can, stores it, and makes it available to other services such as our Cockpit UI.
Scope
[!NOTE]
We have many ideas to extend this later. For our first version, this section describes the scope of what we want to achieve.
A read-only service that ingests OpenLineage events, stores them, and serves one or two clients/APIs:
- Cockpit probably via a custom (but documented and usable by others) API
- Stretch/Optional: OpenLineage out as well (so it can function as some kind of proxy)
- Investigate this briefly as I do not want us to offer a Push API as part of this. If that is the only way for OpenLineage to gather data.
Technical details
- Unless there are really really good arguments otherwise I'd like this to be a Rust based service, using Axum and PostgreSQL as the sole supported backing store.
- I don't know enough about this but from what I understand we probably don't need an operator for this: A Helm chart should be enough
- The "URL space" API should be ours and it should be designed so we can extend it later with more APIs to ingest and export.
/ingest/v1/openlineage, /ingest/v1/<foosource>, /v1/ingest/openlineage or /openlineage are all up for discussion. Two things to keep in mind: for OpenLineage there will be ingest and output at the same time, and it is an open question whether our own API and the external ones share a prefix or get separate ones (/api/ and /ingest/openlineage, or all the same).
No layout will be perfect, so pick one rather than overthinking it. Especially for this very first version. We will mark it experimental so we can learn a bit.
All tools we support can configure the path, so we can pick whatever. The default some of them use is /api/v1/lineage.
Just remember that each of these APIs might behave slightly different also with authentication & authorisation.
OpenLineage - Technical details
OpenLineage is complex and troublesome. Also see our OpenLineage epic and talk to the colleagues if it helps.
It emits several kinds of event and we need to decite how we want to store which.
Here is one example (this might not be 100% correct, AI generated based on it reading the code)
Apache Spark reads raw.orders and writes the Iceberg table sales.orders_daily through a Hive catalog:
{
"eventType": "COMPLETE",
"job": {
"namespace": "prod-lineage",
"name": "orders_etl.execute_replace_table_as_select_command"
},
"run": { "runId": "c7f2…-01b8", "facets": {} },
"inputs": [
{ "namespace": "s3a://warehouse", "name": "raw.db/orders" }
],
"outputs": [{
"namespace": "s3a://warehouse",
"name": "sales.db/orders_daily",
"facets": {
"symlinks": { "identifiers": [
{ "namespace": "hive://hive-metastore:9083",
"name": "sales.orders_daily",
"type": "TABLE" }
]}
}
}]
}
Trino then queries that same table through catalog lakehouse and writes sales.orders_summary:
{
"eventType": "COMPLETE",
"job": {
"namespace": "prod-lineage",
"name": "20260902_101500_00042_ab3kd" // default $QUERY_ID
},
"run": { "runId": "9d13…-77e0" },
"inputs": [{
"namespace": "trino://trino-coordinator-default-0.trino-coordinator-default.prod.svc.cluster.local:8443",
"name": "lakehouse.sales.orders_daily"
}],
"outputs": [{
"namespace": "trino://trino-coordinator-default-0.…:8443",
"name": "lakehouse.sales.orders_summary"
}]
}
Both emitters are configured with the same OpenLineage namespace, and both describe the same physical table. Unfortunately, they do not agree on what it is called:
| Emitted by | Namespace | Name |
|---|---|---|
| Spark, primary | s3a://warehouse |
sales.db/orders_daily |
| Spark, symlink | hive://hive-metastore:9083 |
sales.orders_daily |
| Trino | trino://trino-coordinator-default-0.…:8443 |
lakehouse.sales.orders_daily |
- The scheme differs.
s3aagainsthiveagainsttrino: physical storage, catalog service, query engine. All three correct, but...annoying for us. - The separator differs. Spark writes
sales.db/orders_dailyfrom the warehouse path, its own symlink writessales.orders_daily, Trino writes dots throughout. - Trino prefixes its catalog name.
lakehouse.is a Trino-side concept with no counterpart anywhere in Spark's output. You cannot resolve it without knowing which catalog points at which metastore.
So, even the symlinks thing doesn’t help us. It gives us a hive:// when we need to match a trino://
In this MVP we will not be able to resolve this and it’s out of scope too. We should be able to do clever things later because the information is in TrinoCatalog and so on AND we have the source code for all those OpenLineage things under control so we’ll have to tweak this later but for now this matching between products etc. is out of scope.
That means (namespace, name) can not be the identity of a dataset. It can be used for lookup only. We’ll later have to derive unique/stable keys for datasets somehow.
Why do I have this section here? Because I have a feeling that this is important for our database schema :)
Out of scope for all of it
- A write path: We'll want this later, but not for this.
- Any other APIs than the ones I mentioned above (Cockpit/OpenLineage)
- Any kind of Web UI -> all of that lives in Cockpit
- Any kind of pruning of old data. The store grows without limit for now.
- Ngôn ngữ chính
- Không có dữ liệu ngôn ngữ
- Star
- 2
- Fork
- 0
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của stackabletech/issues
-
Release Retro 26.11.0 Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 50/100
stackabletech/issues#890 ·
-
tracking: SDP Release 26.11.0 Đang mởepic
stackabletech/issues#889 · 2 người được giao ·
-
stackabletech/issues#888 · 1 bình luận · 1 người được giao ·
-
stackabletech/issues#887 · 1 bình luận · 1 người được giao ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 25/100
stackabletech/issues#886 ·
Tất cả issue của stackabletech/issues
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
avniproject/avni-client#2135 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
use-agent-os/agent-os#3276 ·
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
-
good first issue refactor
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100