Observability: metrics, logging, and tracing
- Vorherrschende Sprache
- Go
- Sterne
- 4
- Forks
- 0
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
## Summary
Implement observability infrastructure for apex: Prometheus metrics, structured logging, and OpenTelemetry tracing.
## Metrics
Instrument all subsystems with OTel metrics exposed via Prometheus endpoint.
### Sync engine
- `apex_sync_head` (gauge) — last synced height
- `apex_sync_network_head` (gauge) — upstream network head
- `apex_sync_lag_seconds` (gauge) — time behind network head
- `apex_sync_backfill_duration` (histogram) — per-batch backfill latency
- `apex_sync_errors_total` (counter) — sync errors by type
- `apex_sync_backfill_progress_pct` (gauge) — backfill completion percentage
### API
- `apex_rpc_request_duration` (histogram) — per-method latency
- `apex_rpc_request_total` (counter) — per-method call count
- `apex_rpc_errors_total` (counter) — per-method errors
### Store
- `apex_store_query_duration` (histogram) — SQLite query latency
- `apex_store_insert_duration` (histogram) — insert latency
- `apex_store_size_bytes` (gauge) — DB file size
### Subscriptions
- `apex_subscriptions_active` (gauge) — active subscription count
- `apex_subscription_deliveries` (counter) — messages delivered
- `apex_subscription_drops` (counter) — messages dropped (slow reader)
### Node
- `apex_build_info` (gauge) — with version labels
- `apex_uptime_seconds` (counter)
## Logging
- Use `slog` (Go stdlib) — no heavy dependencies like `ipfs/go-log`
- Structured key-value pairs: height, namespace, method, duration, error
- Configurable log level at startup (and ideally at runtime via admin endpoint)
## Tracing
- OpenTelemetry spans for sync fetch, store operations, and RPC handlers
- Configurable exporter (stdout for dev, OTLP for production)
## API endpoint
Expose a unified `SyncStatus()` endpoint returning current synced height, network head, sync state (backfilling/streaming), lag, and active subscriptions in one call. celestia-node spreads this across 4 separate modules.
## Configuration
```toml
[observability]
metrics_address = "0.0.0.0:9090" # Prometheus endpoint
log_level = "info" # debug, info, warn, error
tracing_enabled = false
tracing_endpoint = "" # OTLP endpoint
```
Beitragsleitfaden
Rechercherichtung
Beginnen Sie damit, die Einstiegspunkte der Sync-Engine, der API, des Stores, der Subscriptions, des Nodes und der Konfiguration abzubilden, und entscheiden Sie dann, wie die Anforderungen an Metriken, slog-Logging und Tracing zusammenpassen. Überprüfen Sie den Prometheus-Endpunkt, die SyncStatus-Antwort, die TOML-Konfiguration sowie jede aufgeführte Metrik und jeden aufgeführten Span; in der Issue werden keine Dateien oder Tests genannt, daher müssen die Testorte während der Recherche ermittelt werden.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- go, prometheus, sqlite
- Bereich
- api, backend, databases, observability
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 25/100