agent-substrate / agent-substrate/substrate
OpenTelemetry Application Telemetry Data Loss and Delivery Delays Across Actor Lifecycle Events
- Lenguaje dominante
- Go
- Estrellas
- 1.8k
- Forks
- 316
- Merge medio
- 2 d 43 min
- PR fusionados (30 d)
- 287
Descripción
---
## 1. Issue Background & Context
User applications running as Agent Substrate actors frequently emit custom metrics and distributed trace spans using OpenTelemetry (OTel) push exporters (`otlpmetricgrpc`, `otlptracegrpc`, or `otlpmetrichttp`).
In typical production workloads:
- **OTel Push Exporter Interval:** e.g., 60 seconds (standard batch export timer).
- **Active Execution Duration Before Idle Suspension:** e.g., 30 seconds (active request handling window before an idle timeout or on-demand pause occurs).
When the push exporter interval (e.g., 60s) is longer than the active execution duration before suspension (e.g., 30s), telemetry events generated during the active phase may remain buffered in the OTel SDK's in-memory queues (`PeriodicReader` accumulators and `BatchSpanProcessor` ring buffers).
This issue summarizes the core challenges, empirical findings across Substrate snapshot modes, short-term mitigations for actor authors, and proposed long-term system-level solutions.
---
## 2. Identified Telemetry Issues Across Actor Lifecycle Events
Through architectural analysis and lifecycle benchmarking, five primary telemetry issues were identified:
### Issue 1: In-RAM Telemetry Loss in Data Snapshot Mode (`onPause: Data`) — **HIGH**
- **Root Cause:** When `snapshotsConfig.onPause` is set to `Data`, Substrate excludes process memory (RAM) from the snapshot and saves only attached `DurableDir` disk volumes. The container process is killed upon pause.
- **Impact:** **100% loss of pre-pause in-memory telemetry**. All metrics and trace spans generated during the 30s before pause that were awaiting the 60s push timer are permanently lost when the process is killed.
### Issue 2: Telemetry Abandonment for Un-Resumed Suspended Actors — **HIGH**
- **Root Cause:** In Full Snapshot Mode (`onPause: Full`), process RAM is saved to snapshot storage. However, if an actor is suspended and subsequently **deleted or garbage-collected** (`kubectl ate delete actor`) without ever being resumed, the buffered in-memory metrics remain stuck in the snapshot file.
- **Impact:** Un-pushed telemetry is never delivered to the OTel Collector.
### Issue 3: Delivery Delays & Wall-Clock Timer Expiration — **MEDIUM**
- **Root Cause:** In Full Snapshot Mode (`onPause: Full`), when an actor is suspended for 5 minutes (from $t = 30\text{s}$ to $t = 330\text{s}$), wall-clock time advances on the host machine.
- **Behavior:** Upon resume at $t = 330\text{s}$, the language runtime (Go/Python/Node) evaluates the 60s timer deadline ($D = 60\text{s}$). Because $330\text{s} > 60\text{s}$, the timer expires **immediately upon resume at $t = 330\text{s}$**, flushing all pre-suspend and post-resume telemetry in a single batch.
- **Impact:** Delivery to the OTel backend is delayed by the suspension duration, though original generation timestamps ($t = 10\text{s}$) remain accurate.
### Issue 4: Backend Retention Window Expiration & Stale Timestamp Rejection — **HIGH**
- **Root Cause:** If an actor remains suspended for hours or days before being resumed, when it wakes up, it flushes telemetry with timestamps from hours/days ago.
- **Impact:** Telemetry backends (Prometheus, Cloud Monitoring, Datadog) **drop incoming samples older than their allowed ingestion window** (e.g., Prometheus drops samples > 2 hours old).
### Issue 5: Telemetry Loss on Sudden Actor Termination (Issue #23) — **HIGH**
- **Root Cause:** When an active or suspended actor is explicitly terminated (e.g., via `kubectl ate delete actor`, container eviction, or TTL cleanup), Substrate sends a `SIGTERM` / `SIGKILL` to the workload container.
- **Impact:** If the application process does not catch `SIGTERM` and call `meterProvider.ForceFlush(ctx)` / `tracerProvider.ForceFlush(ctx)` during the graceful termination period (or if killed abruptly via `SIGKILL`), **100% of un-flushed in-memory metrics and spans are lost** upon container destruction.
---
## 3. Empirical Test Matrix
Empirical test suite results across lifecycle scenarios:
| Test Scenario | Snapshot Mode (`snapshotsConfig`) | In-Memory Telemetry Loss? | Delivery Delay | Socket Recovery | Test Results |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **Suspend & Resume** | `onPause: Full` | **~0% Loss under standard resume conditions** | Delayed by suspend duration; fires **immediately at resume** ($t=330\text{s}$) | Automatic re-dial | Telemetry in RAM preserved across snapshot; delivery delayed by suspension duration until immediate post-resume timer flush. |
| **Pod Migration** | `onPause: Full` | **~0% Loss under standard resume conditions** | Delayed by migration duration | Automatic re-dial | Telemetry in RAM preserved; OTLP gRPC/HTTP exporter automatically re-dials new worker pod IP post-migration. |
| **Suspend & Resume** | `onPause: Data` | **100% Loss** of pre-pause buffer | N/A (Data lost on process kill) | N/A | 100% loss of pre-pause in-memory telemetry because process RAM is excluded from snapshot and container is killed. |
| **Data Mode + Disk Replay** | `onPause: Data` + `DurableDir` | **~0% Loss** (Persisted to disk) | Flushed post-resume from disk | Automatic re-dial | 0% loss; telemetry persisted to DurableDir volume survives process kill and is replayed to collector upon startup. |
| **Actor Termination** | Any (`Full` or `Data`) | **100% Loss** (unless SIGTERM handler flushes) | N/A (Container destroyed) | N/A | 100% loss of buffered in-memory telemetry upon container shutdown unless application handles `SIGTERM` and invokes `ForceFlush()` before exit. |
---
## 4. Short-Term Mitigations (For Actor Developers)
Until native Substrate lifecycle hooks are introduced, actor application developers should adopt these mitigations:
1. **Use `Full` Snapshot Mode for In-Memory OTel Workloads**:
```yaml
snapshotsConfig:
onPause: Full
onCommit: Full
```
2. **Configure a Shorter Push Export Interval**:
Reduce `sdkmetric.WithInterval()` from the default 60s down to 5s or 10s to ensure metrics and spans are exported promptly during short active execution windows before idle suspension.
3. **Application `SIGTERM` Signal Handling & Flush Pattern (Issue #23)**:
Register a POSIX `SIGTERM` / `SIGINT` signal listener in the application entrypoint. Upon receiving termination signals, invoke `meterProvider.ForceFlush(ctx)` and `tracerProvider.ForceFlush(ctx)` before process termination.
4. **Application Self-Suspend Flush Pattern**:
When an actor requests self-suspension, execute `meterProvider.ForceFlush(ctx)` immediately before invoking the suspend API call.
5. **Persist High-Value Events to `DurableDir` Volumes**:
In `Data` mode, append critical events to a `DurableDir` file on disk, replaying and flushing them upon container startup post-resume.
---
## 5. Proposed Long-Term Substrate Solutions (Complementary Lifecycle Contracts)
To achieve 100% telemetry continuity across all actor lifecycle events, Agent Substrate requires **both** of the following complementary system contracts working in tandem:
### Part 1: Suspend/Resume Hooks (`PreSuspend` & `PostResume` — Issue #450)
Handles the **pause / checkpoint / resume** lifecycle:
- **`PreSuspend` Hook**: Substrate sends an HTTP POST request (`/pre-pause`) or POSIX signal (`SIGUSR1`) to the actor container prior to process checkpoint or container pause in `Data` mode. The actor handler invokes `meterProvider.ForceFlush(ctx)` / `tracerProvider.ForceFlush(ctx)` before returning 200 OK.
- **`PostResume` Hook**: Substrate notifies the actor container post-resume, enabling immediate network connection re-verification or flush triggers.
### Part 2: Graceful Termination Contract (`SIGTERM` Grace Period — Issue #23)
Handles the **deletion / teardown / eviction** lifecycle:
- **Graceful Termination Contract**: In `ateom-gvisor` / `atelet`, guarantee a configurable termination grace period (`terminationGracePeriodSeconds: 30s`) with explicit `SIGTERM` notification prior to `SIGKILL`. This guarantees application OTel SDK signal handlers can execute final flushes to the OTel Collector before container destruction when actors are deleted or garbage-collected.
---
## 6. Proposed Action Items
- [ ] Align on `PreSuspend` hook design in Issue #450 and graceful termination contract in Issue #23.
- [ ] Update Substrate documentation guide in `docs/observability.md` advising `onPause: Full` and `SIGTERM` flush handlers for push-exporter workloads.
- [ ] Provide a reference demo for actor telemetry continuity.
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.