aws / aws/bedrock-agentcore-sdk-python

docs: /ping response requires undocumented `time_of_last_update` field — without it AgentCore silently reaps microVMs even when status is HealthyBusy

Open
#471 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
761
Forks
147
Avg merge
1d 23h
Merged PRs (30d)
7

Description

## Problem

The `/ping` contract documented at
[runtime-long-run.html](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-long-run.html)
and [runtime-troubleshooting.html](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-troubleshooting.html)
shows only:

{"status": "HealthyBusy"}

In practice, the platform's idle reaper requires a second field —
`time_of_last_update` — that the AWS-facing docs never mention.
Without it, `idleRuntimeSessionTimeout` fires at the configured
boundary even while `/ping` returns `HealthyBusy`, silently terminating
microVMs mid-execution.

The official SDK has emitted this field since the first public commit
(`4f5c80d`, 2025-07-15) — see
[`runtime/app.py:612`](https://github.com/aws/bedrock-agentcore-sdk-python/blob/main/src/bedrock_agentcore/runtime/app.py#L612).
A third-party post documents the full schema correctly:
https://eashank16.medium.com/demystifying-the-http-protocol-contract-for-amazon-bedrock-agentcore-runtime-30ae130485b4

## Reproduction (controlled, ~10 min)

1. Build a minimal ARM64 container with two endpoints:
- `/invocations`: spawn a background asyncio task, return `{"status":"accepted"}` immediately
2. `CreateAgentRuntime` with `lifecycleConfiguration: {idleRuntimeSessionTimeout: 60, maxLifetime: 28800}`
3. Invoke once. Observe CloudWatch:
- Pings continue every ~2s returning HealthyBusy
- Container is reaped at exactly +60s

Repeat with the only change being the ping body:

{"status": "HealthyBusy", "time_of_last_update": }

Container now survives indefinitely (verified to 256 s in our test).

## Other affected customers

- re:Post (May 2025): https://repost.aws/questions/QUXrgkw_c_QmC2m9fh-6LIzg/aws-bedrock-agentcore-issues — describes identical silent-restart symptom; AWS-bot answer recommends `HEALTHY_BUSY` pattern but doesn't
mention the field
- `aws-samples/sample-host-openclaw-on-amazon-bedrock-agentcore` PR #36 — merged a fix that returns `{"status": "HealthyBusy"}` without the field, almost certainly still broken under load
- Internal: we hit this on N=4 production agent runs across two AWS accounts before isolating it via the test above

## Asks

1. Update [runtime-long-run.html](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-long-run.html) and
[runtime-troubleshooting.html](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-troubleshooting.html)
to document the full ping schema (both fields, with semantics).
2. Either make the field optional on the platform side (so the
currently-documented contract works) or surface a clear error/warning
when a runtime returns a `/ping` body missing the field.
3. Patch the broken reference implementation in
`aws-samples/sample-host-openclaw-on-amazon-bedrock-agentcore` PR #36.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.