apache / apache/trafficserver

Expose cache directory header metadata as stats (create_time, cycle count)

Open
#12,920 0 comments 0 reactions 1 assignee Claimed by @bryancall View on GitHub
Cache Metrics
Dominant language
C++
Stars
2k
Forks
874
Avg merge
6d 15h
Merged PRs (30d)
46

Description

## Problem

The cache currently has a single wrap-related stat: `proxy.process.cache.wrap_count` (and its per-volume variant `proxy.process.cache.volume_N.wrap_count`). This is a runtime counter that increments each time a stripe wraps around and resets to zero on process restart.

The problem is that this stat alone doesn't tell you much that's operationally useful:

1. **You can't determine total wrap history.** After a restart, `wrap_count` is zero. You have no idea if the cache has wrapped 0 times or 500 times. The only way to find out is to run `traffic_cache_tool` and inspect the directory header, which requires stopping traffic or SSHing into the box.

2. **You can't determine cache age.** There's no stat for when the cache was created. Combined with the lack of persistent wrap count, you can't compute wrap frequency (e.g., "this cache wraps every 4 hours" vs "every 4 days").

3. **You can't tell if a cache has ever wrapped.** On a fresh deployment or after a cache clear, there's no way to know from stats alone whether the cache has filled up and started overwriting old content. This matters for capacity planning -- if your caches are wrapping frequently, you may need more disk.

4. **The `Note()` log message is minimal.** When a wrap occurs, the log says:
```
Cache volume 1 on disk '/dev/sda' wraps around
```
No cycle count, no age information. To correlate wrap frequency you'd have to parse timestamps from syslog and count occurrences.

### Real-world scenario

Consider a fleet of proxy servers with 500GB cache disks. You want to answer: "How often are our caches cycling through?" Today, the only way is to either:
- Watch `wrap_count` in real-time and hope ATS doesn't restart during your observation window
- SSH into each host and run `traffic_cache_tool` to read the directory header

Neither scales to a fleet. If `cycle` and `create_time` were exposed as stats, you could query your metrics system and immediately compute wrap frequency across every host.

## The directory header already has the data

The `StripeHeaderFooter` structure persists all of this to disk:

```cpp
struct StripeHeaderFooter {
// ...
time_t create_time; // when the stripe was initialized
uint32_t cycle; // total wrap count (incremented on each wrap, persisted)
uint32_t phase; // toggled on wrap
// ...
};
```

The `cycle` field is already used internally -- `cache_bytes_used()` checks `cycle` to determine whether the stripe has ever wrapped (cycle 0 means it hasn't filled up yet, so bytes used = write_pos - start; otherwise the stripe is full).

The `create_time` is set when the stripe is initialized and never changes.

Neither value is exposed as a metric.

## Proposal

### Add two new gauge stats

| Stat | Type | Source | Meaning |
|------|------|--------|---------|
| `proxy.process.cache.directory.cycle` | Gauge | `header->cycle` | Total historical wrap count (persists across restarts) |
| `proxy.process.cache.directory.create_time` | Gauge | `header->create_time` | Stripe creation time (epoch seconds) |

Per-volume variants:
- `proxy.process.cache.volume_N.directory.cycle`
- `proxy.process.cache.volume_N.directory.create_time`

### Keep existing `wrap_count` unchanged

The existing `wrap_count` counter is still useful for rate-based alerting (wraps/hour). The new gauges serve a different purpose -- persistent state inspection.

### Update in `CachePeriodicMetricsUpdate()`

The existing periodic update function already iterates all stripes every ~5 seconds to compute `bytes_used`. Adding reads of `cycle` and `create_time` from the directory header is trivial -- just a few more lines in the same loop.

For per-volume aggregation:
- `cycle`: sum across stripes in the volume
- `create_time`: minimum across stripes (oldest creation time)

### Improve the wrap log message

Enhance the `Note()` in `agg_wrap()` to include the cycle count and stripe age:

```
Cache volume 1 on disk '/dev/sda' wraps around (cycle 47, created 2025-01-15 08:30:00)
```

This makes log-based analysis much easier without needing to cross-reference metrics.

## Implementation scope

Three files need changes:

- `src/iocore/cache/P_CacheStats.h` -- add `directory_cycle` and `directory_create_time` gauge fields
- `src/iocore/cache/CacheProcessor.cc` -- register new stats, extend `CachePeriodicMetricsUpdate()`
- `src/iocore/cache/StripeSM.cc` -- improve `Note()` log message in `agg_wrap()`

No on-disk format changes. No new config knobs. No behavioral changes to the cache itself.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.