S3 Lifecycle: address review feedback (alert filters, fence tags, weed shell flag, naming)

Pulls the post-review version from seaweedfs/seaweedfs#9491:

- Alerts filter zero-valued gauges (avoids fresh-install false positives)
- LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED documented as proto zero-value
- weed shell -master flag clarified as host:http_port.grpc_port
- Manual cursor edits gated on admin-UI pause of the s3_lifecycle job
- Fenced code blocks tagged with 'text' for markdownlint compliance
- S3 spec rule names used consistently (ExpiredObjectDeleteMarker,
  NewerNoncurrentVersions, AbortIncompleteMultipartUpload)
Chris Lu committed 2026-05-13 15:28:46 -07:00
1 parent 29aa6c5c5f
commit 9dc49c9596
5 files changed
+24 -22

No files matched your search

+4 -4
@@ -6,7 +6,7 @@ High-level overview of the lifecycle worker. For implementation detail, see [`we
The lifecycle worker runs as a scheduled job. Each invocation:
```
```text
┌──────────────────────────────────────────┐
│ dailyrun.Run (one filer subscription) │
│ │
@@ -34,8 +34,8 @@ Once every shard's goroutine returns, the worker tears down the subscription, em
Each shard owns a cursor file on the filer at `/etc/s3/lifecycle/daily-cursors/shard-NN.json`:
```
TsNs — last meta-log event whose matches all dispatched successfully
```text
TsNs — last meta-log event for which all matches dispatched successfully
RuleSetHash — ReplayContentHash of the rule set when this cursor was written
PromotedHash — PromotedHash(retentionWindow) at write time
LastWalkedNs — wall-clock of the last successful walker fire
@@ -79,7 +79,7 @@ The walker throttle decouples walker firing from invocation rate. CI invokes the
Cluster-wide cap allocated per worker at job dispatch:
```
```text
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
```
+8 -6
@@ -10,7 +10,7 @@ All labels are in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem.
| Metric | Labels | What |
|---|---|---|
| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event whose matches all dispatched successfully on this shard |
| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event on this shard for which all matches dispatched successfully |
| `s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire |
Derived queries:
@@ -37,7 +37,7 @@ Zero values mean "not started yet" — distinct from "0s caught up". The heartbe
| `s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter |
| `s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path |
`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run.
`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run. `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED` is the proto zero-value — a healthy worker / server pair should never emit it; a non-zero count there indicates an internal error or a version mismatch between worker and server.
### Histograms
@@ -50,7 +50,7 @@ Zero values mean "not started yet" — distinct from "0s caught up". The heartbe
Emitted at the end of every `dailyrun.Run` invocation, at `glog.V(0)` (default verbosity):
```
```text
daily_run: status=ok shards=16 errors=0 duration=7s cursor_lag_max=2m walked_max_age=3m
```
@@ -67,7 +67,7 @@ Tokens are space-separated `key=value` for grep / log-aggregator filtering. Stab
A healthy production heartbeat looks like:
```
```text
daily_run: status=ok shards=16 errors=0 duration=12.3s cursor_lag_max=45s walked_max_age=58m
```
@@ -87,15 +87,17 @@ Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag
## Suggested alerts
```yaml
# `> 0` filters out shards whose gauge is still the proto zero (never
# started); without it, every fresh-install heartbeat triggers the alert.
- alert: S3LifecycleCursorLagHigh
expr: max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 > 3600
expr: max(time() * 1e9 - (s3_lifecycle_cursor_min_ts_ns > 0)) / 1e9 > 3600
for: 30m
annotations:
summary: "S3 lifecycle replay lag > 1h on shard {{ $labels.shard }}"
runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#stuck-cursor
- alert: S3LifecycleWalkerStuck
expr: max(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 > 86400
expr: max(time() * 1e9 - (s3_lifecycle_daily_run_last_walked_ns > 0)) / 1e9 > 86400
for: 1h
annotations:
summary: "S3 lifecycle walker hasn't run in > 24h"
+3 -3
@@ -97,8 +97,8 @@ After applying a rule:
For testing without waiting, the `weed shell` command supports manual invocation:
```
weed shell -master <addr>
```text
weed shell -master <host:http_port.grpc_port>
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
```
@@ -108,7 +108,7 @@ This is exactly what the CI integration suite uses. See [test/s3/lifecycle/](htt
The cluster delete cap is allocated per-worker at job dispatch:
```
```text
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
```
+8 -8
@@ -20,7 +20,7 @@ sum by (outcome) (rate(s3_lifecycle_dispatch_total[5m]))
Look at worker log for the offending event. The dispatcher logs at `glog.V(1)`:
```
```text
daily_run: RETRY_LATER on <bucket>/<key> EXPIRATION_DAYS
daily_run: BLOCKED on <bucket>/<key> NONCURRENT_DAYS
daily_run: transport error on <bucket>/<key> ABORT_MPU: <err>
@@ -35,7 +35,7 @@ daily_run: transport error on <bucket>/<key> ABORT_MPU: <err>
| `BLOCKED SKIPPED_OBJECT_LOCK` | Object is locked (legal hold, retention) | Wait for lock to expire, or remove lock manually. Cursor advances normally — this isn't stuck. |
| `RPC_ERROR` (sustained) | Transport / network issue | Check S3 server health and filer reachability |
The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it, edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Restart the worker.
The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it: first **pause the `s3_lifecycle` job in the admin UI** so the worker isn't running mid-edit, then edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Resume the job. Editing while the worker is active would race with the worker's own save and either lose your change or overwrite the persisted progress.
## Walker stuck (no progress on walker-only rules)
@@ -66,8 +66,8 @@ For testing, the in-repo integration suite uses a trick: backdate the entry's `M
For ad-hoc verification, invoke the worker manually:
```
weed shell -master <addr>
```text
weed shell -master <host:http_port.grpc_port>
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
```
@@ -98,8 +98,8 @@ This runs the same code path as the scheduled worker, but driven from your shell
Cursors live at `/etc/s3/lifecycle/daily-cursors/shard-NN.json` on the filer. Read with the filer's `read` API or `weed shell`:
```
weed shell -master <addr>
```text
weed shell -master <host:http_port.grpc_port>
> fs.cat /etc/s3/lifecycle/daily-cursors/shard-00.json
```
@@ -125,8 +125,8 @@ Manually editing the cursor is supported as an escape hatch but obviously breaks
If a shard's cursor is corrupted or wedged in an unrecoverable state:
```
weed shell -master <addr>
```text
weed shell -master <host:http_port.grpc_port>
> fs.rm /etc/s3/lifecycle/daily-cursors/shard-07.json
```
+1 -1
@@ -22,7 +22,7 @@ This page is the operator-facing entry point. Developers and architecture reader
## API endpoints
```
```text
PUT /{bucket}?lifecycle # PutBucketLifecycleConfiguration
GET /{bucket}?lifecycle # GetBucketLifecycleConfiguration
DELETE /{bucket}?lifecycle # DeleteBucketLifecycle